← Back to VOLUME 15, ISSUE 10, OCTOBER 2026
This work is licensed under a Creative Commons Attribution 4.0 International License.
AI-Based Semantic Understanding of Code Mixed Tamil-English Text in Informal Digital Communication
Aditya Raman, Sushmita Kundu, Aditya Balaji, G.Paavai Anand
👁 4 views📥 3 downloads
Abstract: Tamil-English code-mixing, popularly known as “Tanglish,” is a pervasive mode of informal digital communication among Tamil speakers, frequently rendered in Romanized script across social media and messaging platforms. Existing natural language processing (NLP) systems, largely trained on monolingual or formally structured corpora, are not optimized to interpret the semantic, sentiment, and intent-level nuances of such hybrid, non-standardized text. This paper proposes a comparative research framework to evaluate the capability of different computational approaches—traditional NLP techniques, BERT-based transformer models, and large language models (LLMs)—in understanding Tamil-English code-mixed text, alongside a proposed preprocessing pipeline designed to normalize Romanized Tamil variants prior to model inference. We outline a methodology for constructing an annotated code-mixed dataset, detail the architecture of each comparative model, and define evaluation metrics spanning accuracy, F1-score, and semantic similarity. The study aims to establish whether targeted preprocessing can meaningfully improve semantic fidelity for low-resource, code-mixed language understanding, contributing toward AI systems that better preserve and process minority and regional language expression in digital spaces.
Keywords: code-mixing, Tanglish, Tamil-English NLP, Romanized Tamil, semantic preservation, low-resource language processing, transformer models
Keywords: code-mixing, Tanglish, Tamil-English NLP, Romanized Tamil, semantic preservation, low-resource language processing, transformer models
How to Cite:
[1] Aditya Raman, Sushmita Kundu, Aditya Balaji, G.Paavai Anand, “AI-Based Semantic Understanding of Code Mixed Tamil-English Text in Informal Digital Communication,” International Journal of Advanced Research in Computer and Communication Engineering (IJARCCE), DOI: 10.17148/IJARCCE.2026.151005
