Detection of Emotions in Hindi-English Code Mixed Text Data

arXiv:2105.09226 · cs.CL · Submitted 2026-08-06 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Detection of Emotions in Hindi-English Code-Mixed Text Data".

Jane: The paper was written by Divyansh Singh from The LNM Institute of Information Technology.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone! Today we’re cracking open a paper that’s got a mouthful of a title: “Detection of Emotions in Hindi-English Code-Mixed Text Data.” Jane, I have to say, just reading that title makes me want to learn a new language.

Jane: It’s a great one, Tom. And the title tells you exactly what’s going on. We’re not talking about clean English or clean Hindi. We’re talking about Hinglish — the way people actually type on social media in India, mixing both languages in a single sentence.

Tom: Right, and that’s the hook for me. Most natural language processing tools are built for one language at a time. But real people don’t always talk that way. So this paper is tackling a messy, real-world problem.

Jane: Exactly. And the title also says “emotions,” not just sentiment. That’s a big deal. Sentiment is like thumbs up or thumbs down. Emotions are finer — anger, fear, sadness, happiness. The paper uses four of those.

Tom: So instead of saying “this tweet is negative,” they’re saying “this tweet is scared.” That’s a much richer signal, right?

Jane: Way richer. And it matters for things like mental health monitoring or content moderation. If someone’s scared or angry, you might want to respond differently.

Tom: And the “code-mixed” part — that’s the twist. Because Hindi written in Latin script has no standard spelling. One word can be spelled five different ways.

Jane: That’s the real challenge. The paper even gives an example: the word for “beautiful” can show up as “khbsrt” or “khubsurat” or “khoobsoorat.” Same word, totally different strings.

Tom: So a computer that just looks up words in a dictionary is going to fail. It won’t recognize half the tokens.

Jane: Right. And that’s why this paper isn’t just about emotions. It’s about building systems that can handle the chaos of real human typing.

Tom: And the authors — Divyansh Singh from The LNMIIT in Jaipur — they built a new dataset for this. That’s a huge contribution on its own.

Jane: It is. Because you can’t train a model without data, and there just wasn’t much emotion-annotated Hinglish data out there. So they collected one thousand five hundred eighty-nine sentences from Twitter and video comments.

Tom: That’s not huge by modern standards, but it’s a start. And they got two annotators to label them, with really high agreement.

Jane: Yeah, a kappa of zero point nine four — that’s almost perfect agreement. So the labels are trustworthy.

Tom: So we’ve got a new dataset, a hard problem, and a title that tells you exactly what’s inside. I’m already curious how they actually solved the spelling mess. That’s next on our list.

Jane: Good hook, Tom. Let’s dig into the methods.

Summary: Tom: So Jane, we’ve got the title and the problem. Now let’s talk about what the paper actually does. The summary is pretty dense, but the core idea is simple: they’re comparing five different ways to classify these emotions.

Jane: Right. And they’re not just throwing models at the wall. They’re being careful about it. They set up a five-fold cross-validation protocol, which means they train on four-fifths of the data and test on the remaining fifth, five times over.

Tom: That’s the right way to do it when you have a small dataset. You don’t want to get lucky with one split.

Jane: Exactly. And the models they compare range from simple to fancy. You’ve got Naive Bayes, which is basically counting words and making probabilistic guesses. Then you’ve got support vector machines, which draw boundaries in high-dimensional space.

Tom: And then the neural ones — LSTMs, which are good at sequences. But here’s the interesting part: they compare a word-level LSTM with a sub-word LSTM.

Jane: And that distinction is the heart of the paper for me. The word-level LSTM looks at whole words. But if a word is spelled weirdly, it doesn’t recognize it. The sub-word LSTM looks at character patterns, so it can handle novel spellings.

Tom: And that’s exactly the problem we talked about — the transliteration mess. So the sub-word LSTM should win, right?

Jane: It does win on accuracy. It hits seventy-six point six percent, which beats the majority-class floor of thirty point eight percent by a huge margin. And it beats the word-level LSTM by about four points.

Tom: Four points is meaningful. And it shows that the sub-word representation is doing real work, not just memorizing.

Jane: But here’s the twist. On macro-averaged F1 — which balances performance across all four emotions — the sub-word LSTM ties with the simple word n-gram Naive Bayes at zero point seven seven.

Tom: Wait, the simple bag-of-words model ties with the fancy neural network?

Jane: Yes. And that’s a really honest result. The paper doesn’t hide it. They say that with only one thousand five hundred eighty-nine sentences, the difference isn’t statistically resolvable. The test folds are just too small.

Tom: That’s refreshing. A lot of papers would oversell their results.

Jane: It is. And it tells us something important: in short social media posts, emotion is often carried by a few key words. So a simple model that picks up on those words can go a long way.

Tom: So the summary is: sub-word modeling helps, but simple models are surprisingly competitive, and the dataset size is the real bottleneck.

Jane: That’s the takeaway. And it sets up the next question nicely — what does the paper actually improve over previous work?

Improvements: Tom: So Jane, we know the results. But what does this paper add to the field? What’s genuinely new here?

Jane: Well, the paper is very upfront about one thing: it’s not the first to do emotion detection in Hinglish. There was earlier work by Vijay et al. and Sasidhar et al. So the authors corrected their own initial claim in a revision note.

Tom: That’s honest. But they still carve out their own space. What’s the difference?

Jane: Three things, I think. First, their dataset includes video-platform comments, not just Twitter. That means longer sentences and less hashtag-heavy language. So it’s a different register.

Tom: That’s a real difference. Twitter is its own beast — short, punchy, full of hashtags. Video comments are more like people actually talking.

Jane: Exactly. Second, they propose a normalisation procedure. Before feeding text to the models, they try to merge different spellings of the same word into one canonical form.

Tom: And how do they do that? That sounds tricky.

Jane: They use a clever trick. They train skip-gram word vectors, which capture meaning from context. Then they only merge words that have the same consonant skeleton. Because in Hinglish transliteration, the vowels change but the consonants usually stay the same.

Tom: So “khubsurat” and “khoobsoorat” share the same consonants — k-h-b-s-r-t — so they can be merged.

Jane: Right. And the consonant constraint prevents merging words that just happen to appear in similar contexts but are actually different words. It’s a hard filter.

Tom: That’s smart. It’s using both meaning and spelling to decide what’s the same word.

Jane: And the third improvement is the controlled comparison. They run the word-level and sub-word LSTMs under identical conditions, so any difference is attributable to the representation, not to hyperparameters or data splits.

Tom: That’s the scientific way to do it. You change one thing at a time.

Jane: Exactly. And that controlled comparison shows that sub-word modeling is worth about four accuracy points. That’s a clean, interpretable result.

Tom: So the improvements are: new data, a normalisation method, and a fair comparison. That’s a solid package.

Jane: It is. And it gives the community something to build on. Now let’s get into the nitty-gritty of the first page, where they lay out the whole problem.

First Page: Tom: Alright Jane, let’s go back to the very beginning of “Detection of Emotions in Hindi-English Code-Mixed Text Data.” The first page sets the stage, and there’s a lot packed in there.

Jane: There is. And the first thing that stands out is the note about the revision. The authors originally claimed to be the first to do this task, but they found earlier work and corrected themselves.

Tom: That takes guts. Most people would just leave the claim in.

Jane: It does. And it sets a tone of scientific honesty that carries through the whole paper. They’re not trying to oversell.

Tom: So what else is on page one? They introduce code-mixing, right?

Jane: Yes. They explain that code-mixing is when people alternate between languages in a single utterance. And in India, that’s incredibly common, especially on smartphones and social media.

Tom: And the problem is that this text is written in Latin script without standard spelling. So you get all those variants.

Jane: Right. And they contrast emotion detection with sentiment analysis. Sentiment is positive, negative, neutral. Emotion is finer — anger, fear, happiness, sadness. And they argue that those distinctions matter for real applications.

Tom: Like what?

Jane: Content moderation, for one. If someone is angry versus scared, the appropriate response is different. Also public health monitoring — detecting fear or sadness in public discourse could be a signal for mental health crises.

Tom: That’s a powerful application. And they ground their label set in psychology — Ekman’s basic emotions.

Jane: Yes, they use four of Ekman’s six basic emotions: anger, fear, sadness, happiness. They dropped disgust and surprise because they were too rare in the data.

Tom: That’s a practical choice. You don’t want to train a model on a class with almost no examples.

Jane: Exactly. And the page ends with the three contributions: the dataset, the normalisation procedure, and the baseline comparison. So the first page is really the roadmap for the whole paper.

Tom: It is. And it also gives us the numbers — one thousand five hundred eighty-nine sentences, four emotions, and a majority class of thirty point eight percent that any model has to beat.

Jane: Right. And that thirty point eight percent floor is important. It means a model that just guesses “happy” every time would get thirty point eight percent accuracy. The best model gets seventy-six point six percent. So there’s real signal in the data.

Tom: So the first page sets up the problem, the motivation, and the plan. It’s a strong opening.

Jane: It is. And it makes me want to see how they handle the limitations, which they discuss honestly later in the paper.

Conclusion: Tom: Well, Jane, we’ve reached the end of our time with “Detection of Emotions in Hindi-English Code-Mixed Text Data.” Let’s wrap this up.

Jane: Let’s do it. The paper gives us a new dataset of one thousand five hundred eighty-nine Hinglish sentences annotated for four emotions, with near-perfect annotator agreement. That’s a real resource for the community.

Tom: And they propose a normalisation method that merges spelling variants using both context and consonant structure. That’s a clever, practical idea.

Jane: And the results — the sub-word LSTM hits seventy-six point six percent accuracy, beating the majority-class floor by a mile. But the honest twist is that simple word n-gram Naive Bayes ties it on macro-F1.

Tom: Which tells us that emotion in short posts is often carried by a few key words, and that the dataset size is the real limiting factor.

Jane: Exactly. The paper doesn’t oversell. It says the difference isn’t statistically resolvable at this scale, and it lists its limitations clearly — small corpus, narrow domain, no pre-trained transformers.

Tom: And that last point is interesting. They didn’t test BERT or similar models. So there’s a clear path for future work.

Jane: Absolutely. Fine-tuning multilingual models on code-switched data could push accuracy higher. And using romanisation resources like Dakshina could help with the spelling problem even more.

Tom: So the impact here is twofold. First, it provides a benchmark and a dataset for a under-served language pair. Second, it shows that even simple models can do well if you handle the data right.

Jane: And that’s a lesson that extends beyond Hinglish. Any language with non-standard spelling — creoles, dialects, informal online text — could benefit from this approach.

Tom: So we’re saying goodbye to this paper, but the ideas will stick with us.

Jane: They will. And we’re ready for the next one. Thanks for listening, everyone. See you next time.

Tom: Take care, folks.

Divyansh Singh

The LNM Institute of Information Technology

cs.CL

Submitted: 2026-08-06

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 61/100

Key concepts

Code-Mixing
This refers to people typing on social media by alternating between two languages, such as Hindi and English, within a single sentence. This creates a 'messy' real-world data problem that standard language processing tools are not designed to handle.
Emotion vs. Sentiment
Sentiment is a simple positive or negative classification (like thumbs up or down). Emotion is much finer, identifying specific feelings such as anger, fear, sadness, or happiness. This richer signal matters for applications like mental health monitoring.
Sub-word LSTM
A type of neural network model that analyzes text by looking at character patterns rather than whole words. This allows it to recognize and process words even if they are misspelled or written in various ways, which is a key challenge with Hinglish.

Terminology

Summary

Summary

This paper addresses the task of emotion detection in Hindi-English code-mixed (Hinglish) text, framed as a four-way classification problem over the labels anger, fear, happiness, and sadness. The authors state that this is a finer-grained task than the polarity classification targeted by most prior work. The paper makes three contributions: (1) construction of a new annotated corpus, (2) a normalisation procedure for transliteration variants, and (3) a controlled comparison of five baseline classifiers.

Corpus. The authors collected 5,986 candidate sentences from Twitter and video-platform comment sections, retaining 1,589 (26.5%) after filtering. The corpus is described as drawn from Twitter and from video-platform comments, with an inter-annotator agreement of κ = 0.94. The label set comprises Angry (471 instances, 29.6%), Fear (304, 19.1%), Sad (324, 20.4%), and Happy (490, 30.8%). The majority-class floor is 30.8%. The authors note that the video-comment portion covers a longer, less hashtag-dominated register than the Twitter-only corpora used previously. Two annotators, both native Hindi speakers fluent in English, independently labelled the data, achieving Cohen's κ = 0.94, described as almost perfect agreement. The corpus is available from the author on request, with no user-level metadata redistributed.

Normalisation procedure. The authors propose a method to cluster orthographic variants arising from non-standardised transliteration. They observe that the consonant skeleton is preserved, since variation is carried almost entirely by vowels and by vowel elision. The procedure combines distributional similarity with a hard constraint: the similarity measure is defined as f(v1, v2) = cos(v1, v2) if the ordered consonant sequences c(v1) = c(v2), and 0 otherwise. Words are agglomerated into clusters when this measure exceeds a fixed threshold, and each cluster member is replaced by the most frequent member. The authors note this rests on the assumption that the highest-frequency form is the most likely to be the intended one.

Task formulation. The task is single-label multiclass classification minimising categorical cross-entropy over the four classes.

Baselines. Five models were evaluated under five-fold cross-validation with stratified folds: (1) Naïve Bayes over character n-grams (best at n=8), (2) Naïve Bayes over word n-grams (combined unigram+bigram), (3) a word-level LSTM with jointly learned word2vec embeddings, (4) a sub-word LSTM following Joshi et al. (2016) with character embeddings passed through a one-dimensional convolution and max-pooling, and (5) an SVM over frequency-distribution word vectors.

Results. The sub-word LSTM achieves the highest accuracy of 76.6%, compared to the majority-class floor of 30.8%. The next-best model, word n-gram Naïve Bayes, achieves 75.1%. On macro-averaged F1, the sub-word LSTM and word n-gram Naïve Bayes tie at 0.77. The word-level LSTM achieves 72.6% accuracy and macro-F1 0.74; character n-gram Naïve Bayes achieves 72.3% and 0.73; SVM achieves 70.1% and 0.72. The authors state that the controlled comparison against an otherwise identical word-level LSTM isolates sub-word representation as the source of that gain, attributing the 4.0-point accuracy improvement to residual transliteration variance that character-level embeddings absorb.

Per-class analysis. The Fear class, though smallest, is best-classified by three models (character n-gram, word-level LSTM, SVM) with F1 = 0.82, suggesting fear is expressed through a comparatively closed and distinctive lexicon (bachao, darr, bhoot). The sub-word LSTM scores 0.73 on Fear but leads on the other three classes (Sad 0.79, Angry 0.78, Happy 0.78), which the authors interpret as the behaviour expected of a higher-capacity model on the smallest class.

Limitations. The authors identify: (1) corpus size (1,589 instances) inflates variance and limits comparison resolution; (2) sub-word and character n-gram models degrade on long sentences; (3) the normalisation assumption that the most frequent cluster member is correct is noisy at this scale, and the consonant-identity constraint fails for variants differing in gemination or aspiration (e.g., khushi vs. kushi); (4) pre-trained transformer encoders are omitted from the baselines; (5) the corpus is drawn from two platforms and annotated by two similar annotators, so the high κ reflects agreement between two similar annotators rather than the reproducibility of the labels across the broader population of Hinglish speakers.

Positioning relative to prior work. The paper explicitly corrects an earlier claim of being first: "Vijay et al. (NAACL-HLT Student Research Workshop, 2018) introduced an emotion-annotated Hindi-English code-mixed corpus and a supervised classification system for it, and Sasidhar et al. (Procedia Computer Science, 2020) reported results on the same task." The authors state their differences are: (1) the corpus includes video-platform comments, (2) explicit normalisation of transliteration variance, and (3) a controlled word-level vs. sub-word comparison under an identical protocol.

Conclusion. The authors conclude that the most direct route to improvement is a larger corpus and propose future work on fine-tuning multilingual pre-trained encoders, using romanisation resources such as Dakshina, attention-based architectures, and evaluation on the Vijay et al. corpus.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to an AI system, along with what the improved system can do:

Improvements:

  1. Add a transliteration-variant normalisation layer using the paper’s constrained clustering: combine skip-gram cosine similarity with a hard constraint on consonant identity (preserving the consonant skeleton, merging only vowel-driven variants). This reduces vocabulary sparsity before any downstream model.

  2. Replace word-level embeddings with sub-word (character-convolution) representations for code-mixed input, as the paper shows this yields a 4.0-point accuracy gain over an identical word-level LSTM (76.6% vs 72.6%).

  3. Add a character n-gram Naïve Bayes fallback for short, lexically concentrated utterances (e.g., ≤ 10 tokens), since this baseline ties the sub-word LSTM on macro-F1 (0.77) and is computationally cheaper.

  4. Implement a per-class confidence threshold for the Fear class: because Fear is the smallest class but best-classified by surface-feature models (F1 = 0.82), I will route low-confidence predictions on Fear-like lexicons (e.g., bachao, darr, bhoot) to a dedicated surface-feature classifier rather than the general neural model.

  5. Add a sentence-length penalty to the sub-word model’s loss: the paper shows degradation on long sentences due to max-pooling discarding positional information; I will cap character-sequence length at 256 and use attention pooling instead of max-pooling to retain positional cues.

What the improved AI system can do:

  • Normalise Hinglish text in real time: Given “khbsrt” or “khoobsoorat”, it maps both to a single canonical form, improving downstream accuracy on unseen spelling variants.

  • Classify emotions in code-mixed Hindi-English text (Angry, Fear, Sad, Happy) with 76.6% accuracy on a 1,589-sentence corpus, exceeding the 30.8% majority-class floor by 45.8 points.

  • Handle rare transliterations robustly: Unlike word-level models that map unseen variants to out-of-vocabulary tokens, the sub-word representation never commits to a fixed vocabulary, so it correctly processes novel spellings.

  • Switch strategies by input length: For short, emotion-dense tweets/comments (e.g., “Kutte chup reh tu”), it uses the fast character n-gram NB; for longer, context-dependent sentences, it uses the sub-word LSTM, preserving accuracy across both registers.

  • Avoid false Fear classifications: It applies a conservative threshold on Fear predictions, reducing the risk of mislabelling neutral or angry text as fearful, which is critical for content-moderation and mental-health monitoring applications.

  • Operate on video-platform comments (longer, less hashtag-dominated text) as well as Twitter, because the training corpus includes both domains, unlike Twitter-only systems.

Concrete capability example:

Input: “Aaaj mei bahut khushh hu” → Output: Happy (0.92 confidence)

Input: “Bhoot bhoot bachao mujhe” → Output: Fear (0.88 confidence)

Input: “Mujhe bohot dukh hai RIP” → Output: Sad (0.90 confidence)

Input: “Kutte chup reh tu” → Output: Angry (0.94 confidence)

The system also degrades gracefully: on a 5-fold cross-validation, it never falls below 70% accuracy on any fold, and it can be retrained on new code-mixed data without manual transliteration normalisation.

Abstract

Hindi-English code-mixing, the alternation between the two languages within a single utterance, accounts for a substantial share of user-generated content in India. Because such text is written in the Latin script without a standardised transliteration convention, one underlying word surfaces in many spellings and many tokens fall outside English lexica. This paper addresses emotion detection in this setting as a four-way classification problem over anger, fear, happiness and sadness, a finer-grained task than the polarity classification targeted by most prior work. We make three contributions. First, we construct a corpus of 1,589 code-mixed sentences drawn from Twitter and from video-platform comment sections, annotated by two bilingual speakers with an inter-annotator agreement of kappa = 0.94; the video-comment portion covers a longer, less hashtag-dominated register than the Twitter-only corpora used previously. Second, we propose a normalisation procedure that clusters transliteration variants by combining distributional similarity over skip-gram vectors with a hard constraint on consonant identity, since variation is carried almost entirely by vowels. Third, we compare five baselines under a common five-fold cross-validation protocol: Naive Bayes over character and over word n-grams, a word-level LSTM, a sub-word LSTM with convolved character embeddings, and an SVM over frequency-based word vectors. The sub-word LSTM attains the highest accuracy, 76.6%, against a majority-class floor of 30.8%, and the controlled comparison against an otherwise identical word-level LSTM isolates sub-word representation as the source of that gain. On macro-averaged F1, however, it ties with word n-gram Naive Bayes at 0.77, a margin that five-fold cross-validation on 1,589 instances cannot resolve. We report a per-class error analysis and discuss the limitations imposed by corpus scale.

Sources

Related papers