Chehre: An Emoji-Prompted Dataset to Explore Perceptual Flexibility in Video Language Models

arXiv:2606.21657 · cs.CV, cs.CL · Submitted 2026-06-19 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Chehre: An Emoji-Prompted Dataset to Explore Perceptual Flexibility in Video Language Models".

Jane: The paper was written by Bita Azari, Zoe Stanley, Avneet Batra, Poorvi Bhatia and Hali Kil, Manolis Savva, Angelica Lim from Simon Fraser University, Canada.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Summary and Core Findings: Tom: So, let’s dive into the summary of this research, as it lays out the scale of what they built and the initial sobering results.

Jane: They collected two thousand one hundred eleven high-quality videos from over two hundred performers using a process that was highly controlled, which gives us a massive amount of confidence in the data quality.

Meng: And as we discussed before, they used reenactment technology to anonymize everything by mapping those real human movements onto synthetic faces before the AI sees them.

Lu: That blending of actual human emotion with a standardized synthetic face is such a powerful way to control for variables while still capturing genuine emotional intent.

Lalam: The summary shows that this massive dataset was meticulously validated by nine hundred two annotators, which gives us incredible confidence in the data quality.

Tom: But the core finding, as they present it, is pretty sobering; they found that even with this diverse data, the best vision-language models still struggled significantly.

Jane: The thirty-two point five percent top-one accuracy on dominant expression recognition shows us that if we just ask a single model to pick the most likely label based on human consensus, it often gets it wrong.

Meng: And when you look at the "False None" rate—where models confidently choose "None" even though humans see a clear expression—that represents a massive practical failure for AI.

Lu: It suggests that current AI is very conservative and hesitant to make any claim about what it sees in a dynamic visual signal, which points to limitations in reasoning.

Lalam: The dataset highlights the fundamental difficulty of reliably extracting nonverbal social signals because it captures the subtle nuances of human interaction.

Improving the Approach: Tom: We’ve seen how they built this massive dataset; now, let's talk about how they approach the problem by introducing "Chehre" and its specific methodology.

Jane: The paper introduces two benchmark tasks that are really clever because they move away from assuming a single answer for a video's expression.

Meng: They call them "Dominant Expression Recognition" and then "Distributional Expression Recognition," which is a huge shift in how we test AI.

Lu: Distributional recognition forces us to look at the entire spectrum of human responses, not just the average consensus, which is incredibly insightful for complex data like facial expressions.

Lalam: It’s about embracing that disagreement—that recognizing diversity *is* the signal—and seeing if we can find a way to capture that in our AI models.

Tom: And this is where they introduce "Persona Prompting," which seems like the next big thing for improving these results, right?

Meng: Persona prompting allows the model to see the video through a specific lens, like an "unfriendly" or "friendly" annotator, which is a tangible way to test different viewpoints.

Jane: By letting us control that observer perspective, we can try to see if we can induce more variety in the AI's predictions than if we just use random sampling.

Lu: It’s a way of injecting human subjectivity into the model's reasoning, which is a fascinating theoretical approach for guiding LLM output toward specific interpretations.

Lalam: It helps us explore how cultural or interpersonal differences in perception might manifest when we ask an AI to adopt specific personas, providing that data about how people see the world.

Practical Implications and Future Work: Tom: We’ve seen the methods; let's talk about what this means for the industry and future work surrounding "Chehre: An Emoji-Prompted Dataset to Explore Perceptual Flexibility in Video Language Models."

Jane: The data is standardized through synthetic faces, which minimizes visual noise, but it also raises questions about generalizing to real-world messy environments where things are much more chaotic.

Meng: I wonder if this approach scales up well—can we apply this dynamic mapping and persona prompting to a much larger, more chaotic dataset without losing the level of quality they achieved in their controlled environment?

Lu: The focus on inter-individual diversity is key here; it’s a microcosm of human variability that applies far beyond just a single controlled study, showing us how unique each person's perception can be.

Lalam: We have to be careful not to mistake what *perceived* emotion is for what the performer *feels*, and this framework helps us frame that distinction clearly for the broader audience in AI development.

Tom: The paper also mentions future work, like using demographic data to create unique personas or exploring different ways people interpret faces based on their own background.

Meng: That sounds like a massive project, but I think it's necessary to understand the full breadth of human perception before we can build truly robust and accountable AI agents.

Lu: It' is about recognizing that the tools we are building must be adaptable and accountable for what they see, not just reliable in a fixed way when faced with human complexity.

Lalam: We need to ensure this benchmark supports future studies on how AI interprets emotion accurately, rather than just providing a single answer that doesn't capture the full picture.

Conclusion: Tom: As we wrap up our discussion of "Chehre: An Emoji-Prompted Dataset to Explore Perceptual Flexibility in Video Language Models," I think the biggest realization is that the goal isn't just accuracy, but capturing human diversity.

Jane: It’s a much more nuanced view of AI performance than simply hitting a percentage score; we are seeing a new standard for what constitutes successful perception.

Lu: The framework forces us to think about perception as a spectrum rather than just discrete boxes, which is vital for progress in AI research and understanding human behavior.

Meng: I'm excited to see how this structure performs when integrated into larger, real-world applications, especially with the personas being so well-defined for everyone else.

Lalam: It offers a powerful way for us to improve cultural intelligence in our models by understanding how people genuinely perceive visual signals.

Lu: I think we've seen a lot of creative potential here; it’s a foundation for new ways to build systems that really feel more human and less mechanical.

Meng: I hope the engineering community can take this framework and turn those diverse ratings into something practical for real-time system deployment in everyday life.

Lalam: It truly helps us move toward building an AI that sees and understands the world with more emotional depth than any previous model.

Tom: Well, what a conversation! Thank you all for sharing your insights today, Lu, Meng, Lalam. We'll be back soon to discuss another fascinating paper from ArXiv!

Simon Fraser University, Canada

cs.CV, cs.CL

Submitted: 2026-06-19

Updated: 2026-09-02

Project page: https://chehre-dataset.github.io

Importance score: 90/100

The gist: The paper introduces "Chehre," an innovative emoji-prompted dataset meticulously designed to advance the study of perceptual flexibility within Video Language Models (VLMs).

Key concepts

Chehre Dataset
A massive, controlled dataset of over 2100 videos from many performers. It was created using reenactment technology to map real human movements onto synthetic faces, allowing for high data quality and variable control.
Dominant Expression Recognition
One of the benchmark tasks introduced in the paper. Instead of relying on a single average consensus, this task helps test how well AI models can interpret video expressions by looking at the full spectrum of human responses.
Persona Prompting
A methodology that allows an AI model to view a video through a specific, controlled perspective (e.g., 'unfriendly' or 'friendly'). This injects human subjectivity into the model's reasoning, testing different viewpoints.
Distributional Expression Recognition
A key benchmark task that shifts testing away from assuming one correct answer. It forces AI to consider the entire range of human responses, which is crucial for accurately interpreting complex nonverbal social signals.

Terminology

Summary

The paper introduces Chehre, an innovative emoji-prompted dataset meticulously designed to advance the study of perceptual flexibility within Video Language Models (VLMs). By systematically combining controlled facial synthesis with structured emotional and interpersonal prompting, Chehre provides a rich resource for evaluating how well advanced AI models can interpret subtle, context-dependent emotional cues from video data. This work is crucial because it moves beyond simple categorization, forcing models to reconcile explicit visual evidence with abstract linguistic and social prompts.

Dataset Generation and Synthesis

The dataset leverages sophisticated generative adversarial networks (GANs) to create highly controlled synthetic face samples. The methodology begins by utilizing two source faces drawn from the Chicago Face Database (Ma et al., 2015). To generate the composite material, the eyes of one face are patched onto the remaining regions of another face. These patched composite faces are then projected into a latent space using StyleGAN2 (Karras et al., 2020). The resulting synthetic face is generated by finding and utilizing the nearest latent representation within this constrained space. This process ensures that the visual inputs for the dataset maintain high fidelity while allowing researchers to manipulate specific facial components independently.

Prompt Engineering via Interpersonal Circumplex

To provide linguistic context, the study employs a structured prompting system inspired by the Interpersonal Circumplex (Leary, 2004). This framework defines personas based on two axes: affiliation (friendliness) and dominance status. The friendliness options range from very unfriendly to very friendly, while dominance status ranges from very submissive to very dominant. The prompt construction follows strict rules to ensure comprehensive coverage of social dynamics:

  • If friendliness is neutral and dominance status is balanced, no new sentence is added to the prompt.

  • If friendliness is neutral, the prompt structure becomes: You are someone who is [dominance status].

  • If dominance status is balanced, the prompt structure becomes: You are someone who is [friendliness] toward people.

  • Otherwise, the full format applies: You are someone who is [friendliness] toward people and is [dominance status].

Emoji Labeling and Ambiguity Analysis

The emotional labeling component of Chehre utilizes a comprehensive set of 40 emojis, cataloged in Table 4. Each emoji code is associated with a specific set of candidate labels, allowing for granular analysis of emotional nuance. The dataset quantifies the ambiguity inherent in these labels by tracking video counts across two subsets: high-ambiguity and low-ambiguity. For instance, the emoji 1f624 (Smiling) has 59 videos in total, with a distribution of 15 videos in the high-ambiguity subset and 44 in the low-ambiguity subset. This quantitative approach allows researchers to pinpoint which emotional expressions—such as those associated with Awkward, Excited, Happy—are most challenging for VLMs to interpret accurately.

Exploring Perceptual Flexibility

The final dataset structure is designed not merely for classification but for exploring perceptual flexibility. The inclusion of diverse and sometimes conflicting cues—such as a visual face generated by StyleGAN paired with a prompt describing an unfriendly persona—challenges the model's ability to reconcile multimodal inputs. The researchers aim to test if VLMs can maintain consistency when interpreting subtle emotional shifts, thereby advancing the understanding of how humans and machines process complex social communication embedded within video media.

Improvements for AI systems

Suggested AI System Improvements and Capabilities

Based on the current methodology, which relies heavily on discrete labeling schemes (Table 4) and pre-defined persona matrices, the system can be significantly advanced by implementing advancements in continuous modeling, multimodal fusion, and causal prediction.

Improvement: Replace the current reliance on discrete label sets (e.g., Angry, Excited, Happy) with a latent space representation of emotion that maps onto established psychological continua (e.g., Valence-Arousal-Dominance, or dimensional models like Russell's Circumplex Model).

Capability: The improved system will no longer merely classify an expression; it will quantify the emotional state as a precise coordinate within a continuous space. This allows for the detection of nuanced blends (e.g., Mildly anxious excitement) that are currently forced into single, potentially inaccurate labels like Awkward or Excited. It can provide quantifiable metrics (e.g., Valence score of +0.6, Arousal score of-0.3) rather than just a label tag.

Improvement: Develop a unified Transformer-based fusion module capable of simultaneously processing and weighting inputs from multiple modalities: video (facial action units/geometry), audio (prosody, pitch contour, speaking rate), and textual context (semantic embedding).

Capability: The system will achieve holistic emotional understanding by resolving conflicts or ambiguities across modalities. For example, if the facial expression registers as Neutral but the vocal tone exhibits high pitch variability and rapid speech (indicating excitement/anxiety), the fusion module can assign a weighted composite score, overriding the purely visual reading. This vastly improves robustness in real-world, noisy communication environments.

Improvement: Evolve the static persona matrix into a dynamic state-tracking model that incorporates temporal dependencies and conversational history (Dialogue State Tracking). The system must model not just who the person is (the assigned persona), but how their emotional state shifts in response to dialogue turns or environmental stimuli.

Capability: The system can predict the likely emotional trajectory of an individual within a conversation. Instead of stating, You are someone who is [friendliness] and is [dominance status], it will output a probability distribution predicting the shift: Given the previous statement and positive feedback loop, Person A's dominance status is predicted to increase by 15% in the next utterance. This moves from description to predictive social modeling.

Improvement: Integrate a causal inference layer (e.g., using Structural Causal Models or Granger Causality testing on time-series data) atop the existing classification framework. This module must learn the directional relationships between emotional states, observed cues, and subsequent actions/utterances.

Capability: The system will move beyond reactive analysis to proactive prediction. Given a sequence of inputs (e.g., Speaker A exhibits mild disappointment to Speaker B responds with exaggerated sympathy), the model can predict:

  1. The probable emotional state of the next speaker (P(Emotion t+1 Inputs t)).

  2. The most appropriate empathetic or corrective response required to de-escalate negative emotional trajectories or enhance positive engagement.

Sources

Related papers