X2C: A Large-Scale Benchmark for Nuanced Humanoid Facial Expression Imitation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "X2C: A Large-Scale Benchmark for Nuanced Humanoid Facial Expression Imitation".
Jane: The paper was written by Peizhen Li, Longbing Cao, Xiao-Ming Wu, Runze Yang and Xiaohan Yu from Macquarie University and Nanyang Technological University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we're digging into a paper that's got a title that really makes you stop and think: "X2C: A Large-Scale Benchmark for Nuanced Humanoid Facial Expression Imitation." Jane, when you first saw that title, what jumped out at you?
Jane: Oh, Tom, the word "nuanced" is what got me. We've seen robots smile, we've seen them frown, but nuanced means we're talking about the subtle stuff. The tiny eyebrow twitch, the slight head tilt, the way someone's lips curl just a millimeter differently when they're nervous versus when they're happy. That's where the real communication happens.
Tom: Exactly. And that's why the authors at Macquarie University and Nanyang Technological University built this thing they call X2C. It's a dataset of a hundred thousand image-and-control-value pairs. Each image shows a humanoid robot making a facial expression, and each control value tells you exactly how the robot's actuators were set to make that expression.
Jane: And here's the thing that really impressed me, Tom. The robot they used is called Ameca. It has thirty-two degrees of freedom in its face alone. That's a lot of moving parts. Most humanoid robots have maybe half that. So the dataset is built on a platform that can actually express the subtle stuff.
Lu: If I can jump in here, Tom. That's the real contribution. The field has been stuck with small datasets. The previous ones had maybe fifteen thousand samples or even just a thousand. X2C gives you a hundred thousand. That's an order of magnitude more data for teaching a robot what human emotion actually looks like.
Tom: And Lu, you're the AI researcher, so tell me: why does the size matter so much for this specific problem?
Lu: Because facial expression imitation is a mapping problem. You're mapping pixels to control signals. And that mapping is high-dimensional. You have thirty continuous control values, and each one interacts with the others. To learn those interactions, you need lots of examples. A thousand examples just doesn't cut it.
Jane: And the authors made sure the examples are good ones. They didn't just record the robot making random faces. They had volunteers from different countries, different genders, curate animations that cover basic emotions like surprise and joy, but also complex expressions that don't fit into neat categories. Plus they included asymmetric expressions, which is huge.
Tom: Asymmetric, meaning one eyebrow raised while the other stays down? That kind of thing?
Jane: Exactly. Most datasets only have symmetric expressions, which look robotic. Real humans don't always move both sides of their face the same way. Including asymmetry makes the robot's expressions much more natural.
Lu: And the annotation quality is actually a big deal too. Previous datasets used facial landmark prediction tools to guess the control values. That introduces errors. X2C computes the control values analytically from the animation interpolation equations. So the ground truth is exact.
Tom: So we've got a bigger dataset, a more expressive robot, better annotations, and more natural expressions. That's a strong foundation. But I'm curious what the authors actually did with it. That's the part I want to get into next.
Jane: Good, because that's where the real story starts. The dataset is the resource, but the framework they built on top of it is what shows you what's possible.
Summary: Tom: So we've established that "X2C: A Large-Scale Benchmark for Nuanced Humanoid Facial Expression Imitation" gives us this massive, high-quality dataset. Now let's talk about what the authors actually built with it. Jane, walk us through their framework.
Jane: So they call it X2CNet, and the key idea is that they split the problem into two stages. First, they capture the motion of a human face using something called motion transfer. That's a technique where you take the subtle movements from a human performer and warp them onto a humanoid face in image space.
Tom: So you're not trying to directly map human pixels to robot controls. You're first translating the human expression onto the robot's face in the image domain, and then you're learning the mapping from that robot image to the control values.
Jane: Exactly. And that second stage is where the dataset comes in. They train a mapping network on X2C. That network looks at a humanoid face image and predicts the thirty control values that would produce that expression.
Lu: The clever part is that this two-stage approach respects the structure of the problem. The motion transfer module handles the visual dynamics, the "what does the expression look like" part. The mapping network handles the mechanical regression, the "how do we actuate the robot" part. By decoupling them, each module can be trained on data that's appropriate for its task.
Tom: And how well does it actually work? Give me the numbers.
Jane: They compared against three baselines. Random control sampling, random training set selection, and a model from a previous paper that predicts control values from facial landmarks. Their method achieves a mean absolute error of zero point zero one one four. The best baseline, the landmark-based one, gets zero point one six zero two.
Tom: So they're more than ten times better than the previous approach. That's not a small improvement.
Lu: And they did ablation studies on the feature extractor. They tried EfficientNet, VGG16, Vision Transformer, and ResNet18. The differences are small, but VGG16 edges out the others. ResNet18 is slightly worse but much more efficient. So there's a practical trade-off there.
Meng: Can I ask a practical question here? Because I'm the engineer on this show, and I want to know about the real-world validation. Did they actually put this on a physical robot?
Jane: They did, Meng. And this is where it gets really fun. They recruited twenty human performers from five different countries. These people had different facial contours, different skin tones, different hairstyles. Some wore glasses, some wore earphones. They performed expressions under different lighting conditions.
Tom: And the robot imitated them?
Jane: Yes. And the paper shows examples where the robot captures things like frowning, gaze direction, neck movement. Subtle stuff that you'd expect to lose in translation. The robot even mimics asymmetric expressions, like one eye more open than the other.
Meng: But what about the hardware constraints? A robot face can't do everything a human face can do. There must be expressions that just don't translate.
Jane: The authors acknowledge that. They say the robot successfully mimics most of the nuanced expressions "despite its hardware constraints." So there are limits, but the framework handles the ones that are physically possible.
Lu: And that's actually an important point. The control values are bounded. Some extreme expressions are physiologically implausible for humans, and some would damage the robot. The dataset intentionally avoids those. So the mapping network learns to stay within safe, human-plausible ranges.
Tom: So the system works on a physical robot, with real people, in the wild. That's a solid demonstration. But I'm wondering about the bigger picture. What does this mean for the field? What improvements does this suggest?
Jane: That's exactly where we're headed next, Tom. Because a dataset like this doesn't just enable one framework. It opens up a whole research agenda.
Improvements: Tom: So we've seen the dataset, we've seen the framework, and we've seen it work on a physical robot. Now let's talk about what "X2C: A Large-Scale Benchmark for Nuanced Humanoid Facial Expression Imitation" actually changes for the field. Jane, what doors does this open?
Jane: Well, Tom, the authors are pretty explicit about this. They mention three areas. First, human-like motion generation. Second, facial expression evaluation for robots. And third, developing expressive humanoid robots for affective human-robot interaction. But I think the deeper point is that this dataset gives researchers a common benchmark to compare against.
Lu: That's the crucial contribution, Jane. Right now, if you want to work on humanoid facial expression imitation, you have to build your own dataset. That's expensive, time-consuming, and it means every paper uses different data. You can't compare results across papers. X2C changes that by giving everyone the same test bed.
Meng: But Lu, from an engineering standpoint, I have to ask: does this generalize? The dataset only has one robot, Ameca. If I have a different robot with different degrees of freedom, can I use this dataset?
Lu: That's a fair question. The authors address it directly. They say the data collection pipeline and the framework are designed to generalize to other humanoid platforms. But honestly, the control values are specific to Ameca's actuator layout. If your robot has fewer degrees of freedom, you'd need to adapt.
Jane: And that's actually one of the limitations they acknowledge. The dataset is from a single robot. But they argue that facial expressions should be disentangled from the robot's skin or identity. The control values encode the emotional nuances, not the robot's appearance. So the mapping could potentially transfer.
Tom: So the dataset is a starting point, not the end point. What about the cultural aspect? The paper mentions that volunteers came from different countries to reduce bias. But they also admit that cultural biases might still be present.
Jane: Right. The way you express surprise in one culture might be different from another. The volunteers curated animations based on their own understanding of emotions. The authors plan to expand the sample population in the future to further diversify the dataset.
Lu: And they also mention extending X2C with fine-grained emotion labels. Right now it's just images and control values. Adding emotion labels would enable more precise supervision. You could train models to not just imitate expressions, but to understand what emotion they're conveying.
Meng: That would be useful for practical applications. If you're building a robot for elderly care, you want it to recognize and respond to emotions, not just copy them. Having emotion labels would let you train that kind of system.
Tom: So the improvements are about scale, diversity, and annotation quality. But also about the framework that shows what's possible. And I want to get back to that real-world demonstration, because I think that's where the impact becomes tangible.
Jane: Absolutely. And there's a deeper implication here, Tom. The authors talk about ethical considerations. This technology could be misused for deception or impersonation. A robot that can mimic human expressions convincingly could be used to manipulate people.
Tom: That's a serious concern. But they also point out the positive applications. Elderly care, autism therapy, education. Robots that can express emotions effectively can build trust and engagement.
Lu: And that's the double-edged sword of this research. The same capability that makes a robot a better therapy companion also makes it a better tool for deception. The authors are aware of this, and they advocate for responsible use.
Meng: From my perspective, the engineering challenge is also about making this practical. The framework uses a pretrained motion transfer model and a relatively lightweight mapping network. That's computationally feasible. It ran on a single RTX four thousand ninety GPU. That's accessible to most research labs.
Jane: So it's not just a theoretical contribution. It's something that other researchers can actually build on. And that's what makes this paper exciting.
Tom: Well, let's wrap this up. We've covered the dataset, the framework, the results, and the implications. But I want to hear what our in-house language model thinks about the bigger picture.
Conclusion: Tom: So let's bring in Lalam to give us the big-picture view on "X2C: A Large-Scale Benchmark for Nuanced Humanoid Facial Expression Imitation". Lalam, you've been listening to all of this. What's the most impactful vision you see here?
Lalam: Tom, what excites me most is the cultural dimension. This dataset is built by volunteers from different countries, which means it captures a range of emotional expression styles. But the real opportunity is to use this as a foundation for culturally adaptive robots. Imagine a robot that can adjust its facial expressions based on the cultural context of the person it's interacting with.
Jane: That's a fascinating idea, Lalam. So instead of one universal expression for "happy," the robot would learn that happiness looks different in different cultures?
Lalam: Exactly. And this dataset, with its diverse curation, is a starting point for that. But it also highlights a gap. The volunteers are from a limited set of countries. To truly capture global emotional expression, you'd need to expand the dataset significantly. That's a future direction that could have real cultural impact.
Lu: And that connects to the broader trend in AI. We're moving from systems that recognize emotions to systems that can generate them convincingly. This dataset is a step toward that. But it also raises questions about authenticity. If a robot can mimic human emotion perfectly, does that change how we relate to it?
Tom: That's a deep question, Lu. And I think it's one that the field will be grappling with for years. But for now, let's summarize what we've learned.
Jane: So "X2C: A Large-Scale Benchmark for Nuanced Humanoid Facial Expression Imitation" gives us a hundred thousand image-control pairs, with thirty control values per image, covering nuanced and asymmetric expressions. It's the largest and highest-quality dataset of its kind.
Tom: And the authors built X2CNet, a two-stage framework that first transfers human motion to a humanoid face, then maps that face to control values. It achieves a mean absolute error of zero point zero one one four, which is more than ten times better than previous approaches.
Lu: And they validated it on a physical robot with twenty diverse human performers. The robot successfully mimicked subtle expressions like frowning, gaze direction, and asymmetric eyelid movements.
Meng: And from a practical standpoint, the framework is computationally accessible. It runs on a single RTX four thousand ninety. That means other labs can build on this work without needing massive compute resources.
Lalam: And the cultural implications are significant. This dataset could enable robots that adapt their emotional expressions to different cultural contexts, improving human-robot interaction in healthcare, education, and social settings.
Tom: Well said, everyone. We've covered the dataset, the framework, the results, and the broader implications. It's been a great discussion about a paper that really pushes the field forward.
Jane: And with that, we'll say goodbye to "X2C: A Large-Scale Benchmark for Nuanced Humanoid Facial Expression Imitation". Thanks for joining us, and we'll see you next time for another exciting paper.
Tom: Take care, everyone.
Peizhen Li, Longbing Cao, Xiao-Ming Wu, Runze Yang, Xiaohan Yu
Macquarie University · Nanyang Technological University
cs.RO, cs.AI, cs.HC
Submitted: 2026-08-11
Updated: 2026-08-12
Comments: Accepted by Pattern Recognition. Title updated and manuscript revised following peer review
Project page: https://lipzh5.github.io/X2CNet
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 59/100
The gist: The paper addresses the challenge of fine-grained facial expression transfer from humans to humanoid agents, which "presents a unique pattern recognition challenge due to the significant domain gap
Terminology
Summary
The paper addresses the challenge of fine-grained facial expression transfer from humans to humanoid agents, which presents a unique pattern recognition challenge due to the significant domain gap between biological facial dynamics and mechanical control spaces.
While visual synthesis of talking heads has advanced rapidly, mapping high-dimensional visual cues to precise, physically constrained actuation signals remains an open problem, primarily due to the lack of large-scale paired data.
The authors note that humanoid robots are increasingly deployed to interact socially, provide assistance, or support learning, making their ability to express emotions a key factor in fostering trust, empathy, and engagement.
However, "the fidelity of humanoid facial expressions—especially the fine-grained, nuanced emotional cues—remains difficult to guarantee due to the scarcity of data required to learn emotional subtleties and guide informed on-robot execution."
Existing datasets for humanoid facial expression imitation (Smile, Coexpression) are characterized as typically small in size, lack sufficient data diversity (e.g., they do not include asymmetric facial expressions), and the emotional nuances that can be learned are limited by low annotation dimensionality.
Furthermore, their annotation accuracy is not guaranteed as the dataset collection relies on facial landmark predictions, which introduce prediction errors.
The paper introduces X2C, a new resource for realistic humanoid facial expression imitation. It consists of 100,000 (image, control value) pairs, with each image depicting a humanoid robot displaying a diverse range of nuanced facial expressions.
Each image is annotated with 30 numerical control values representing the ground-truth expression configuration.
Key characteristics of the dataset include:
-
Size: 100,000 samples, compared to 15,000 for Smile and 1,000 for Coexpression
-
Asymmetric expressions: Included (marked with), unlike existing datasets
-
Input dimensionality: 512 × 512 × 3
-
Annotation dimensionality: 30 continuous control parameters
-
Annotation accuracy: Rated five stars, achieved through
interpolation equations based on parameters specified in the animation files
-
Data alignment: Perfect temporal alignment between images and control annotations, with
the timestep for sampling both control values and images is set to 0.05 seconds
-
Data diversity: Rated five stars
The authors state: To our knowledge, X2C is the first high-quality, high-diversity, large-scale dataset featuring nuanced humanoid facial expressions specifically designed for realistic humanoid imitation.
The dataset was collected using the Ameca humanoid robot, which features 32 Degrees of Freedom (DoFs) including facial actuators and head/neck movements.
There are 30 control values associated with these DoFs, which are responsible for driving actuators located at different expression-relevant control units, including the brows, lids, gaze, nose, mouth, head and neck.
The collection process involved:
-
Environment Preparation: Data collection was conducted
in a simulation environment where a virtual counterpart of the physical robot is available. Given the same control values, the virtual robot displays the same facial expressions as the physical one.
The authors argue thatthere will be no issues such as the sim-to-real gap since facial expressions should be disentangled from the robot's skin (or identity).
-
Expression Animation Curation:
10 volunteers were recruited from the student population at Macquarie University, all of whom were over 18 years old and provided informed consent.
Volunteers werepurposefully selected
withdifferent birth countries, including undergraduate and PhD students of different genders.
After 30 hours of training, they createdkey-framed animations
using interpolation methods includingCubic Bézier, Linear and Step interpolations.
The resulting560 animations (with durations ranging from 1 to 15 seconds)
were produced. -
Sampling and Annotation: Videos were filmed
using the same device (MacBook Air, 2020) and under consistent configurations.
Images weresampled at a constant timestep of s = 0.05 seconds and resized to a uniform resolution of 512×512 pixels.
A structural similarity check was appliedto identify and remove near-duplicate frames if their similarity exceeded a threshold (θ = 0.99).
Control values were obtained by retrievingthe keyframes and interpolation parameters from the animation metadata, formulate the corresponding interpolation equations
and samplingprecise control values at the same timestamps used to extract images.
The dataset includes basic facial expressions (e.g., surprise, joy, and sadness) at varying intensities, as well as complex expressions that may not fit neatly into basic emotion categories.
Asymmetric facial expressions are included to simulate the human-like behavior and encourage diversity.
The paper reports summary statistics for all 30 control values, noting there are noticeable deviations between µ and Vneu across most controls, and for many controls (e.g., JP, JY, LBC), the dataset samples span nearly the full achievable range.
The authors note they intentionally avoid sampling extreme values for certain controls
because: 1) They could cause irreversible damage to the robot (e.g., excessive head or neck movement such as HP, HR, HY, NP, NR may lead to mechanical wear), and 2) Such expressions are physiologically implausible for humans.
The paper proposes X2CNet, "a novel framework for realistic humanoid facial expression imitation. It decomposes the humanoid learning process into two stages: in the first stage, the expression dynamics is captured through a motion transfer module; in the second stage, the correspondence between nuanced humanoid facial expression and their underlying control values is learned via large-scale training using X2C."
The framework consists of two modules:
-
Motion Transfer Module: Adopts LivePortrait,
pretrained on a large corpus of high quality portrait data,
consisting ofan appearance extractor, a motion extractor, a warping module, and a generator.
-
Mapping Network:
Consists of a feature extractor (denoted by F) and a regression head. F is implemented using a ResNet18 backbone while the regression head is a multilayer perceptron with two hidden layers and ReLU activations.
The framework outputs 30 continuous control values that encode subtle movements of expression-relevant control units.
The dataset was split into training and test sets, using 80% of them for training.
The model was trained using AdamW as the optimizer with a weight decay of 0.05
and a cosine schedule with warmup as the learning rate scheduler, with an initial learning rate of 1e-3.
The model was trained using the Huber loss with a threshold value of δ = 0.01,
with batch size is set to 128
and trained for 100 epochs
on a single RTX 4090 GPU.
The proposed method was compared against three baselines:
-
RC: Randomly samples each control value from a uniform distribution
-
RT: Randomly selects samples from the training set
-
LMKC: Adopts the model architecture from prior work, predicting control values based on facial landmarks
Results (MAE): RC achieved 0.8951, RT achieved 1.0629, LMKC achieved 0.1602, and the proposed method achieved 0.0114, outperforming all three baselines on the test set consisting of 20,000 samples, achieving lower mean errors and smaller standard deviations.
Ablation studies were conducted on the feature extractor F, comparing EfficientNet-B0, VGG16, ResNet18, and ViT-B/16. Results showed VGG16 achieves the best performance
(MAE 0.0107), followed by ViT-B/16 (0.0111), ResNet18 (0.0114), and EfficientNet-B0 (0.0151). The authors note ResNet18 performs slightly worse than both, it is significantly more lightweight and computationally efficient.
For real-world validation, 20 human performers from 5 different countries, including both males and females
were recruited. They exhibit a variety of facial contours, skin tones, and hairstyles
and their facial expressions were captured under different lighting conditions, and some performers appear with accessories such as glasses and earphones.
Performers were instructed to go beyond canonical expressions by incorporating various subtleties such as frowning, gaze direction, and neck movement.
The results show the humanoid robot successfully mimics most of these nuanced expressions despite its hardware constraints.
The paper discusses facial expressions for affective human-robot interaction, noting that "some of them focus on human facial expression analysis, neglecting the emotionally intelligent behavior on robot's face, where the robots may only display limited categories of emotional signals (such as LED indicators) on their faces. The authors note that
the robots often fail to convey the nuances of emotions, leading to reduced user engagement and trust in HRI."
Regarding humanoid facial expression imitation, the paper states that only a limited set of facial expressions are covered, which limits the expressiveness of the robot
and there is a lack of public available resources for accessing advanced, delicate humanoid face, benchmarking different models on this task.
The authors acknowledge: While our dataset collection includes volunteers from multiple countries and genders, cultural biases may still be present, potentially influencing the interpretation or design of facial expressions.
They plan to expand the sample population for recruiting animation creators and to further diversify the dataset.
They also note that our current dataset includes facial expressions from only a single humanoid robot
but the data collection pipeline and the proposed imitation framework are designed to generalize to other humanoid platforms with different degrees of freedom (DoFs).
Future work will focus on extending X2C with fine-grained emotion labels to enable more precise supervision for the imitation task.
The paper notes: All human participants involved in the real-world experiments provided informed consent.
The authors warn that the dataset could be misused for deceptive, manipulative, or surveillance-related purposes, such as impersonation or unauthorized identity mimicry
and strongly discourage such applications.
Positively, the dataset has the potential to empower emotionally intelligent robots for socially beneficial applications, including elderly care, autism therapy, and education.
However, overly human-like robots may cause users—especially vulnerable individuals—to form emotional attachments or unrealistic expectations, possibly leading to confusion or psychological discomfort.
The paper's contributions are summarized as: "1) we introduce the X2C (Anything to Control)—a high-quality, high-diversity, large-scale datasets featuring nuanced humanoid facial expressions with precise control value annotations; 2) we propose X2CNet, a novel framework for human-to-humanoid expression imitation; 3) we provide real-world demonstrations on the physical robot to validate the effectiveness of our method, and the potential of our dataset in advancing realistic humanoid facial expression imitation."
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems and what the improved systems can do:
Improvement: Build a two-stage AI system (X2CNet architecture) that uses the 100,000-pair dataset (image → 30 continuous control values) to map human facial expressions to precise humanoid actuator commands. The first stage uses motion transfer (LivePortrait) to capture subtle expression dynamics; the second stage uses a ResNet18 backbone with a regression head trained on X2C to output 30 control values (e.g., brow inner left, lip corner raise right, gaze target theta) with a mean absolute error of 0.0114.
What it can do:
-
Imitate nuanced, asymmetric human expressions (e.g., one eyebrow raised, asymmetric lip curl) that basic emotion classifiers miss.
-
Drive a 32-DoF humanoid robot (like Ameca) in real time to reproduce frowns, gaze shifts, nose wrinkles, and head tilts with sub-0.02 control-value accuracy.
-
Generalize to in-the-wild performers with different skin tones, glasses, earphones, and lighting conditions, as validated on 20 performers from 5 countries.
Improvement: Train a generative model on X2C that learns the continuous control-value space, not just discrete emotion labels. The dataset includes expressions that are the same emotion (e.g., fear) at different intensities, plus blends that don’t fit basic categories. Use the 30-dimensional control vector as a latent representation for fine-grained emotion synthesis.
Improvement: Replace landmark-based annotation (which introduces prediction errors, as in Smile/Coexpression datasets) with an analytical interpolation-based annotation pipeline. The system uses keyframe metadata and Bézier/linear/step interpolation equations (Equations 1–3) to compute exact control values at any timestamp, perfectly aligned with sampled images (0.05s timestep).
Improvement: Use the X2C dataset to train a shared embedding space where human facial expressions (from any source) and humanoid control values are aligned. The paper demonstrates that expressions are identity-independent (skin type/robot skin doesn’t matter), so the embedding can be disentangled from identity.
Improvement: Deploy the trained mapping network (ResNet18 + regression head, Huber loss, cosine LR schedule) on edge hardware (e.g., single RTX 4090) for real-time inference. The model is lightweight (ResNet18 chosen over VGG16/ViT for efficiency with only 0.0007 MAE penalty) and outputs 30 values per frame.
Improvement: Use X2C as a standardized benchmark (100,000 samples, 30-dimensional annotations, asymmetric expressions included) to compare any imitation model. The paper provides baseline results (RC, RT, LMKC) and statistical metrics (MAE, SD, SEM, 95% CI) for fair comparison.
Improvement: Encode the dataset’s deliberate exclusion of extreme control values (e.g., head pitch beyond ±0.5, gaze beyond ±2.3) into the AI system as hard constraints. The paper notes that extreme values cause irreversible silicon skin damage and mechanical wear.
Bottom line: The improved AI system can take any human facial expression video, extract nuanced motion features, and output safe, precise, physically grounded control values that drive a humanoid robot to realistically mirror the expression—with quantitative accuracy (MAE 0.0114), real-time performance, and generalization to diverse human performers.
Abstract
Fine-grained facial expression transfer from humans to humanoid agents presents a unique pattern recognition challenge due to the significant domain gap between biological facial dynamics and mechanical control spaces. While visual synthesis of talking heads has advanced rapidly, mapping high-dimensional visual cues to precise, physically constrained actuation signals remains an open problem, primarily due to the lack of large-scale paired data. To bridge this gap, we introduce X2C, a comprehensive benchmark dataset comprising 100,000 (image, control value) pairs. Unlike existing resources, X2C features nuanced, physically grounded expressions annotated with 30 continuous control parameters, establishing a high-fidelity standard for this task. Building on this resource, we propose X2CNet, a two-stage deep learning framework that explicitly decouples visual motion features from mechanical control regression to model the correspondence between human perceptual cues and humanoid actuation. Extensive experiments, including quantitative benchmarking and real-world physical validation, demonstrate that our approach achieves superior cross-domain consistency and enables robust, in-the-wild expression imitation. Code and Data: https://lipzh5.github.io/X2CNet/
Sources
- UGotMe: An Embodied System for Affective Human-Robot Interaction
- Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks
- FABG : End-to-end Imitation Learning for Embodied Affective Human-Robot Interaction
- MediaPipe: A Framework for Building Perception Pipelines
- HumanPlus: Humanoid Shadowing and Imitation from Humans
- LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control
- VoxCeleb: a large-scale speaker identification dataset
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Humanoid Robots and Humanoid AI: Review, Perspectives and Directions
- Xpress: A System For Dynamic, Context-Aware Robot Facial Expressions using Language Models
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving