FitAQA: A Benchmark of Fitness Action Quality Assessment for Multimodal Large Language Models
cs.AI
Submitted: 2026-08-09
Updated: 2026-09-17
License: http://creativecommons.org/licenses/by/4.0/
The gist: Fitness Action Quality Assessment (AQA) is important for intelligent sports training, yet the capabilities of Multimodal Large Language Models (MLLMs) in this setting remain underexplored.
Terminology
Abstract
Fitness Action Quality Assessment (AQA) is important for intelligent sports training, yet the capabilities of Multimodal Large Language Models (MLLMs) in this setting remain underexplored. Existing benchmarks rely on action-specific annotation schemes and focus primarily on final assessment outputs, offering limited insight into how models assess exercise quality. We introduce FitAQA, a systematic benchmark for evaluating MLLMs in fitness AQA, containing 2,219 videos and 5,512 QA instances across 30 bodyweight exercises. In collaboration with experts in sports science, we develop a unified form error taxonomy that defines 38 recurring form errors within six complementary quality dimensions: alignment, symmetry, stability, coordination, tempo, and completeness. This taxonomy provides a shared assessment framework across different exercises. FitAQA further formulates three evaluation tasks: perception for recognizing relevant visual evidence, judgement for combining that evidence with domain knowledge to assess execution correctness, and temporal grounding for localizing form errors over time. Extensive evaluation shows that current MLLMs still struggle to assess exercise quality comprehensively and localize form errors precisely. Controlled experiments further indicate that visual perception is a key bottleneck, as judgement performance improves substantially when ground-truth perceptual evidence is provided. The dataset and evaluation code will be made publicly available.
Sources
- Qwen3-VL Technical Report
- HumanVideo-MME: Benchmarking MLLMs for Human-Centric Video Understanding
- A Short Note on the Kinetics-700 Human Action Dataset
- Can Vision Language Models Judge Action Quality? An Empirical Evaluation
- GPT-4o System Card
- Real-Time Feedback and Benchmark Dataset for Isometric Pose Evaluation
- From 3D Pose to Prose: Biomechanics-Grounded Vision--Language Coaching
- Explainable Action Form Assessment by Exploiting Multimodal Chain-of-Thoughts Reasoning
- OpenAI GPT-5 System Card
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- FLEX: A Largescale Multimodal, Multiview Dataset for Learning Structured Representations for Fitness Action Quality Assessment
- VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
- TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
- FineSkiing: A Fine-grained Benchmark for Skiing Action Quality Assessment
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection