MyoMechanix: Biomechanically-Grounded Compositional Skilled Activity Understanding and Coaching
cs.CV, cs.AI, cs.ET, cs.HC, cs.LG
Submitted: 2026-08-26
Updated: 2026-08-26
Terminology
Sources
- FitAQA: A Benchmark of Fitness Action Quality Assessment for Multimodal Large Language Models
- Learning to Visually Connect Actions and their Effects
- Uncertainty-Driven Action Quality Assessment
- SkillSight: Efficient First-Person Skill Assessment with Gaze
- Video Action Differencing
- RICA2: Rubric-Informed, Calibrated Assessment of Actions
- TVQA+: Spatio-Temporal Grounding for Video Question Answering
- Explainable Action Form Assessment by Exploiting Multimodal Chain-of-Thoughts Reasoning
- MiniCPM-V: A GPT-4V Level MLLM on Your Phone
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- Qwen2.5-VL Technical Report
- LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models