REMEDY: How Far Is Video Generation from Medical Education World Models?
cs.CV
Submitted: 2026-09-26
Updated: 2026-09-26
Code: https://github.com/ultralytics/ultralytics
Terminology
Sources
- Safety is Contextual, LLM-Judges Are Not: Navigating the Rigid Priors of Evaluators
- VideoPhy: Evaluating Physical Commonsense for Video Generation
- EduVQA: Towards Concept-Aware Assessment of Educational AI-Generated Videos
- MoHallBench: A Benchmark for Motion Hallucination in Video Large Language Models
- How Far Are Surgeons from Surgical World Models? A Pilot Study on Zero-shot Surgical Video Generation with Expert Assessment
- Same Verdict, Different Reasons: LLM-as-a-Judge and Clinician Disagreement on Medical Chatbot Completeness
- HumanScore: Benchmarking Human Motions in Generated Videos
- LTX-2: Efficient Joint Audio-Visual Foundation Model
- VLM Judges Can Rank but Cannot Score: Task-Dependent Uncertainty in Multimodal Evaluation
- MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation
- Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation
- Do generative video models understand physical principles?
- Cosmos World Foundation Model Platform for Physical AI
- Right Answer, Wrong Method: Shortcut Hacking Misleads the Evaluation of LLM Reasoning on Frontier Science Benchmarks
- T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation
- Bora: Biomedical Generalist Video Generation Model
- Towards Accurate Generative Models of Video: A New Metric & Challenges
- LLM Judges Have Dark Current: A Psychometric Datasheet for LLM-as-a-Judge Evaluation
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models