MedHorizon: Towards Long-context Medical Video Understanding in the Wild
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "MedHorizon: Towards Long-context Medical Video Understanding in the Wild".
Jane: Long medical video understanding requires models to navigate highly redundant, sparse, and context-dependent evidence distributed across full clinical procedures.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Okay, moving on to the title and authors of "MedHorizon: Towards Long-context Medical Video Understanding in the Wild." The team includes Bodong Du, Bowen Liu, Yang Yu, Xinpeng Ding, Zhiheng Wu, Shuning Wang, Shuo Nie, Naiming Liu, Qifeng Chen1 through Xiaomeng Li1. It’s a big group of experts coming together on this specific challenge.
Jane: And the title itself really tells you what it's about: pushing multimodal large language models to handle long-context medical video understanding in real-world situations. It’s not just about looking at pictures; it’s about understanding an entire procedure from start to finish, which is a completely different level of complexity than standard video analysis.
Lu: The authors clearly understand the gap they are trying to fill by emphasizing that existing benchmarks often assume the evidence is already localized, but MedHorizon forces models to do the retrieval before they can reason. That’s a critical distinction in their problem framing for full-procedure medical video understanding in the wild (MedHorizon: Towards Long-context Medical Video Understanding in the Wild Bodong Du1†, Bowen Liu1†, Yang Yu1, Xinpeng Ding3, Zhiheng Wu2, Shuning Wang2, Shuo Nie2, Naiming Liu2, Qifeng Chen1, Yangqiu Song1, Xiaomeng Li1 1The Hong Kong University of Science and Technology, 2Baidu Inc <ref:2605.06537#pg0,MedHorizon: Towards Long-context Medical Video Understanding in the Wild Bodong Du1>., threeXidian University †Equal Contribution Medical multimodal large language models (MLLMs) have advanced image understanding and shortvideo analysis, but real clinical review often requires full-procedure video understanding <ref:2605.06537#pg0,3Xidian University †Equal Contribution Medical multimodal large language models (MLLMs) have advanced>. Unlike general long videos, medical procedures contain highly redundant anatomical views, while decisive evidence is temporally sparse, spatially subtle, and context dependent. Existing benchmarks often assume this evidence has already been localized through images, short clips, or pre-segmented videos).
Meng: I’m interested in that part about "in the wild." That implies they aren't just using perfectly curated datasets; they’re testing how models hold up when the data is messy and representative of actual clinical practice. That makes the results much more meaningful for practical applications.
Lalam: It really highlights that our vision systems need to stop being good at isolated object detection and start being good at maintaining a coherent, long-term view of a complex process, which is something this paper points directly toward with its focus on full-procedure video understanding.
The paper's summary: Tom: So what’s the actual core of what they did? MedHorizon is an in-the-wild benchmark composed of three hundred forty public full-length videos spanning seven organs and two clinical scenarios: diagnostic examination and surgery, totaling seven hundred fifty-nine hours of clinical video. It provides one thousand two hundred fifty-three evidence-grounded multiple-choice questions that test both sparse evidence understanding and multi-hop clinical reasoning <ref:2605.06537#pg0,provides 1,253 evidence-grounded multiple-choice questions that>.
Jane: Essentially, the paper summarizes that this benchmark forces models to search through the entire raw video stream before they can interpret what they see. They aren't just given a short clip; they have to find things across the whole procedure, which is a much harder task than most current evaluations.
Lu: The structure of their evaluation pipeline is key here; it involves retrieving evidence from full videos, building candidate questions around annotated events or findings, filtering those candidates through automatic checks, strengthening distractors with language rewriting from GPT-five point four, and then finally having three non-physician data reviewers audit everything to get those one thousand two hundred fifty-three final QA items.
Meng: From an engineering standpoint, that pipeline sounds incredibly complex to build and maintain. Getting the annotations right for seven hundred fifty-nine hours of video and ensuring the human verification protocol is consistent across all those items must have been a massive undertaking before you even start testing the models <ref:2605.06537#pg0>.
Lalam: That complexity speaks to how deep these problems are; it’s not just about having a big model, it’s about building a robust infrastructure that can handle this level of structured, multi-step evidence evaluation for medical video understanding in the wild.
The paper's improvements: Tom: Now we get to the findings they pulled from testing these models. They found four major things: first, performance doesn't scale reliably with more frames because adding frames can just introduce "near-duplicate anatomy" and degrade accuracy once the redundant context overwhelms the useful signal.
Jane: That’s a tough pill to swallow for model builders, Tom. It suggests that simply feeding them more video isn't a magic fix; in fact, it can hurt their ability to focus on what actually matters because of all that visual noise.
Lu: The second finding points out that evidence retrieval and clinical interpretation are the main bottlenecks, meaning we need both better temporal retrieval based on evidence and better visual verification grounded in clinical knowledge.
Meng: That tells me the practical application isn't just about making the vision model bigger; it’s about building a system that can intelligently decide *when* to look at what part of the video stream, which is a much harder engineering challenge to solve than just increasing parameters.
Lalam: And they also found a weakness in weak procedural reasoning and attention drift under redundancy, where the model’s focus drifts toward visually repetitive but clinically non-decisive frames instead of the sparse answer-relevant evidence.
Conclusion: Tom: So, to wrap up on "MedHorizon: Towards Long-context Medical Video Understanding in the Wild," the main point is that for medical video understanding, simply scaling context length isn't enough; models need stronger mechanisms specifically for sparse evidence perception and contextual reasoning.
Jane: Exactly. The biggest obstacle they found is that there's a coupling between finding sparse evidence and performing cross-temporal reasoning across different parts of the procedure. That’s what we need to work on next, not just longer inputs for existing models.
Lu: The authors suggest that future medical MLLMs need stronger mechanisms for sparse evidence perception and contextual reasoning rather than just scaling context length (MedHorizon: Towards Long-context Medical Video Understanding in the Wild Bodong Du1†, Bowen Liu1†, Yang Yu1, Xinpeng Ding3, Zhiheng Wu2, Shuning Wang2, Shuo Nie2, Naiming Liu2, Qifeng Chen1, Yangqiu Song1, Xiaomeng Li1 1The Hong Kong University of Science and Technology) <ref:2605.06537#pg0,MedHorizon: Towards Long-context Medical Video Understanding in the Wild Bodong Du1>.
Meng: From an engineering standpoint, this suggests we need to build more sophisticated attention mechanisms that can dynamically filter out the visual redundancy they identified so models can concentrate on those subtle decisive frames.
Lalam: I think what this paper shows is that medical AI needs to evolve from just pattern recognition to true clinical reasoning by incorporating structured procedural knowledge into how the AI processes time and space.
Tom: It’s a heavy paper, but it gives us a very clear roadmap for where the research needs to go next in making these models genuinely useful in complex clinical environments. We’ll be looking at what comes next on arXiv soon.
The Hong Kong University of Science and Technology
cs.CV
Submitted: 2026-05-07
Updated: 2026-10-06
Comments: NeurIPS 2026
Code: https://github.com/QwenLM/Qwen3-VL
Project page: https://alibaba-damo-academy.github.io/lingshu
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 83/100
The gist: Long medical video understanding requires models to navigate highly redundant, sparse, and context-dependent evidence distributed across full clinical procedures.
Key concepts
- In-the-wild Benchmark
- This benchmark uses 340 public, full-length clinical videos spanning various organs and procedures. It tests models on real, messy data rather than pre-cut clips, forcing them to search through the entire video stream to find relevant information.
- Sparse Evidence Understanding
- This refers to the model's ability to correctly interpret medical findings when only small, scattered pieces of evidence are available. The benchmark specifically challenges models that must reason using minimal, non-contiguous visual cues from a long procedure.
- Multi-hop Clinical Reasoning
- This is the complex task where a model must connect multiple pieces of sparse evidence across different points in time or different parts of the video to form one complete clinical judgment. It requires aggregating scattered observations into a single, coherent understanding.
- Non-monotonic Frame Scaling
- This finding means that simply increasing the number of frames in a video does not always improve accuracy. In fact, adding too many frames can introduce redundant or near-duplicate visual information that confuses the model instead of helping it understand the procedure.
Terminology
Summary
Long medical video understanding requires models to navigate highly redundant, sparse, and context-dependent evidence distributed across full clinical procedures. MedHorizon introduces an in-the-wild benchmark designed to rigorously test whether multimodal large language models (MLLMs) can perform full-procedure retrieval and multi-hop clinical reasoning using extremely sparse evidence. The best model evaluated reached only 41.1% accuracy, indicating that current systems remain far from robust full-procedure understanding.
How it works
MedHorizon is an in-the-wild benchmark comprising 340 public full-length videos spanning 7 organs and 2 clinical scenarios: diagnostic examination and surgery, totaling 759 hours of clinical video. It provides 1,253 evidence-grounded multiple-choice questions that jointly evaluate sparse evidence understanding and multi-hop clinical reasoning. The benchmark preserves the full procedural stream rather than pre-cut clips, requiring models to search noisy procedural streams before interpreting and aggregating findings.
Benchmark Construction
The benchmark construction follows an evidence grounded pipeline
(Fig. 2), which involves several rigorous steps:
-
Source Annotations + Task Templates are collected covering phases, anatomical regions, examination targets, lesion properties, instruments, and procedure level events.
-
Each QA item is generated from explicit video evidence in expert annotations by retrieving evidence from the full video (including sparse key frames and temporally relevant segments), building candidate questions around the annotated event or finding.
-
Candidates are filtered through
Automatic filtering
to remove temporally invalid, ambiguous, duplicate, or malformed items (removing 81 invalid candidates). -
Distractors are strengthened using
Language rewriting with GPT-5.4
and by drawing distractors fromclinically plausible alternatives.
-
The final stage is a
Human verification
protocol where three non-physician data reviewers audit every retained item against the source annotations to ensure evidence support and label consistency, resulting in 1,253 final QA items.
Diagnostic Findings
The evaluation reveals four key findings regarding model performance:
-
Performance does not scale reliably with more frames, showing
non-monotonic frame scaling,
where adding frames can introducenear-duplicate anatomy
and degrade accuracy once redundant context dominates the useful signal. -
Evidence retrieval and clinical interpretation remain primary bottlenecks, exposing a dual bottleneck requiring both
evidence-aware temporal retrieval and clinically grounded visual verification.
-
There is a weakness in
weak procedural reasoning and attention drift under redundancy,
where model attention drifts towardvisually repetitive but clinically non-decisive frames and patches
rather than sparse answer-relevant evidence. -
Generic sampling methods show limited ability, as they often fail to address the challenge of finding
subtle decisive frames
and maintainingglobal procedural continuity or multi-step aggregation.
Task Hierarchy and Bottlenecks
MedHorizon categorizes tasks into fine-grained understanding (PR, SD, RL, BPQ, LLS, Site, Hist, Size) and semantic reasoning (IR, WR, RSR, CPR). The hierarchy requires models to move from one-hop visual-semantic tasks
like Instrument Recognition (IR) to three-hop aggregative tasks,
such as Count-Proportion Reasoning (CPR) and Lesion Analysis (LA), which demand locating repeated events or lesion-relevant moments and aggregating them into a unified, procedure-level clinical judgment. The results show that semantic reasoning tasks, particularly CPR and LA, remain difficult because they require counting, comparing, and aggregating distributed observations across the same procedure.
Sampling Strategies
The paper analyzes the impact of different sampling strategies on performance. Uniform sampling remains stable by preserving procedural coverage but is evidence-agnostic. Conversely, change-focused samplers improve short-window tasks like RL and Size but hinder workflow-sensitive tasks (e.g., PR). Specialized methods, such as AKS, ViLAMP, and WFS-SB (using 128 frames), show modest gains; for instance, WFS-SB reaches 32.8% on Site,
reflecting its design to employ semantic-boundary selection to retain structured temporal context.
However, these improvements remain task-dependent and still underperform on tasks requiring global procedural continuity or multi-step aggregation like WR and RSR.
Conclusion
MedHorizon provides a rigorous testbed for MLLMs by explicitly evaluating the full-procedure setting, demonstrating that the primary obstacle is the coupling between sparse evidence localization and cross-temporal reasoning.
Future medical MLLMs need stronger mechanisms for sparse evidence perception and contextual reasoning
rather than simply scaling context length. The benchmark remains limited by public-video diversity and annotation quality but offers a clear path for research evaluation in this challenging domain.
Improvements for AI systems
As a fastidious research assistant, I have analyzed the MedHorizon paper and identified critical weaknesses in current Medical Multimodal Large Language Models (MLLMs) when applied to long-context clinical video understanding.
The core limitation is that current models fail because they treat long medical procedures like general long videos—assuming salient evidence is dense and easy to retrieve—when, in reality, the decisive evidence is extremely sparse, spatially subtle, and temporally distributed across redundant frames.
Here are the specific improvements I propose for AI systems based on the MedHorizon findings:
) 1. Implement Evidence-Aware Temporal Retrieval Mechanisms (Addressing RQ2).
The system must move beyond simple frame sampling or uniform coverage.
-
Specific Improvement: Develop a retrieval module that uses clinical context (e.g.,
during resection phase,
post-withdrawal
) to guide sparse evidence search, rather than relying on visual saliency alone. This requires integrating procedural knowledge graphs or temporal segmentation models with the vision encoder. -
Improved AI Capability: The system will be able to reliably locate rare, subtle clinical events (e.g., a specific lesion appearance at a precise 15-minute mark) that are visually indistinguishable from redundant frames, significantly improving accuracy in tasks like Lesion Analysis (LA) and Step Discrimination (SD).
) 2. Develop Sparse Evidence Perception and Compression Modules (Addressing RQ4).
The model needs to learn to ignore redundancy while preserving clinically relevant information.
-
Specific Improvement: Integrate a mechanism for
evidence compression
orcontextual denoising
during the attention aggregation phase. This module should be trained to suppress attention on patches exhibiting high spatial redundancy with keyframes, focusing computational resources only on regions where the visual features deviate significantly from the established procedural norm (i.e., identifying outliers). -
Improved AI Capability: The system will reduce
attention drift,
allowing it to maintain focus on sparse, decisive moments even when surrounded by highly similar anatomical views. This directly addresses the failure of generic sampling methods, leading to more stable performance across different video inputs.
) 3. Establish Procedural Reasoning and Contextual Binding Layers (Addressing RQ3).
The system must learn the why
and where
behind observations, not just the what.
-
Specific Improvement: Introduce a dedicated reasoning layer that explicitly models clinical workflows (e.g., gastric bypass steps, colonoscopy phases) as structured temporal dependencies. This layer should be tasked with binding localized visual cues (e.g., an instrument) to its correct procedural phase or anatomical region before performing cross-temporal comparisons (RSR).
-
Improved AI Capability: The system will transition from pattern recognition to true clinical reasoning. It will be able to correctly answer multi-hop questions that require linking a finding in the initial phase to a subsequent event in a later phase, mimicking the holistic judgment of an experienced clinician.
) 4. Implement Non-Monotonic Frame Scaling Strategies (Addressing RQ1).
The system must stop assuming that more input always equals better performance.
-
Specific Improvement: Design an adaptive sampling strategy that dynamically adjusts frame budget based on the identified evidence density and the task complexity, rather than using a fixed cap or uniform sampling. The system should prioritize retaining frames with high information gain for specific reasoning tasks (e.g., prioritizing frames around known procedural transitions).
-
Improved AI Capability: The system will achieve robust performance across a wide range of video lengths without suffering catastrophic degradation in accuracy when context length increases, overcoming the
context-dilution
effect observed in models like MedGRPO.
In summary, the improved AI system will transition from a general video analyzer to a specialized clinical reasoning engine capable of performing surgical or endoscopic review by:
-
Locating needle-in-a-haystack sparse evidence accurately.
-
Filtering out visual redundancy to maintain focus on subtle findings.
-
Applying structured procedural knowledge to connect temporally distant observations into coherent clinical judgments.
Abstract
Medical multimodal large language models (MLLMs) have advanced image understanding and short-video analysis, but real clinical review often requires full-procedure video understanding. Unlike general long videos, medical procedures contain highly redundant anatomical views, while decisive evidence is temporally sparse, spatially subtle, and context dependent. Existing benchmarks often assume this evidence has already been localized through images, short clips, or pre-segmented videos, leaving the retrieval-before-reasoning problem under-tested. We introduce MedHorizon, an in-the-wild benchmark for long-context medical video understanding. MedHorizon preserves 759 hours of full-length clinical procedures and provides 1,253 evidence-grounded multiple-choice questionsthat jointly evaluate sparse evidence understanding and multi-hop clinical reasoning. Its evidence is extremely sparse, with only 0.166% evidence frames on average, requiring models to search noisy procedural streams before interpreting and aggregating findings. We evaluate representative general-domain, medical-domain, and long-video MLLMs. The best model reaches only 41.1% accuracy, showing that current systems remain far from robust full-procedure understanding. Further analysis yields four key findings: performance does not scale reliably with more frames, evidence retrieval and clinical interpretation remain primary bottlenecks; these bottlenecks are rooted in weak procedural reasoning and attention drift under redundancy, and generic sampling methods only partially balances local detail with global coverage. MedHorizon provides a rigorous testbed for MLLMs that retrieve sparse evidence and reason over complete clinical workflows.
Sources
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale
- Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understanding
- Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation
- Hulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding
- VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
- Divide-then-Diagnose: Weaving Clinician-Inspired Contexts for Ultra-Long Capsule Endoscopy Videos
- SUREON: A Benchmark and Vision-Language-Model for Surgical Reasoning
- How Well Can General Vision-Language Models Learn Medicine By Watching Public Educational Videos?
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning
- MiniCPM-V: A GPT-4V Level MLLM on Your Phone
- VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
- Long Context Transfer from Language to Vision
- LLaVA-Video: Video Instruction Tuning With Synthetic Data
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models