One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "One Ranking, Any Budget".
Jane: Frame selection is crucial for applying Large Multimodal Models (LMMs) to long videos due to severe frame redundancy and limited context windows,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Let's talk about the title and who came up with this work. It’s "One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding." The authors are Wang Chen, Yu Chen, Xiang Wang, Shuai Li, Jinfa Huang, and Xiawu Zheng.
Jane: Those names suggest a team with diverse expertise because they're tackling this complex problem from different angles. It seems like the core idea revolves around using a Matryoshka ranking to manage frame selection for long videos.
Lu: The authors are clearly digging into the structure of how we need to prioritize information across different temporal scales, which is really interesting when you consider the complexity of video data.
Meng: I'm looking at their background, and it shows they've been working on related things like sparse probing and visual token compression before coming up with this specific ranking mechanism.
Lalam: It’s impressive that they managed to combine those ideas into one training-free framework; that suggests a very robust approach that doesn't rely on specific model knowledge.
The paper's summary: Tom: So, what’s the actual core idea behind this paper, "One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding"? Basically, they are proposing a training-free framework that builds a single priority sequence which is position-adaptive.
Jane: That means instead of optimizing separate frame subsets for every budget you want to use, MEC creates one ranking structure where the small prefixes focus on specific query evidence while the larger prefixes build up broader temporal context.
Lu: The mathematical objective they define, "R⋆ ∈ arg max R∈ΠM ∑ K∈B πK QK(R:K; V, q)," is what really drives this; it rewards concentrating query-conditioned evidence for tight budgets and then progressively rewards complementary temporal and visual context as the prefix grows.
Meng: It sounds like they’re trying to solve the problem where if you switch from a short context window to a long one, the frames you picked before might not be optimal anymore because they were optimized for that specific constraint.
Lalam: That ability to serve multiple frame budgets with just one ranking is what really stands out; it simplifies the deployment side immensely.
The paper's improvements: Tom: One of the main improvements they highlight is moving away from optimizing isolated subsets for each budget, which was a problem with existing methods. MEC constructs one Matryoshka ranking whose prefixes preserve early evidence and progressively extend temporal context, unlike previous methods that might replace previously selected evidence when the budget changes.
Jane: They also built a reusable sparse video index first, which involves creating low-resolution appearance encodings and local visual change estimates before scoring the frames to discover query-conditioned evidence using a multi-cue scoring mechanism.
Lu: The discovery mechanism uses an evidence score defined by "En(q) = λrρn(q) + λ∆∆n + λoOn," which balances relevance, visual change, and observability cues, followed by segment activation and anchor-centered local zooming to form a compact candidate pool.
Meng: That multi-cue scoring sounds like a practical way to filter out frames that are just noisy or irrelevant before the model even gets to look at them in detail.
Lalam: And then they refine this with a position-adaptive surrogate score, "uk(n) = we(k)En(q) + wc(k)Hn(k) + wdDn(k)," which dynamically balances the evidence, temporal coverage, and visual diversity based on the current rank position.
Conclusion: Tom: So to wrap up, the paper "One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding" shows how a single ranking can handle multiple frame budgets by having prefixes that concentrate evidence and then progressively expand context. It’s a training-free method that works well across four long-video benchmarks and six frame budgets.
Jane: Essentially, this means LMMs can operate more efficiently under varying input constraints because they get the best possible selection tailored to their needs without needing a new selector for every scenario.
Lu: The implications for video understanding are that we can now expect better performance across different reasoning demands and latency requirements with a single, consistent configuration.
Meng: Practically, this means we can reduce the computational overhead of processing long videos by using this selection strategy instead of exhaustively sampling everything.
Lalam: I think what really matters is how this system decouples the frame selection from the specific LMM architecture; it allows us to transfer one robust configuration across different models.
Tom: It’s a solid piece of work focusing on making long-video reasoning practical and efficient, and we’ll be keeping an eye on how this influences our next steps.
Amap, Alibaba Group Company
cs.CV
Submitted: 2026-08-06
Updated: 2026-09-30
Comments: 19 pages
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 90/100
The gist: Frame selection is crucial for applying Large Multimodal Models (LMMs) to long videos due to severe frame redundancy and limited context windows, and this work introduces Matryoshka
Key concepts
- Matryoshka Ranking Problem
- This is the core idea where a single ranking is built such that its small prefixes highlight specific query-related evidence. As the prefix grows, it simultaneously retains that initial evidence while incorporating more general temporal and visual context from the video, making it versatile for various selection needs.
- Reusable Sparse Video Index
- This index is pre-built to efficiently represent the video without needing a query. It uses low-resolution encodings and estimates of local visual change between adjacent segments to quickly map out important temporal locations in the video.
- En(q) Score
- This score evaluates how useful a frame is for a specific query, combining three factors: relevance (how much it relates to the question), visual change (if it looks different from neighbors), and observability (how easily it can be seen). This allows the system to prioritize frames that are both informative and visually distinct.
- Position-Adaptive Matryoshka Ranking
- This final step builds the actual selection sequence by calculating a score for every frame based on its position in the ranking. This score balances query evidence, temporal coverage, and visual diversity, ensuring that any shorter part of the sequence is an accurate subset of any longer one.
Terminology
Summary
Frame selection is crucial for applying Large Multimodal Models (LMMs) to long videos due to severe frame redundancy and limited context windows, and this work introduces Matryoshka Evidence-to-Context (MEC) Frame Selection, a training-free framework that constructs a single priority sequence capable of serving multiple frame budgets. This method addresses the challenge of optimizing an isolated frame subset for each budget by formulating long-video frame selection as a Matryoshka ranking problem,
where small prefixes concentrate query-conditioned evidence, while progressively larger prefixes preserve this evidence and add broader temporal context, allowing a single ranking to be truncated to any target budget without rerunning the selector.
The Matryoshka Ranking Problem
The core idea is to construct a single priority sequence whose small prefixes concentrate query-conditioned evidence, while progressively larger prefixes preserve this evidence and add broader temporal context.
This contrasts with existing methods that optimize an isolated frame subset for each predefined budget, which may lead to previously selected evidence being replaced when the budget changes. The objective is defined as:
R⋆ ∈ arg max R∈ΠM ∑ K∈B πK QK(R:K; V, q)
where the utility function rewards concentrated query-conditioned evidence for tight budgets and progressively rewards complementary temporal and visual context as the prefix grows.
Reusable Sparse Video Index
To efficiently construct this ranking without substantial overhead, MEC first builds a reusable sparse video index.
This index is independent of the query and caches low-resolution appearance encodings, local visual change estimates, and observability priors. The construction involves:
-
Duration-Adaptive Sparse Probing to create temporally ordered segments.
-
Low-Resolution Appearance Encoding using L2-Norm on resized grayscale frames to reduce storage and computation overhead.
-
Local Visual-Change Estimation by comparing adjacent probes on the sparse grid to highlight visually distinctive moments, defined as
∆n = (1/2 max m∈AdjP (n) fn − fm squared.
Query-Conditioned Evidence Discovery
Once the index is built, MEC discovers query-conditioned evidence through a multi-cue scoring mechanism. The evidence score for a frame is defined as:
En(q) = λrρn(q) + λ∆∆n + λoOn
where Relevance (λr=0.7), Visual Change (λ∆=0.2), and Observability (λo=0.1) are weighted cues. Following this scoring, the framework refines promising regions using:
-
Evidence-Guided Segment Activation, which retains segments with the largest mean evidence scores based on an activation ratio η=0.25 to form the active set A(q).
-
Anchor-Centered Local Zooming, where the highest-evidence probe serves as an anchor, and
Uni(a−j, aj, Lz)
samples distinct interior frames around this anchor to form a compact candidate pool C(q).
Position-Adaptive Matryoshka Ranking
The final stage involves greedily constructing the ranking using a position-dependent surrogate score:
uk(n) = we(k)En(q) + wc(k)Hn(k) + wdDn(k)
This score balances query-conditioned evidence (we), temporal coverage (wc, where wc = 1 - we - wd), and visual diversity (wd). The ranking is constructed by appending the candidate with the highest surrogate utility at each rank k, ensuring that every shorter ranking is an exact prefix of every longer one.
Experimental Validation and Efficiency
Extensive experiments across four long-video benchmarks (Video-MME, LongVideoBench, MLVU) and six frame budgets demonstrate MEC's effectiveness. MEC improves average accuracy over uniform sampling by 3.77 percentage points across the benchmarks. Furthermore, it significantly reduces end-to-end selection latency by 47.37–51.19% compared to state-of-the-art selectors like AKS and WFS-SB, while remaining competitive with a favorable accuracy–latency trade-off across all tested downstream LMMs (Qwen3.5, LLaVA, InternVL). The framework is training-free and uses a shared configuration across all experiments.
Limitations
The paper notes limitations in sparse discovery (sparse-discovery blind spots
), where the grid may miss brief, visually isolated events that fall between probes. Additionally, reliance on lightweight proxies like BLIP-2 ITM scores and fixed evidence mixtures means the system may fail to capture evidence emerging from ordered pairs or long temporal dependencies. Finally, the greedy construction is a surrogate for global optimization; it does not guarantee global optimality for any individual budget or weighted aggregate of utilities.
Improvements for AI systems
Here are the specific improvements to AI systems based on the Matryoshka Evidence-to-Context (MEC) Frame Selection framework, and what these improved systems can achieve:
- 】Multi-Budget Adaptability for Constrained Context Windows
The MEC framework allows a single, reusable frame selection sequence to be truncated into any target budget (K) without rerunning the entire complex selector.
- 】Specific Improvement:
Improve Long-Video Understanding Models (LMMs) by enabling them to operate efficiently under varying input constraints (e.g., different context window sizes or processing budgets).
- 】Improved AI System Capability:
- Reduce end-to-end selection latency by up to 51% compared to fixed budget optimizers, allowing real-time processing of long videos with limited computational resources.
- 】Generalization Across Downstream Architectures and Models
The MEC selector is training-free and architecture/model agnostic, relying on a reusable sparse index and query-conditioned scoring (BLIP-2 ITM).
- 】Specific Improvement:
Decouple the frame selection mechanism from the specific Large Multimodal Model (LMM) architecture or internal tokenization strategy.
- 】Improved AI System Capability:
- Transfer a single, robust frame selection configuration across different LMMs (e.g., Qwen3.5-9B, LLaVA-OneVision-2), ensuring consistent performance regardless of the downstream model's specific input representation or compression method.
- 】Position-Adaptive Evidence Prioritization
The system constructs a Matryoshka ranking where early ranks emphasize query-conditioned evidence, and later ranks progressively favor temporal coverage and visual diversity.
- 】Specific Improvement:
Implement a dynamic, position-dependent weighting scheme that shifts the optimization focus from purely relevance (early) to context expansion (later) based on the required budget K.
- 】Improved AI System Capability:
- Achieve superior reasoning for complex queries by ensuring that critical, fleeting evidence is concentrated in the initial frames while maintaining a broad temporal understanding of the video's narrative structure in subsequent frames.
10.】Efficient Candidate Pool Construction via Sparse Probing and Local Zooming
The framework builds a reusable sparse index (caching low-resolution appearance, local visual change, and observability) and uses query-conditioned discovery through multi-cue scoring, segment activation, and anchor-centered zooming to form a compact candidate pool.
11.】Specific Improvement:
Develop an efficient pre-processing pipeline that generates a query-independent visual scaffold (sparse index) independent of the specific question asked. The system then refines the search space by focusing only on temporally relevant and visually distinctive regions around discovered anchors.
12.】Improved AI System Capability:
- Significantly reduce the computational overhead associated with exhaustively sampling long videos, enabling faster query processing while maintaining high fidelity for key evidence discovery.
13.】Robustness Against Sparse Discovery Blind Spots (Mitigation via Contextual Refinement)
The framework explicitly identifies limitations related to sparse probing missing critical, temporally isolated events (e.g., the final state). The greedy ranking mechanism is designed to leverage temporal coverage and diversity terms to mitigate these omissions where possible.
14.】Specific Improvement:
Inference time, if necessary, can be extended by utilizing the structure of the Matryoshka ranking (where prefixes are exact subsets) to ensure that even if a specific frame is missed in an initial sparse probe, it can be recovered by subsequent ranks that prioritize temporal or diversity coverage.
15.】Component Ablation and Hyperparameter Optimization
The MEC framework allows for detailed ablation studies on its core components: evidence scoring mixture, position-adaptive schedules (evidence/coverage/diversity weights), and ranking horizon (M).
16.】Specific Improvement:
System designers can precisely tune the trade-off between evidence concentration and temporal context expansion by adjusting the rank weight schedule.
17.】Improved AI System Capability:
- Tailor the frame selection strategy dynamically based on whether the task requires immediate, precise evidence (low budget) or broader narrative understanding (high budget), leading to optimal accuracy for any deployment scenario.
18.】Enhanced Visual Quality Priors (Observability Scoring)
The system incorporates a query-independent observability prior that combines frame sharpness and well-exposedness, suppressing severely degraded frames without relying on semantic content scores.
19.】Specific Improvement:
Integrate a dual-metric visual quality assessment into the evidence scoring process to filter out noise
(e.g., blurry or poorly exposed frames) before they even enter the query relevance calculation, ensuring the model focuses its attention on visually viable data points.
Sources
- DivPrune: Diversity-based Visual Token Pruning for Large Multimodal Models
- LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence
- Qwen3-VL Technical Report
- Matryoshka Multimodal Models
- Event-Anchored Frame Selection for Effective Long-Video Understanding
- Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understanding
- Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
- MatFormer: Nested Transformer for Elastic Inference
- Agentic Keyframe Search for Video Question Answering
- Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders
- Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding
- FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models
- Matryoshka Query Transformer for Large Vision-Language Models
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
- QuoTA: Query-oriented Token Assignment via CoT Query Decouple for Long Video Comprehension
- Adaptive Keyframe Sampling for Long Video Understanding
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
- WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models