VidHarness: Evolving Agent Harnesses for Cost-Efficient Long Video Understanding
cs.CV
Submitted: 2026-09-29
Updated: 2026-09-29
Terminology
Sources
- Qwen2.5-VL Technical Report
- FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance
- HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry
- Exploring the Effect of Reinforcement Learning on Video Understanding: Insights from SEED-Bench-R1
- Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- MetaVideoAgent: Automated Video-Agent Evolution for Long-Form Video Understanding
- Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution
- From Question Answering to Task Completion: A Survey on Agent System and Harness Design
- GPT-4o System Card
- AIDE: AI-Driven Exploration in the Space of Code
- Meta-Harness: End-to-End Optimization of Model Harnesses
- VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
- AutoHarness: improving LLM agents by automatically synthesizing a code harness
- Recursive Agent Harnesses
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
- MR. Video: "MapReduce" is the Principle for Long Video Understanding
- LLM-as-Code: Agentic Programming for Agent Harness
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models