DataVista: Diagnosing Multimodal LLMs on Data Video Understanding
cs.CV, cs.AI, cs.CL
Submitted: 2026-10-08
Updated: 2026-10-08
Code: https://github.com/HKUSTDial/DataVista
Terminology
Sources
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- You Don't Know Until You Click:Automated GUI Testing for Production-Ready Software Evaluation
- HumanVideo-MME: Benchmarking MLLMs for Human-Centric Video Understanding
- Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding
- VisTR: Visualizations as Representations for Time-series Table Reasoning
- DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces
- VideoAesBench: Benchmarking the Video Aesthetics Perception Capabilities of Large Multimodal Models
- nvBench 2.0: Resolving Ambiguity in Text-to-Visualization through Stepwise Reasoning
- Exploring Agentic Visual Analytics: A Co-Evolutionary Framework of Roles and Workflows
- Neptune: The Long Orbit to Benchmarking Long Video Understanding
- Data Playwright: Authoring Data Videos with Annotated Narration
- How Does Empirical Research Facilitate Creation Tool Design? A Data Video Perspective
- DeepVIS: Bridging Natural Language and Data Visualization Through Step-wise Reasoning
- IGenBench: Benchmarking the Reliability of Text-to-Infographic Generation
- sketch-plot: Progressive Editing for Text-to-Image Academic Figures
- ViviDoc: Generating Interactive Documents through Human-Agent Collaboration
- Demonstrating chart-plot: Closing the Last Mile of Academic Chart Generation
- TableTale: Reviving the Narrative Interplay Between Data Tables and Text in Scientific Papers
- LVBench: An Extreme Long Video Understanding Benchmark
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models