Hard Vision, Easy Vision: What GPT-6 Astra Reveals Across Computer Vision
cs.CV
Submitted: 2026-09-28
Updated: 2026-09-28
Terminology
Sources
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models
- UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and Texture
- Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget
- BabyVision: Visual Reasoning Beyond Language
- Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios
- TRACE: Temporal Grounding Video LLM via Causal Event Modeling
- Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs
- WildDet3D: Scaling Promptable 3D Detection in the Wild
- OpenAI o1 System Card
- GenEval 2: Addressing Benchmark Drift in Text-to-Image Evaluation
- Depth Anything 3: Recovering the Visual Space from Any Views
- PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models
- VLRS-Bench: A Vision-Language Reasoning Benchmark for Remote Sensing
- DINOv2: Learning Robust Visual Features without Supervision
- The 2017 DAVIS Challenge on Video Object Segmentation
- ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
- VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain Knowledge
- CountQA: How Well Do MLLMs Count in the Wild?
- Gemini: A Family of Highly Capable Multimodal Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models