CaC: Advancing Video Reward Models via Hierarchical Spatiotemporal Concentrating
cs.CV, cs.AI
Submitted: 2026-05-12
Updated: 2026-09-25
Terminology
Sources
- A Survey on Video Anomaly Detection via Deep Learning: Human, Vehicle, and Environment
- Qwen3-VL Technical Report
- VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation
- PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation
- CropVLM: Learning to Zoom for Fine-Grained Vision-Language Perception
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Investigating Text Insulation and Attention Mechanisms for Complex Visual Text Generation
- Sculpting Features from Noise: Reward-Guided Hierarchical Diffusion for Task-Optimal Feature Transformation
- RePer-360: Releasing Perspective Priors for 360$^\circ$ Depth Estimation via Self-Modulation
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning
- LTX-2: Efficient Joint Audio-Visual Foundation Model
- VideoScore2: Think before You Score in Generative Video Evaluation
- EVE: Verifiable Self-Evolution of MLLMs via Executable Visual Transformations
- GPT-4o System Card
- HunyuanVideo: A Systematic Framework For Large Video Generative Models
- Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning
- DiffPCN: Latent Diffusion Model Based on Multi-view Depth Images for Point Cloud Completion
- DRL4AOI: A DRL Framework for Semantic-aware AOI Segmentation in Location-Based Services
- Flow-GRPO: Training Flow Matching Models via Online RL
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models