Token-Level Video Reinforcement Learning
cs.CV
Submitted: 2026-10-01
Updated: 2026-10-01
Terminology
Sources
- Qwen2.5-VL Technical Report
- Directly Fine-Tuning Diffusion Models on Differentiable Rewards
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Improving Dynamic Object Interactions in Text-to-Video Generation with AI Feedback
- VideoScore2: Think before You Score in Generative Video Evaluation
- HunyuanVideo: A Systematic Framework For Large Video Generative Models
- TeleBoost: A Systematic Alignment Framework for High-Fidelity, Controllable, and Robust Video Generation
- Flow-GRPO: Training Flow Matching Models via Online RL
- Improving Video Generation with Human Feedback
- Seeing What Matters: Visual Preference Policy Optimization for Visual Generation
- Movie Gen: A Cast of Media Foundation Models
- Aligning Text-to-Image Diffusion Models with Reward Backpropagation
- Video Diffusion Alignment via Reward Gradients
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Wan: Open and Advanced Large-Scale Video Generative Models
- Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models
- Unified Reward Model for Multimodal Understanding and Generation
- Diffusion-DRF: Free, Rich, and Differentiable Reward for Video Diffusion Fine-Tuning
- HunyuanVideo 1.5 Technical Report
- DenseDPO: Fine-Grained Temporal Preference Optimization for Video Diffusion Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models