Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence
Aman Tyagi, Hemanth Boinpally, Jonathan Chen, Douglas Gebert, Steven Hickson
Google Cloud · Google DeepMind
cs.CV, cs.AI, cs.MM
Submitted: 2026-08-12
Updated: 2026-08-13
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 95/100
The gist: The paper introduces the “Agentic Self-Improvement” framework, a closed-loop, goal-directed optimization system designed to address the lack of fine-grained control and reliability in modern
Terminology
Summary
The paper introduces the “Agentic Self-Improvement” framework, a closed-loop, goal-directed optimization system designed to address the lack of fine-grained control and reliability in modern black-box Image-to-Video (I2V) models. The framework reframes video synthesis from a speculative, trial-and-error process into a systematic optimization problem by targeting the two primary sources of variance: the textual prompt and the model’s hyperparameters.
The framework operates in two stages. In the first stage, an iterative prompt optimization loop uses a multimodal Large Language Model (mLLM) to refine the input prompt. This refinement is guided by two automated evaluations: Davidsonian Scene Graph (DSG) queries, which ensure semantic adherence by deconstructing the scene into verifiable questions about agents, actions, and locations, and Common Mistake Questions (CMQ), which detect typical generation artifacts such as anatomical inconsistencies, object permanence issues, and unnatural transitions. The mLLM’s reliability as an automated evaluator was validated against human-annotated ground truth, achieving 92% accuracy on DSG questions and 82% on CMQ questions, with an overall accuracy of 87%.
In the second stage, Bayesian optimization is used to efficiently co-optimize stochastic seeds and Classifier-Free Guidance (CFG) scales, with CFG explored within a range of [1 to 15]. This search is guided by a multi-objective reward function composed of quality metrics, including perceptual quality models (RAHF and UVQ), temporal coherence metrics (shotcut detection and loop detection), and a novel Video-Text Adherence (VTA) score derived from the DSG and CMQ evaluations. The VTA metric aggregates the mLLM’s binary responses across the question trees, with a recursive masking scheme that ensures foundational concepts exert a higher penalty upon failure.
The framework was evaluated through human preference studies using 100 image-prompt pairs from V-Bench, with Veo 2.0 as the experimental vehicle. The prompt optimization module alone showed that videos generated from optimized prompts were preferred 27% of the time versus 5% for the baseline, with 68% ties. In the end-to-end evaluation, the framework consistently outperformed unguided random search: with a 10-generation budget, win rates reached 42-47% versus 14% for the baseline, and with a 100-generation budget, win rates climbed to 60-69%. Against a stronger “Best-of-Random” baseline, the Bayesian optimizer still decisively outperformed it, achieving win rates of 38-42% versus 8-15% for the baseline. The paper notes that sorting by the RAHF score yielded the highest human preference (69% win rate), suggesting it is the metric most aligned with human perception.
The paper concludes that this structured, closed-loop process significantly improves predictability and control of state-of-the-art video generation models, moving the field beyond speculative tools toward reliable, production-ready systems. The authors acknowledge limitations including significant computational cost and the two-stage nature of the optimization, suggesting future work could explore joint optimization of prompts and parameters, as well as distilling the mLLM critic into a smaller reward model.
Improvements for AI systems
Improvements to AI Systems:
-
Add a Self-Correcting Prompt Refinement Module – Integrate the iterative mLLM-based prompt optimization loop into text-to-video and image-to-video pipelines. The system will automatically rewrite user prompts to reduce ambiguity and preempt common generation artifacts, improving semantic adherence without user intervention.
-
Implement a Dual-Objective Quality Evaluator – Use the Davidsonian Scene Graph (DSG) queries and Common Mistake Questions (CMQ) as a built-in, automated critic. The AI system will generate a structured question tree for each scene, then score its own output on agent-action-location consistency and artifact presence (e.g., anatomical errors, object permanence), enabling real-time rejection or regeneration of low-quality frames.
-
Add a Bayesian Hyperparameter Co-Optimizer – Replace manual seed and CFG tuning with a Bayesian search that jointly optimizes stochastic seeds and classifier-free guidance scales (range 1–15). The system will automatically adapt these parameters per input, maximizing a multi-objective reward (perceptual quality, temporal coherence, and video-text adherence) within a fixed generation budget.
-
Introduce a Video-Text Adherence (VTA) Scoring Layer – Incorporate the recursive masking scheme from the paper into any video generation model. The system will compute a VTA score that penalizes failures of foundational scene concepts more heavily, allowing downstream modules (e.g., ranking, filtering, or reinforcement learning) to prioritize semantically faithful outputs over visually appealing but misaligned ones.
-
Enable Budget-Aware Generation Strategy – Use the paper’s findings on win rates versus generation budget to implement an adaptive search: with a low budget (e.g., 10 generations), the system will prioritize prompt optimization over parameter search; with a higher budget (e.g., 100), it will allocate more resources to Bayesian parameter co-optimization, improving final output quality.
-
Build a Distilled Reward Model for Efficiency – Train a smaller, faster reward model that mimics the mLLM critic’s DSG/CMQ judgments (87% accuracy). This distilled model can be embedded into real-time video generation systems, providing continuous feedback during inference without the computational overhead of a large multimodal LLM.
What the Improved AI System Can Do:
-
Generate videos from a single image and a vague prompt, automatically refining the prompt to ensure the output matches the intended agents, actions, and locations.
-
Self-evaluate each generated video for semantic correctness and common visual artifacts, then regenerate only the failing segments.
-
Dynamically adjust its internal sampling parameters (seed and CFG) per input, achieving higher human preference (up to 69% win rate) without manual tuning.
-
Rank multiple candidate videos using a perception-aligned score (RAHF-based) to select the most preferred output, even when the user provides no explicit preference.
-
Operate under strict computational budgets, intelligently balancing prompt refinement and parameter search to maximize output quality within a given number of generations.
-
Run on edge devices or real-time pipelines by using the distilled reward model for fast, approximate quality checks, while reserving the full mLLM critic for offline fine-tuning or high-stakes generation.
Abstract
Modern black-box Image-to-Video (I2V) models offer powerful capabilities in automated content creation, yet their lack of fine-grained control and reliability presents significant challenges in professional workflows. Their inherent stochasticity causes minor variations in textual prompts or hyperparameters to yield drastically different outputs often necessitating inefficient, brute-force trial-and-error processes. To address these limitations, we introduce the ``Agentic Self-Improvement" framework, which reframes video synthesis into a closed-loop, goal-directed optimization. Our framework systematically navigates the generation parameter space using a novel two-stage approach. In the first stage, an iterative prompt optimization loop uses a multimodal Large Language Model (mLLM) to refine the input prompt. This refinement implements two automated evaluations: Davidsonian Scene Graph (DSG) queries ensure semantic adherence, and Common Mistake Questions (CMQ) for artifact detection. At the second stage, we use Bayesian optimization to efficiently co-optimize stochastic seeds and CFG scales. This search is guided by a suite of quality metrics, including the novel Video-Text Adherence (VTA) score derived from the DSG and CMQ evaluations. Our framework significantly outperforms unguided search methods: in human preference studies, videos generated via our agentic approach were strongly preferred over baseline outputs, achieving win rates up to 69%. This work provides a practical and extensible methodology for enhancing the predictability and control of state-of-the-art video generation models, moving the field beyond speculative curiosities toward reliable, production-ready tools.
Sources
- Sora as a World Model? A Complete Survey on Text-to-Video Generation
- From Sora What We Can See: A Survey of Text-to-Video Generation
- Video-to-Video Synthesis
- Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k
- We need to talk about random seeds
- Prompt-to-Prompt Image Editing with Cross Attention Control
- Control-A-Video: Controllable Text-to-Video Diffusion Models with Motion Prior and Reward Feedback Learning
- Automated Black-box Prompt Engineering for Personalized Text-to-Image Generation
- Classifier-Free Diffusion Guidance
- MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action
- DialogGen: Multi-modal Interactive Dialogue System for Multi-turn Text-to-Image Generation
- The Rise and Potential of Large Language Model Based Agents: A Survey
- The Vizier Gaussian Process Bandit Algorithm
- VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models
- Gemini Robotics: Bringing AI into the Physical World
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models