Order Matters: LVLMs as Judges for Temporal Reasoning in Image Sequences
Martina Ianaro, Guilherme Fernandes, Maurizio Gabbrielli, Joao Magalhaes
University of Bologna · NOVA School of Science and Technology · NOVA Laboratory for Computer Science and Informatics
cs.CV, cs.CL
Submitted: 2026-08-11
Updated: 2026-08-12
Comments: 34 pages, camera-ready
License: http://creativecommons.org/licenses/by-sa/4.0/
Importance score: 100/100
The gist: This paper identifies a foundational bottleneck in generative multimedia evaluation: the "judgment crisis." As the authors state, "While human perception naturally synthesizes the temporal and
Terminology
Summary
This paper identifies a foundational bottleneck in generative multimedia evaluation: the judgment crisis.
As the authors state, "While human perception naturally synthesizes the temporal and logical flow of a story, automated evaluation systems remain largely 'blind' to sequential continuity, often failing to distinguish between a coherent narrative and a semantically shuffled or contradictory sequence."
The paper argues that current evaluation paradigms suffer from the Bag of Frames
Fallacy: "the evaluation of multi-image sequences remains tethered to 'pointwise' metrics, such as FID or CLIPScore, which assess images in total isolation. These paradigms are effectively 'blind' to order; they cannot distinguish a logically sound visual progression from a semantically identical set of frames that have been randomly shuffled."
The central discovery is a startling performance dichotomy
: "even state-of-the-art LVLMs, when acting as judges, exhibit a profound 'reasoning chasm.' While these models assign plausible scores to individual sequences in isolation, their ability to discriminate between ordered and shuffled sequences collapses when required to perform pairwise discrimination."
This failure is rooted in structural blindness: positional asymmetries, specifically primacy and recency effects, significantly influence a model's judgment, often more than the actual logical flow of the images.
The paper provides theoretical grounding for these biases in transformer architectures:
-
Causal masking induces primacy bias:
as network depth L increases, hidden representations tend to converge exponentially toward the first token,
renderingearly content structurally more salient.
-
Rotary Position Embeddings (RoPE) exert recency bias: RoPE
suppress interactions between distant tokens, leaving the logical core of a narrative, in particular the middle steps of a sequence, theoretically invisible to the judge.
The authors note these biases are further amplified by instruction tuning and RLHF,
and as these biases operate at the representation level and are not surfaced during CoT reasoning, they remain resistant to systematic mitigation via reasoning augmentation.
The paper releases two benchmarks:
PRISM (Perturbation-based Reasoning samples for Image Sequence Modeling): designed to evaluate temporal and logical consistency through a multi-level perturbation strategy
using cooking procedures. It includes:
-
PRISM-Semantic:
1–3 frames are replaced with out-of-context images
-
PRISM-Temporal:
frames are reordered via consecutive or non-consecutive swaps
-
600 gold-standard instances yielding
over 4,620 model evaluations
MIRAGE (Model-generated Image sequences with Real And hallucinated Generated Examples): provides a realistic evaluation setting grounded in actual generative model outputs
with negatives from seven distinct pipelines (e.g., GILL, StoryDiffusion, BeamDiffusion) across FLUX.1 and SD 2.1 backbones,
spanning MIRAGE-Recipes (procedural) and MIRAGE-Vist (narrative), totaling 672 pairwise comparisons.
The Scoring Illusion: LLaVA-Critic tracks the Gemini oracle most closely but diverges from human raters, while LLaVA-OneVision shows the reverse pattern across every prompt.
The authors note Gemini concentrates 91% of its scores in the 1, 2 range... a pessimistic collapse.
The Accuracy Collapse: On MIRAGE, LLaVA-Critic and LLaVA-OneVision reach comparable F1 scores (0.56)
despite a massive response-length disparity (825 vs. 19 tokens).
On PRISM, LLaVA-OneVision takes a clear lead (Acc 0.64 vs. 0.43), indicating superior sensitivity to controlled visual inconsistencies, yet both remain far below human-level discrimination.
SFT and Calibration Gains: "Fine-tuning LLaVA-OneVision (OneVision†) yields consistent gains across all Gemini-referenced metrics, reducing MAEG from 2.21 to 1.96. Against human judgments, OneVision† reaches rH=0.34, perfectly recovering the human interrater ceiling."
SFT Scaling and the Limits of Mitigation: The multi-prompt variant (OneVision‡) outperforms single-prompt training by +0.03 F1, achieving the benchmark peak (Acc 0.72, F1 0.47).
However, fine-tuning does not fully close the gap with human performance, suggesting that the underlying architectural biases... remain a fundamental bottleneck that requires new model designs rather than simply larger datasets.
Structural Blindness: "Many models show disproportionate reliance on the first frame to determine overall quality, ignoring subsequent logical contradictions (primacy bias); conversely, middle-sequence information is often invisible to the judge (recency bias and lost-in-the-middle effect). Critically,
these asymmetries persist even after fine-tuning, confirming that SFT consolidates rather than mitigates structural bias."
-
Semantic Robustness vs. Temporal Fragility:
All model variants fall below the F1=0.40 threshold
on PRISM-Temporal, withSFT providing only marginal gains over zero-shot baselines.
-
Positional Bias:
On PRISM-Semantic, we observe a consistent primacy effect: performance peaks at the initial position (pos0, Acc=0.70) and degrades monotonically toward later positions.
Conversely,PRISM-Temporal exhibits a recency effect, where swaps at later positions are marginally easier to detect.
-
Scaling vs. Reasoning:
Zero-shot evaluations on Gemma4-31B and Qwen3.6-27B reproduce every failure mode observed at 7B, from severe answer invalidity to the same RoPE-consistent positional signature.
The paper includes a diagnostic task where models must order four historical images by date using visual content alone, with no textual scaffolding.
Results show:
-
LLaVA-OneVision: 0/10, LLaVA-Critic: 1/10
-
Qwen3.6-27B: 3/10, Gemma-4-31B: 7/10
The authors note: LLaVA-Critic returns D-C-B-A, the reverse of the label order, in 8 of 10 sequences regardless of what the images depict. This is the positional heuristic... in its purest form.
The authors conclude: "We have challenged the prevailing 'pointwise' evaluation paradigm in generative multimedia: by treating multi-image sequences as a 'bag of frames,' the community prioritises visual fidelity while remaining blind to the temporal failures that undermine narrative coherence."
They call for "Temporally-Aware Evaluation: metrics that assess the interleaved logic of a sequence as a whole, judges that ground decisions in per-step causal rationales rather than visual similarity, and architectural redesign rather than data scaling to mitigate positional drift."
The paper's three key contributions are:
-
Analysis of the Reasoning Chasm: We uncover the performance dichotomy between pointwise and pairwise judging
-
Identification of Structural Biases: We reveal that systematic positional asymmetries persist even after fine-tuning
-
A Visual-Order Reasoning Tasks: We release the first dataset to assess LVLM-judges accuracy in sequential multi-image reasoning tasks
The final message: The era of judging generative AI by its beauty must give way to judging it by its logic.
Improvements for AI systems
Improvements to AI Systems:
-
Order-Aware Evaluation Module: Implement a dedicated temporal-consistency checker that explicitly compares frame sequences against shuffled permutations during inference. The system can detect narrative incoherence by computing pairwise order-discrimination scores, flagging sequences where the model cannot distinguish original vs. perturbed order.
-
Positional Bias Correction Layer: Add a post-hoc calibration mechanism that reweights token contributions based on learned positional importance profiles. The system can compensate for primacy/recency biases by down-weighting first/last frame influences and up-weighting middle-sequence information during judgment tasks.
-
Causal Rationale Grounding: Require the judge to output explicit per-step causal justifications (e.g.,
Frame 3 contradicts Frame 2 because ingredient X appears before preparation step Y
) before producing a final score. This forces the model to surface structural reasoning that would otherwise remain implicit and biased. -
Multi-Prompt Ensemble with Conflict Detection: Run evaluations under multiple prompt formulations (e.g.,
rate quality
vs.detect order violation
) and flag cases where responses diverge significantly. The system can then trigger deeper analysis or abstain from judgment when prompt-induced variance exceeds a threshold. -
Temporal Attention Redistribution: Modify the transformer's attention mechanism during fine-tuning to explicitly penalize attention patterns that ignore middle-sequence tokens. The system learns to allocate more attention to intermediate frames, reducing lost-in-the-middle effects.
-
Synthetic Perturbation Training: Augment training data with automatically generated order-perturbed sequences (swaps, replacements) and train the judge to output both a score and a binary
order-valid
flag. This dual-output training forces the model to internalize temporal consistency as a separate axis from visual quality. -
Human-Ceiling Calibration: During fine-tuning, optimize not just for oracle agreement but for matching human interrater reliability distributions. The system adjusts its score dispersion to avoid pessimistic collapse (e.g., all scores in 1,2) or overconfident extremes.
-
Architectural Bias Profiler: Add a diagnostic tool that runs the chronological ordering probe (e.g., sorting historical images by date) before deployment. The system can report its positional heuristic tendency (e.g.,
reverses order 80% of the time
) and automatically adjust its judgment confidence accordingly.
What the Improved System Can Do:
-
Reliably distinguish a coherent narrative from a semantically identical but shuffled sequence, achieving near-human discrimination on temporal tasks (target: >0.85 F1 vs. current 0.56).
-
Explain its judgments with per-frame causal rationales, allowing users to audit why a sequence was deemed inconsistent.
-
Self-correct for positional bias: When asked to judge a sequence, it can explicitly state
I am over-weighting the first frame; re-evaluating middle frames now
and adjust its score. -
Detect its own blindness: If middle-sequence information is invisible, the system flags the evaluation as low-confidence rather than producing a misleading score.
-
Generalize across domains: The order-aware evaluation works for procedural tasks (recipes), narrative tasks (stories), and even historical chronological ordering, without task-specific fine-tuning.
-
Resist prompt manipulation: The multi-prompt ensemble ensures that a single phrasing change cannot flip the judgment from
coherent
toincoherent.
-
Quantify its own bias: The system can output a
positional bias index
(e.g., 0.7 primacy, 0.2 recency) alongside each judgment, enabling downstream users to interpret reliability. -
Achieve human-level interrater agreement (rH ≈ 0.34) while maintaining robustness to architectural scaling, meaning the improvements persist from 7B to 31B parameter models.
Abstract
As generative multimedia evolves from static image synthesis to complex, interleaved visual narratives, a foundational bottleneck has emerged: the judgment crisis. While human perception naturally synthesizes the temporal and logical flow of a story, automated evaluation systems remain largely "blind" to sequential continuity, often failing to distinguish between a coherent narrative and a semantically shuffled or contradictory sequence. This work identifies a critical structural gap in current multimodal evaluation paradigms, arguing that the reliance on Large Vision-Language Models (LVLMs) as judges is fundamentally limited by architectural biases. Our analysis reveals a profound performance dichotomy: while models may appear competent in isolated pointwise scoring, they suffer a catastrophic collapse when required to perform pairwise discrimination of temporal order. We demonstrate that this is not merely a data-scarcity issue but a structural one. Through a series of diagnostic probes, we uncover systematic positional asymmetries, specifically primacy and recency effects, where a model's judgment of a story is significantly influenced by the placement of a frame, often more than by its semantic consistency. These biases, potentially rooted in causal masking and rotary embeddings, suggest that current transformer-based judges are inherently ill-equipped for long-form visual reasoning. By exposing these blind spots, we challenge the multimedia community to move beyond snapshot-centric metrics and instead pioneer Temporally-Aware Evaluation paradigms that treat visual sequences as unified logical structures rather than unordered collections of frames.
Sources
- Fairness and Bias in Multimodal AI: A Survey
- Beyond Single-Point Judgment: Distribution Alignment for LLM-as-a-Judge
- From Captions to Rewards (CAREVL): Leveraging Large Language Model Experts for Enhanced Reward Modeling in Large Vision-Language Models
- More is Less? A Simulation-Based Approach to Dynamic Interactions between Biases in Multimodal Models
- Latent Beam Diffusion Models for Generating Visual Sequences
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Gemma 4 Technical Report
- Unblocking Fine-Grained Evaluation of Detailed Captions: An Explaining AutoRater and Critic-and-Revise Pipeline
- When Attention Sink Emerges in Language Models: An Empirical View
- MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos
- LoRA: Low-Rank Adaptation of Large Language Models
- A Survey on Evaluation of Multimodal Large Language Models
- Bridging and Modeling Correlations in Pairwise Data for Direct Preference Optimization
- Generating Images with Multimodal Language Models
- Prometheus-Vision: Vision-Language Model as a Judge for Fine-Grained Evaluation
- LLaVA-OneVision: Easy Visual Task Transfer
- Generative Judge for Evaluating Alignment
- Who Judges the Judge? LLM Jury-on-Demand: Building Trustworthy LLM Evaluation Systems
- Evaluating Text-to-Visual Generation with Image-to-Text Generation
- Lost in the Middle: How Language Models Use Long Contexts
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models