Dual-Stream Cross-Anchor Correction Grounding Long-Form Captions and the Domain Limits of Object-Level Anchors

arXiv:2608.12746 · cs.CV, cs.CL · Submitted 2026-08-19 · Read on arXiv

Lingkai Bu, Qian Gao, Jun Fan, Guohui Ding, Zhenyu Yang, Yuteng Xiao, Jinyi Liang

Qilu University of Technology (Shandong Academy of Sciences)

cs.CV, cs.CL

Submitted: 2026-08-19

Updated: 2026-08-20

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 75/100

The gist: Dual-Stream Cross-Anchor Correction (DSCC) is a fine-tuning framework proposed to address object hallucination in multimodal large language models (MLLMs), where models confidently describe objects

Terminology

Summary

Dual-Stream Cross-Anchor Correction (DSCC) is a fine-tuning framework proposed to address object hallucination in multimodal large language models (MLLMs), where models confidently describe objects not present in the image. The paper argues that existing mitigation methods are trapped in a short-caption regime (roughly 90-105 words) and that supervised fine-tuning (SFT) on detail-rich corpora lengthens captions but leaves hallucination rates high, with 41.60% of captions still containing a hallucinated object. The central research question is how can a model say more while getting less wrong?

DSCC is the first method to inject object-level visual anchors into the language model itself during fine-tuning, rather than intervening at decoding time. It adds two auxiliary streams on top of standard visual instruction tuning:

  1. Perception stream: At layer 16, an object-level InfoNCE loss aligns local features with frozen CLIP text anchors, constructing fine-grained visual anchors.

  2. Cognition stream: In deeper layers (24 and 28), a cross-attention mechanism is introduced where the deep hidden state actively queries the visual anchors at every forward step of autoregressive generation.

  3. Curriculum-gated fine-tuning (CGFT): A two-stage schedule couples the streams progressively, with the perception stream establishing stable anchors before the cognition stream begins querying them.

The training objective is L = L SFT + α·L perc, where L SFT alone governs the output distribution (length, object density, coverage), while the dual-stream modules refine precision without reshaping length or coverage.

Experiments use LLaVA-1.5-7B as backbone, trained on ShareGPT4V long captions intersected with COCO object annotations (about 95k samples), with four configurations compared: D (both streams off, equivalent to vanilla SFT), A (perception-only), B (cognition-only), and C (full DSCC).

Key results include:

  • CHAIR-500: DSCC (C) attains CHAIR S = 38.80, 2.8 percentage points below the vanilla SFT baseline D (41.60), with caption length staying at roughly 170 words across all configurations. DSCC produces captions about 1.9 times longer than baselines (171.5 words vs. 89.5-104.9) while achieving the highest precision per object mention under a density-independent criterion: 88.19% against OPERA's 86.73%.

  • POPE Adversarial subset: Precision rises monotonically as streams are added: A (0.8315) < D (0.8510) < B (0.8638) < C (0.8839), with C improving on D by 3.3 percentage points. The paper notes F1 is essentially flat (0.838), making precision the informative metric.

  • Length-aware analysis: DSCC is the only method landing in the long-caption, low-hallucination region of the length-quality plane, driving CHAIR S to its lowest value while producing captions about 1.9 times as long.

The ablation reveals a synergy: the perception stream alone degrades precision (A has the highest YesRatio 0.5023 and lowest precision), yet stacked on the cognition stream it pushes adversarial precision to the highest value. The paper describes this as harmful on its own, yet stacked on the cognition stream it pushes precision to its highest value, demonstrating an interaction rather than linear additivity.

The paper makes a two-level attribution: the lengthening of captions comes from the SFT data paradigm (ShareGPT4V), not the dual-stream architecture, while the net architectural gain (D → C) is real but modest: CHAIR S −2.8, CHAIR I −0.61, and POPE Adversarial precision +3.3.

Crucially, the paper claims no universal superiority. Out-of-domain evaluations reveal a predictable, falsifiable domain-conditionality: the synergy holds only within the COCO object semantic domain. On MME (out-of-distribution but same semantic domain), C is best (588.33 vs. D's 460.00). On HallusionBench (charts and optical illusions, genuinely out of domain), the synergy breaks and B (cognition-only) overtakes C. On MMHal-Bench (Flickr and abstract scenes), the four configurations are statistically indistinguishable, reported as a null result.

The paper concludes that the cognition stream is the domain-independent workhorse bringing positive gains on every benchmark, while the perception stream is domain-conditional, bound hard to COCO object semantics by its CLIP-COCO text anchors. The value of the work lies not in state-of-the-art numbers but in the combination of a new perspective, a complete empirical account and a predictable boundary of validity.

Acknowledged boundaries include: the model adopts a conservative stance (POPE recall 0.797 vs. D's 0.826; trading 15 percentage points of object recall against OPERA for higher precision per mention), the full model is not strictly best on instance-level CHAIR I (B beats C by 0.47), the synergy breaks outside COCO semantics, logical hallucination receives no explicit signal, no layer-by-layer sweep was performed, all differences are point estimates without significance testing, and OOD evaluations use a different scoring protocol from official leaderboards.

Improvements for AI systems

Improvements to AI systems:

  1. Add a dual-stream architecture with object-level visual anchors injected into the language model – The perception stream (InfoNCE loss at layer 16 aligning local features with frozen CLIP text anchors) and cognition stream (cross-attention at layers 24/28 querying those anchors during autoregressive generation) let the model actively verify each generated object mention against visual evidence, reducing hallucination without changing output length or coverage.

  2. Implement curriculum-gated fine-tuning (CGFT) – Train the perception stream first to establish stable visual anchors, then activate the cognition stream. This staged coupling prevents the cognition stream from querying unstable anchors, yielding the observed synergy where perception alone hurts precision but boosts it when stacked on cognition.

  3. Decouple length control from precision control – Use standard SFT on detail-rich corpora (e.g., ShareGPT4V) to drive caption length (up to 170 words), while the dual-stream modules refine precision independently. The loss L = L SFT + α·L perc ensures the output distribution (length, object density, coverage) is governed solely by SFT, not by the auxiliary streams.

  4. Add a domain-conditionality guard – The system should detect whether the input image falls within the training semantic domain (e.g., COCO objects). If yes, activate both streams (full DSCC). If out-of-domain (e.g., charts, optical illusions), fall back to cognition-only mode (stream B), which is the domain-independent workhorse. This prevents the perception stream’s COCO-bound anchors from degrading performance on unseen semantics.

  5. Enable a conservative precision-recall trade-off knob – The system can explicitly trade object recall for higher precision per mention (e.g., POPE recall 0.797 vs. 0.826 baseline, but precision 88.19% vs. 86.73%). This is useful for applications where false positives are more costly than missed objects (e.g., medical imaging, surveillance).

What the improved AI system can do:

  • Generate long, detailed image captions (170 words) with significantly lower object hallucination rates (CHAIR S 38.80 vs. 41.60 baseline) and higher precision per object mention (88.19%).

  • Maintain flat F1 on adversarial object existence questions while improving precision by 3.3 percentage points over vanilla SFT (POPE Adversarial).

  • Automatically adapt its verification strategy: use full dual-stream synergy for COCO-like scenes, cognition-only for out-of-domain images, and gracefully degrade to baseline performance on abstract or illusion-heavy inputs (no catastrophic failure).

  • Provide a tunable precision-recall trade-off, allowing deployment in high-stakes domains where hallucinated objects are unacceptable, even at the cost of missing some real objects.

  • Avoid the short-caption trap: it can say more without getting more wrong, breaking the prior trade-off between detail and accuracy.

Abstract

Object hallucination in multimodal large language models arises when language priors and corpus co-occurrence bias outweigh the visual evidence, with nothing tying an individual object mention to what the image shows. Most remedies intervene at decoding time without training, yet under a unified protocol their benefit is confined to short captions;supervised fine-tuning (SFT) on a detail- rich corpus lengthens captions, but over forty percent still name absent objects. This paper proposes Dual-Stream Cross-Anchor Correction (DSCC). Unlike work that post-processes decoding, DSCC is the first to inject object-level visual anchors into the language model itself during fine- tuning: a perception stream aligns object-level hidden states at an intermediate layer to frozen text anchors by a bidirectional contrastive objective; a cognition stream lets deeper layers query those anchors by cross-attention at every generation step; and a two-stage curriculum gate couplesthem, making evidence retrieval a structural constraint at each autoregressive step. Under one backbone and one scoring protocol, experiments span long-caption hallucination, object-existence discrimination and cross-domain generalisation, with vanilla SFT on the same corpus and schedule as a length- and density-matched control, so gains are attributed layer by layer. DSCC is the only method reaching the long-caption, low-hallucination region: captions roughly 1.9 times the baseline length at 88.19% precision per object mention, the highest under a density-independent criterion. Ablations expose a synergy: the perception stream alone degrades precision yet reverses sign when stacked on the cognition stream. No universal superiority is claimed: three out-of- domain benchmarks yield a predictable, falsifiable domain-conditionality, the synergy being bound to the anchors' semantic domain and breaking on charts and optical illusions.

Sources

Related papers