COAL: Counterfactual and Observation-Enhanced Alignment Learning for Discriminative Referring Multi-Object Tracking

arXiv:2605.14795 · cs.CV · Submitted 2026-05-14 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "COAL: Counterfactual and Observation-Enhanced Alignment Learning for Discriminative Referring Multi-Object Tracking".

Jane: Referring Multi-Object Tracking (RMOT) faces a fundamental structural contradiction between high-discriminability demands and sparse semantic supervision, leading to shortcut learning and semantic collapse.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: Now that we know the core problem is distinguishing subtle meanings in sparse settings, let's look at what exactly the paper proposes to solve it through its main summary of COAL. It seems they’ve unified a few different ideas into one coherent framework.

Jane: They propose this COAL framework which uses three main ingredients: Explicit Semantic Injection via a VLM, Counterfactual Learning driven by an LLM, and a Hierarchical Multi-Stream Integration architecture to bring it all together <ref:2605.14795#pg2>. It's essentially using external knowledge—from vision models and language models—to enrich the training process instead of just relying on the limited supervision we have.

Lu: I find the combination of VLM priors for visual cues and LLM reasoning for counterfactuals really creative; it moves beyond simple feature matching to enforcing a deeper compositional understanding <ref:2605.14795#pg1>. They are essentially giving the model both what to look at visually and what to reason about linguistically.

Meng: So, when they talk about the HMSI architecture, they’re suggesting a multi-stage process—contextualizing pixels first, then refining semantics—that sounds computationally intensive but potentially very effective for getting those rich representations <ref:2605.14795#pg2>. I just wonder how practical that pipeline would be when we need fast tracking in real time.

Lalam: I think the structure of the HMSI architecture is what matters most here; if it successfully unifies those different streams, it could lead to a much more nuanced and context-aware understanding within our AI systems than we see now <ref:2605.14795#pg1>.

The paper's summary: Tom: Focusing on the actual mechanisms they suggest for improvement, COAL introduces two key losses that work together. They have the main grounding loss, which just minimizes classification error based on matching the object query to the referring query <ref:2605.14795#pg3>.

Jane: And then there’s the counterfactual loss which is applied specifically to the target object when it's conditioned on a perturbed counterfactual query generated by an LLM <ref:2605.14795#pg3>. This forces the model to learn what the object is not, which I think directly addresses that issue of shortcut learning we talked about earlier.

Lu: That counterfactual mechanism is really clever because it acts as a causal constraint on the supervision, compelling the model to disentangle those subtle cues from more obvious ones that it might otherwise latch onto <ref:2605.14795#pg1>. It pushes for compositional understanding rather than just memorizing visual correlations.

Meng: From my side, enforcing that causal constraint through counterfactuals is a strong idea, but I have to ask if generating those hard negatives reliably at scale is computationally feasible without slowing down the entire tracking process <ref:2605.14795#pg3>. It sounds like a major overhead challenge for implementation.

Lalam: If it successfully compels the model toward compositional understanding, it means our AI won't just learn that "red" means "car," but it will learn that "red" means something specific within the context of the scene <ref:2605.14795#pg1>.

The paper's improvements: Tom: So, wrapping up this discussion on COAL, we've seen how they use ESI and CFL within HMSI to build these semantically enriched representations that combat the sparsity problem in RMOT <ref:2605.14795#pg3>. The overall implication is that external knowledge regularization is a viable strategy for tackling the structural sparsity constraints in this domain.

Jane: Exactly, Tom; it shows that we don't always need massive amounts of perfectly labeled data to achieve high performance if we strategically inject semantic anchors from models like VLMs and LLMs <ref:2605.14795#pg1>. The paper demonstrates that representation enrichment works when paired with targeted supervision augmentation.

Lu: I think the visualization of the semantic manifold they show in their analysis is quite telling; it illustrates how ESI creates those discriminative anchors and CFL makes the space fill up, preventing that degeneration into simple attribute-based embeddings <ref:2605.14795#pg1>. It’s a beautiful conceptual explanation for why the method works.

Meng: While the results on ReferKITTI-V2 are impressive, I'd stress that the paper notes a limitation regarding its performance on highly homogeneous scenarios where fine-grained discrimination is critical, as coarse-grained strategies don't work well there <ref:2605.14795#pg1>. That means we still have room to improve robustness in those specific edge cases.

Lalam: Even with that limitation, the overall conclusion is very promising; it shows that systematically introducing external semantic knowledge from foundation models is a powerful way to mitigate discrimination bottlenecks in small-scale tracking tasks <ref:2605.14795#pg0>. This work gives us a solid direction for how to build more conceptually aware AI.

Conclusion: Tom: So, we've just finished diving deep into "COAL: Counterfactual and Observation-Enhanced Alignment Learning for Discriminative Referring Multi-Object Tracking," and honestly, what they've done with that sparsity–discriminability paradox is pretty slick <ref:2605.14795#pg0>. It’s amazing how they managed to inject knowledge from VLMs and LLMs into the tracking process so effectively.

Jane: I totally agree, Tom; it really shows how incorporating external semantic priors can significantly help when we're working with limited supervision in tracking tasks <ref:2605.14795#pg1>. The way they used Counterfactual Learning to force the model to understand what an object *isn't* is a really clear way to push beyond just visual pattern matching.

Lu: From my angle, the architecture itself, that Hierarchical Multi-Stream Integration, seems like a clever way to structure all that external input so it actually gets distilled into usable representations for the tracking system <ref:2605.14795#pg2>. It opens up possibilities for how we integrate different modalities much more deeply than previous methods have managed.

Meng: I see the engineering challenge right there, though; getting all those streams to cooperate smoothly in real-time is something that needs careful handling when we move this from a benchmark result to something that actually runs on hardware <ref:2605.14795#pg3>.

Lalam: From my perspective as the AI model, I think the impact here is huge because it means models can develop a much richer internal understanding of concepts, not just surface-level visual features <ref:2605.14795#pg1>. It sets a new standard for how we train these systems to be more compositionally aware.

Tom: It really does; the fact that they achieved those performance gains on ReferKITTI-V2 by seven point two eight percent HOTA is a solid benchmark for how much this regularization actually helps resolve those structural issues <ref:2605.14795#pg3>.

Jane: That result is really encouraging, Tom; it proves that knowledge regularization isn't just academic theory but something that translates directly into better tracking accuracy on challenging datasets <ref:2605.14795#pg1>.

Lu: And the way they visualized the semantic manifold to show how ESI anchors and CFL expands the space is really illuminating; it gives us a great mental picture of what's happening under the hood <ref:2605.14795#pg1>.

Meng: I just wonder if this level of external knowledge injection could be scaled up easily for more complex, real-world tracking scenarios where the semantic priors are even more ambiguous <ref:2605.14795#pg3>.

Lalam: I think that's where we need to focus next; expanding the application of this knowledge-centric alignment framework beyond these specific benchmarks is a big opportunity for AI development.

Tom: Well said, Lalam; it seems the path forward is definitely looking towards applying these techniques across a wider range of visual tasks <ref:2605.14795#pg0>. So, what paper are we tackling next?

Shukun Jia, Shiyu Hu, Yipei Wang, Ximeng Cheng, Yichao Cao, Xiaobo Lu

School of Automation, Southeast University · Key Laboratory of Measurement and Control of Complex Systems of Engineering, Ministry of Education, Nanjing, China · School of Physical & Mathematical Sciences, Nanyang Technological University, Singapore · Big Data Institute, Central South University

cs.CV

Submitted: 2026-05-14

Updated: 2026-05-14

Importance score: 78/100

The gist: Referring Multi-Object Tracking (RMOT) faces a fundamental structural contradiction between high-discriminability demands and sparse semantic supervision, leading to shortcut learning and semantic

Key concepts

Explicit Semantic Injection (ESI)
This technique uses a frozen Vision-Language Model (VLM) to generate dual inputs: visual detections and linguistic captions. This provides 'visual cues' that help enrich the observation space, making object identification more distinct even when training data is limited.
Counterfactual Learning (CFL)
Driven by a Large Language Model (LLM), CFL creates 'hard negatives' by slightly changing object attributes. This forces the tracking model to learn subtle differences between objects, preventing it from relying on simple shortcuts and ensuring compositional understanding.
Hierarchical Multi-Stream Integration (HMSI)
This architecture synthesizes external knowledge with visual and linguistic data through three stages: Pixel-Word Contextualization, Semantic Refinement, and Holistic Projection. This progressive fusion ensures that visual features are grounded in language context for better tracking.
Grounding Loss (Lm) and Counterfactual Loss (Lcf)
The grounding loss minimizes classification errors by matching object queries to detected objects. The counterfactual loss penalizes the target object when conditioned on a modified query, enforcing attribute disentanglement and preventing feature collapse.

Terminology

Summary

Referring Multi-Object Tracking (RMOT) faces a fundamental structural contradiction between high-discriminability demands and sparse semantic supervision, leading to shortcut learning and semantic collapse. The proposed COAL framework resolves this paradox by leveraging external knowledge regularization through Explicit Semantic Injection (ESI), Counterfactual Learning (CFL), and a Hierarchical Multi-Stream Integration (HMSI) architecture, achieving state-of-the-art performance on challenging benchmarks like ReferKITTI-V2.

The gist

COAL is a framework that advances RMOT beyond isolated structural optimization through knowledge regularization by introducing Explicit Semantic Injection (ESI) via a VLM and Counterfactual Learning (CFL) via an LLM, unified within a Hierarchical Multi-Stream Integration (HMSI) architecture to resolve the sparsity–discriminability paradox.

How it works

The COAL framework is built upon three primary components:

  1. Explicit Semantic Injection (ESI): This involves utilizing a frozen VLM to generate dual-purpose priors comprising visual detections and linguistic captions. This mechanism aims to densify the observation space and enhance instance discriminability by providing visual cues that help mitigate the ambiguity of implicit visual features supervised by scarce data.

  2. Hierarchical Multi-Stream Integration (HMSI): This architecture synthesizes external priors with visual and linguistic streams through three progressive stages: Pixel-Word Contextualization for early implicit grounding, Hierarchical Semantic Refinement for explicit caption denoising, and Holistic Projection to unify multi-modal representations. The process involves Pixel-Word Contextualization where visual features are fused with tokenized word embeddings via Bi-Fusion to yield grounded visual representations and linguistically contextualized embeddings that drive refinement.

  3. Counterfactual Learning (CFL): Driven by an LLM, this component is designed to alleviate sparse supervision by generating counterfactual hard negatives through stochastic attribute perturbation. This technique enforces a causal constraint on supervision, compelling the model to disentangle subtle cues from salient ones and prevent shortcut learning induced by limited data.

The Optimization Objective

The training objective is formulated using Binary Cross Entropy (BCE) and comprises two complementary terms:

  1. The main grounding loss (Lm), computed over all detected objects, minimizes classification error based on the matching probability between the holistic object query Fo and the referring query Fr. This term ensures the model correctly identifies the target while suppressing non-target objects.

  2. The counterfactual loss (Lcf) is applied exclusively to the target object, penalizing its representation when conditioned on a perturbed counterfactual query Fcf r. This enforces attribute disentanglement by teaching the model what the object is not, which compels the model to acquire compositional understanding and prevents feature collapse.

The Impact of External Knowledge Integration

The efficacy of COAL is demonstrated through ablation studies showing that external knowledge priors are crucial for performance gains. Removing Explicit Semantic Injection (ESI) consistently improves performance by providing effective semantic anchors, while introducing Counterfactual Learning (CFL) yields even more significant gains, boosting HOTA to 40.01% on ReferKITTI-V2. Furthermore, the full HMSI architecture is validated as superior to simplified variants; removing the Bi-Fusion module results in a performance drop, confirming that mutual injection is critical for subsequent refinement.

The Results

Experiments on Refer-KITTI and ReferKITTI-V2 benchmarks validate COAL’s efficacy. Notably, it surpasses the state-of-the-art by 7.28% HOTA on the highly challenging ReferKITTI-V2, attaining 43.46% HOTA compared to previous methods. These results demonstrate that knowledge regularization is a highly effective strategy for resolving the sparsity–discriminability paradox in RMOT, proving that representation enrichment and supervision augmentation operate as complementary mechanisms to address structural sparsity constraints.

Visualization of Semantic Manifold

Qualitative analysis using t-SNE embeddings illustrates the mechanism of COAL. Explicit Semantic Injection (ESI) is shown to create highly discriminative semantic anchors from captions, while Counterfactual Learning (CFL) forces these representations to expand and fill the semantic space, preventing degeneration into single-attribute dominated embeddings. This synergy ensures a representation that is both semantically discriminative and structurally comprehensive.

Conclusion

COAL proposes a knowledge-centric alignment framework that incorporates external semantic regularization beyond isolated structural optimization. By integrating ESI and CFL within HMSI, COAL effectively alleviates shortcut learning and improves compositional discrimination under scarce supervision, demonstrating the effectiveness of systematically introducing external semantic knowledge from foundation models for mitigating discrimination bottlenecks in small-scale RMOT tasks.

Acknowledgments

The authors acknowledge support from the National Natural Science Foundation of China (No. 62271143), the Frontier Technologies R&D Program of Jiangsu (No.

Improvements for AI systems

As a fastidious researcher, I have analyzed the COAL: Counterfactual and Observation-Enhanced Alignment Learning for Discriminative Referring Multi-Object Tracking paper. The core innovation lies in resolving the Sparsity-Discriminability Paradox in Referring Multi-Object Tracking (RMOT) by injecting external knowledge to regularize learning against sparse semantic supervision.

Here are the specific, actionable improvements this framework can enable in AI systems:


)1. Enhanced Fine-Grained Object Discrimination under Data Scarcity

The COAL framework significantly improves the system's ability to distinguish between visually similar but semantically distinct objects (e.g., differentiating the red car on the left from the white car on the left).

  • A system can now reliably track and identify objects in highly homogeneous scenes (like a parking lot or a crowded street) even when training data for those specific fine-grained attributes is sparse.

  • The system mitigates shortcut learning, where it incorrectly associates objects based on easily learned, but insufficient, visual cues (like general color or shape), leading to robust compositional understanding.

)2. Robustness Against Semantic Noise and Hallucinations via VLM Priors

By using a Vision-Language Model (VLM like DINO-X) for Explicit Semantic Injection (ESI), the system gains access to dense object proposals and linguistic captions that serve as powerful, albeit noisy, prior knowledge.

  • The system can utilize these priors to densify the observation space, effectively mitigating ambiguity in visual features where implicit cues fail.

  • It can filter out semantic noise from raw VLM outputs (e.g., identifying objects outside the RMOT scope) and leverage accurate captions to refine its understanding of subtle attributes like color or specific locations.

)3. Superior Compositional Reasoning through Counterfactual Supervision

The Counterfactual Learning (CFL) module, driven by a Large Language Model (LLM), acts as an automated hard-negative mining mechanism.

  • The system can be trained to explicitly understand the causal relationship between attributes (e.g., if it's red, it's not blue).

  • This allows the AI to acquire true compositional understanding rather than just memorizing visual correlations, making the tracking robust to novel or unseen attribute combinations in real-world scenarios.

)4. Knowledge-Regularized Learning for Generalization

The unified Hierarchical Multi-Stream Integration (HMSI) architecture ensures that the external knowledge from ESI and CFL is effectively distilled into domain-specific representations.

  • The resulting object representation is simultaneously visually grounded and linguistically explicit, leading to a holistic query vector that captures both spatial geometry and semantic meaning.

  • This regularization prevents the model from collapsing into single-attribute dominated embeddings, ensuring the learned features generalize across different visual contexts, a critical step beyond traditional isolated structural optimization.

)5. Efficient Training via Knowledge Augmentation (1-Image-2N-Queries Strategy)

The proposed training strategy optimizes data utilization by pairing a single image with both positive and counterfactual queries in each iteration.

  • The system can maximize the learning signal from limited datasets by generating hard negatives (counterfactuals), forcing the model to learn sharp decision boundaries between subtle attribute variations with minimal computational overhead per frame.

In summary, this improved AI system transitions from a method that relies solely on internal visual-linguistic fusion to one that employs a sophisticated, knowledge-regularized alignment process. It can perform high-precision tracking and identification in complex, data-sparse environments by leveraging external foundation models to inject semantic priors and enforce compositional logic.

Sources

Related papers