Ranking Infrared-Visible Fusion the Way Humans Do: A Learned Pairwise Preference Measure

arXiv:2608.01301 · cs.CV, cs.AI · Submitted 2026-08-02 · Read on arXiv

Sichuan Technology and Business University · Chengdu University of Technology · Sigray, Incorporated

cs.CV, cs.AI

Submitted: 2026-08-02

Updated: 2026-09-16

Comments: 23 pages, 8 figures

Code: https://github.com/HaoranLiu507/LPIFM

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 69/100

The gist: The paper "Ranking Image Fusion the Way Humans Do: A Learned Pairwise Preference Measure for Infrared–Visible Fusion Assessment" addresses the fundamental challenge in infrared–visible image

Terminology

Summary

The paper Ranking Image Fusion the Way Humans Do: A Learned Pairwise Preference Measure for Infrared–Visible Fusion Assessment addresses the fundamental challenge in infrared–visible image fusion (IVIF) that no ideal fused reference exists, leading the field to rely on scalar objective metrics that formalize proxies for information transfer, structure, or source similarity. The authors note that these proxies often disagree with the judgment that ultimately matters: given the same sources, which of two fused results does a human prefer? While human pairwise comparison is the reference standard, its cost grows quadratically with the number of algorithms.

To solve this, the authors present the Learned Perceptual Image Fusion Measure (LPIFM), described as a source-conditioned model that operationalizes the human A/B/Tie comparison protocol as a repeatable, scalable surrogate. LPIFM jointly observes the two sources and two fused candidates and predicts whether A is better, B is better, or the two are perceptually equivalent.

Dataset and Supervision

The model is trained on a new dense preference corpus covering all 6,300 unordered comparisons among 25 fusion methods on the 21 scenes of the VIFB benchmark. This corpus was labeled under a blinded, randomized, two-stage protocol with expert adjudication. For training, the authors use 13,125 augmented training records, which include every non-self comparison... represented in both candidate orders with the label mirrored plus 525 self-pairs labeled Tie.

Architecture

The LPIFM architecture consists of three primary components:

  1. Shared image encoder with context projection: "The infrared and visible sources are concatenated along the channel dimension and mapped by a learned context projection... The same hierarchical encoder is then applied three times: once to this projected context and once each to the fused candidates A and B. The authors utilize a pretrained ConvNeXt V2 Base backbone at 384 × 384 input resolution."

  2. Triadic interaction module: This module models [human behavior] with bidirectional cross-attention (times N) among the context stream C and the two candidate streams A and B, allowing each candidate to encode how it preserves or distorts source information relative to both sources and relative to its competitor.

  3. Shared preference scorer: A shared scoring head maps each interacted candidate representation to a scalar preference score, s A and s B, where the decision is based on the difference = s A - s B.

Training Objective

The authors employ a tie-aware, source-conditioned comparator with an objective defined on the score difference. The loss function L has three terms:

  • A preference term L pref that acts on decisive samples.

  • A margin term L margin that requires to reach a margin on decisive samples.

  • A tie-band term L tie that acts only on tie samples and compresses into the band [-epsilon, epsilon].

This approach separates two questions that scalar metrics conflate: which candidate is better, and whether the pair is decidable at all.

Experimental Results

LPIFM demonstrates high performance across various settings:

  • Pairwise Accuracy: Across scene- and method-generalization settings, LPIFM attains pairwise accuracy of 0.792–0.840. On full 25-method pools, it exceeds the strongest conventional metric by 0.163–0.211 in accuracy.

  • Ranking Fidelity: The model achieves a Spearman correlation of 0.941–0.977 with human-derived tie-aware Bradley–Terry rankings. In the UI/All-M setting, LPIFM reproduced the human top choice (U2Fusion), matched the human ranking to within a mean absolute rank difference of 1.68 positions, and kept every top-10 human method inside its own top 10.

  • Measure-like Consistency: The model's verdicts are antisymmetric under candidate swap, free of preference cycles, and fully transitive, matching or exceeding the internal consistency of the human panel. Specifically, LPIFM produces no cycles: 0 of 10,054 decisive triplets are cyclic, and achieves hard and weak transitivity of 1.0000.

Cross-Protocol Validation

The authors performed cross-protocol external validation on EVAFusion. They found that Zero-shot transfer to a corpus annotated under a different preference protocol is weak, with LPIFM attaining only 0.3656 accuracy. However, three epochs of light fine-tuning... make LPIFM the strongest of 21 assessors on that corpus's held-out split. This suggests that the instrument learns whichever preference regime it is given.

Improvements for AI systems

1. RLHF-Driven Generative Fusion Models

  • The Improvement: Integrate LPIFM as a learned Reward Model within a Reinforcement Learning from Human Feedback (RLHF) framework to train generative models (such as Diffusion Models or GANs) specifically for image fusion.

  • What the improved system can do: Instead of humans manually ranking results to evaluate algorithms, the generative AI will autonomously iterate and optimize its own fusion parameters. It will learn to synthesize infrared–visible images that maximize human-centric perceptual qualities (like texture preservation and thermal saliency) without requiring a ground-truth reference image.

2. Meta-Learning for Rapid Protocol Adaptation

  • The Improvement: Implement a meta-learning architecture (e.g., Model-Agnostic Meta-Learning) or visual prompting techniques to address the weak zero-shot transfer identified in the paper.

  • What the improved system can do: The system will become a Universal Perceptual Evaluator. It can instantly adapt to new, specialized preference regimes—such as a military protocol prioritizing thermal detection or a medical protocol prioritizing anatomical structure—using only a handful of labeled examples (few-shot) rather than requiring extensive fine-tuning.

3. Explainable Perceptual Diagnostics (XAI)

  • The Improvement: Leverage the Triadic Interaction module’s cross-attention maps to generate spatial saliency maps that explain the why behind a preference score.

  • What the improved system can do: When the system prefers Candidate A over Candidate B, it will provide a visual heat map indicating exactly which regions or features (e.g., loss of visible edge detail in the bottom-left or thermal washout in the center) caused the lower score. This transforms the metric from a black box into a diagnostic tool for researchers to debug fusion algorithms.

4. Uncertainty-Guided Active Learning for Dataset Construction

  • The Improvement: Use the LPIFM’s predicted (score difference) and tie-band probability to drive an active learning loop for human annotators.

  • What the improved system can do: Instead of performing expensive, exhaustive quadratic comparisons, the system will automatically identify informative pairs—those where the model is most uncertain or where the human preference is likely to be most contested. This will drastically reduce the human labor cost required to build high-quality perceptual datasets for any multi-modal task.

5. Multi-Source/Multi-Candidate Scalable Evaluator

  • The Improvement: Generalize the Triadic Interaction module from a pairwise (A vs. B) architecture to a multi-stream architecture capable of handling N sources and M candidates.

  • What the improved system can do: The system will be able to perform Global Ranking and Multi-Source Fusion Assessment (e.g., fusing RGB, Thermal, and Depth simultaneously). It can evaluate a pool of many fusion results at once and provide a single, consistent, and transitive ranking of all candidates relative to multiple input modalities.

Abstract

Infrared-visible image fusion (IVIF) has no ideal fused reference, so algorithms are ranked by scalar objective metrics that formalize proxies for information transfer, structure, or source similarity. These proxies often disagree with the judgment that ultimately matters: given the same sources, which of two fused results does a human prefer? Direct pairwise comparison is an established protocol for relative subjective assessment, but its cost grows quadratically with the number of algorithms. We present the Learned Perceptual Image Fusion Measure (LPIFM), a source-conditioned model that operationalizes the human A/B/Tie comparison protocol as a repeatable, scalable surrogate. LPIFM jointly observes the two sources and two fused candidates and predicts whether A is better, B is better, or the two are perceptually equivalent. Supervision comes from a new dense preference corpus covering all 6,300 unordered comparisons among 25 fusion methods on the 21 scenes of the VIFB benchmark, labeled under a blinded, randomized, two-stage protocol with expert adjudication. Across scene- and method-generalization settings, LPIFM attains pairwise accuracy of 0.792-0.840 and Spearman correlation of 0.941-0.977 with human-derived tie-aware Bradley-Terry rankings; on full 25-method pools it exceeds the strongest conventional metric by 0.163-0.211 in accuracy. Its verdicts are also antisymmetric under candidate swap, free of preference cycles, and fully transitive, matching or exceeding the internal consistency of the human panel. We publicly release the dataset, model weights, and code. LPIFM offers a practical instrument for human-aligned comparison and ranking of IVIF methods at scale.

Sources

Related papers