Ranking Infrared-Visible Fusion the Way Humans Do: A Learned Pairwise Preference Measure
Sichuan Technology and Business University · Chengdu University of Technology · Sigray, Incorporated
cs.CV, cs.AI
Submitted: 2026-08-02
Updated: 2026-09-16
Comments: 23 pages, 8 figures
Code: https://github.com/HaoranLiu507/LPIFM
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 69/100
The gist: The paper "Ranking Image Fusion the Way Humans Do: A Learned Pairwise Preference Measure for Infrared–Visible Fusion Assessment" addresses the fundamental challenge in infrared–visible image
Terminology
Summary
The paper Ranking Image Fusion the Way Humans Do: A Learned Pairwise Preference Measure for Infrared–Visible Fusion Assessment
addresses the fundamental challenge in infrared–visible image fusion (IVIF) that no ideal fused reference exists,
leading the field to rely on scalar objective metrics that formalize proxies for information transfer, structure, or source similarity.
The authors note that these proxies often disagree with the judgment that ultimately matters: given the same sources, which of two fused results does a human prefer?
While human pairwise comparison is the reference standard,
its cost grows quadratically with the number of algorithms.
To solve this, the authors present the Learned Perceptual Image Fusion Measure (LPIFM), described as a source-conditioned model that operationalizes the human A/B/Tie comparison protocol as a repeatable, scalable surrogate.
LPIFM jointly observes the two sources and two fused candidates and predicts whether A is better, B is better, or the two are perceptually equivalent.
Dataset and Supervision
The model is trained on a new dense preference corpus covering all 6,300 unordered comparisons among 25 fusion methods on the 21 scenes of the VIFB benchmark.
This corpus was labeled under a blinded, randomized, two-stage protocol with expert adjudication.
For training, the authors use 13,125 augmented training records,
which include every non-self comparison... represented in both candidate orders with the label mirrored
plus 525 self-pairs labeled Tie.
Architecture
The LPIFM architecture consists of three primary components:
-
Shared image encoder with context projection: "The infrared and visible sources are concatenated along the channel dimension and mapped by a learned context projection... The same hierarchical encoder is then applied three times: once to this projected context and once each to the fused candidates A and B.
The authors utilize a
pretrained ConvNeXt V2 Base backbone at 384 × 384 input resolution." -
Triadic interaction module: This module
models [human behavior] with bidirectional cross-attention (times N) among the context stream C and the two candidate streams A and B,
allowing each candidate toencode how it preserves or distorts source information relative to both sources and relative to its competitor.
-
Shared preference scorer:
A shared scoring head maps each interacted candidate representation to a scalar preference score, s A and s B,
where the decision is based on the difference = s A - s B.
Training Objective
The authors employ a tie-aware, source-conditioned comparator
with an objective defined on the score difference. The loss function L has three terms
:
-
A preference term L pref that
acts on decisive samples.
-
A margin term L margin that
requires to reach a margin on decisive samples.
-
A tie-band term L tie that
acts only on tie samples and compresses into the band [-epsilon, epsilon].
This approach separates two questions that scalar metrics conflate: which candidate is better, and whether the pair is decidable at all.
Experimental Results
LPIFM demonstrates high performance across various settings:
-
Pairwise Accuracy:
Across scene- and method-generalization settings, LPIFM attains pairwise accuracy of 0.792–0.840.
On full 25-method pools, itexceeds the strongest conventional metric by 0.163–0.211 in accuracy.
-
Ranking Fidelity: The model achieves a
Spearman correlation of 0.941–0.977 with human-derived tie-aware Bradley–Terry rankings.
In the UI/All-M setting, LPIFMreproduced the human top choice (U2Fusion), matched the human ranking to within a mean absolute rank difference of 1.68 positions, and kept every top-10 human method inside its own top 10.
-
Measure-like Consistency: The model's verdicts are
antisymmetric under candidate swap, free of preference cycles, and fully transitive, matching or exceeding the internal consistency of the human panel.
Specifically, LPIFMproduces no cycles: 0 of 10,054 decisive triplets are cyclic,
and achieveshard and weak transitivity of 1.0000.
Cross-Protocol Validation
The authors performed cross-protocol external validation on EVAFusion.
They found that Zero-shot transfer to a corpus annotated under a different preference protocol is weak,
with LPIFM attaining only 0.3656 accuracy. However, three epochs of light fine-tuning... make LPIFM the strongest of 21 assessors on that corpus's held-out split.
This suggests that the instrument learns whichever preference regime it is given.
Improvements for AI systems
1. RLHF-Driven Generative Fusion Models
-
The Improvement: Integrate LPIFM as a learned Reward Model within a Reinforcement Learning from Human Feedback (RLHF) framework to train generative models (such as Diffusion Models or GANs) specifically for image fusion.
-
What the improved system can do: Instead of humans manually ranking results to evaluate algorithms, the generative AI will autonomously iterate and optimize its own fusion parameters. It will learn to synthesize infrared–visible images that maximize human-centric perceptual qualities (like texture preservation and thermal saliency) without requiring a ground-truth reference image.
2. Meta-Learning for Rapid Protocol Adaptation
-
The Improvement: Implement a meta-learning architecture (e.g., Model-Agnostic Meta-Learning) or visual prompting techniques to address the
weak zero-shot transfer
identified in the paper. -
What the improved system can do: The system will become a
Universal Perceptual Evaluator.
It can instantly adapt to new, specialized preference regimes—such as a military protocol prioritizing thermal detection or a medical protocol prioritizing anatomical structure—using only a handful of labeled examples (few-shot) rather than requiring extensive fine-tuning.
3. Explainable Perceptual Diagnostics (XAI)
-
The Improvement: Leverage the Triadic Interaction module’s cross-attention maps to generate spatial saliency maps that explain the
why
behind a preference score. -
What the improved system can do: When the system prefers Candidate A over Candidate B, it will provide a visual heat map indicating exactly which regions or features (e.g.,
loss of visible edge detail in the bottom-left
orthermal washout in the center
) caused the lower score. This transforms the metric from ablack box
into a diagnostic tool for researchers to debug fusion algorithms.
4. Uncertainty-Guided Active Learning for Dataset Construction
-
The Improvement: Use the LPIFM’s predicted (score difference) and tie-band probability to drive an active learning loop for human annotators.
-
What the improved system can do: Instead of performing expensive, exhaustive quadratic comparisons, the system will automatically identify
informative
pairs—those where the model is most uncertain or where the human preference is likely to be most contested. This will drastically reduce the human labor cost required to build high-quality perceptual datasets for any multi-modal task.
5. Multi-Source/Multi-Candidate Scalable Evaluator
-
The Improvement: Generalize the Triadic Interaction module from a pairwise (A vs. B) architecture to a multi-stream architecture capable of handling N sources and M candidates.
-
What the improved system can do: The system will be able to perform
Global Ranking
andMulti-Source Fusion Assessment
(e.g., fusing RGB, Thermal, and Depth simultaneously). It can evaluate a pool of many fusion results at once and provide a single, consistent, and transitive ranking of all candidates relative to multiple input modalities.
Abstract
Infrared-visible image fusion (IVIF) has no ideal fused reference, so algorithms are ranked by scalar objective metrics that formalize proxies for information transfer, structure, or source similarity. These proxies often disagree with the judgment that ultimately matters: given the same sources, which of two fused results does a human prefer? Direct pairwise comparison is an established protocol for relative subjective assessment, but its cost grows quadratically with the number of algorithms. We present the Learned Perceptual Image Fusion Measure (LPIFM), a source-conditioned model that operationalizes the human A/B/Tie comparison protocol as a repeatable, scalable surrogate. LPIFM jointly observes the two sources and two fused candidates and predicts whether A is better, B is better, or the two are perceptually equivalent. Supervision comes from a new dense preference corpus covering all 6,300 unordered comparisons among 25 fusion methods on the 21 scenes of the VIFB benchmark, labeled under a blinded, randomized, two-stage protocol with expert adjudication. Across scene- and method-generalization settings, LPIFM attains pairwise accuracy of 0.792-0.840 and Spearman correlation of 0.941-0.977 with human-derived tie-aware Bradley-Terry rankings; on full 25-method pools it exceeds the strongest conventional metric by 0.163-0.211 in accuracy. Its verdicts are also antisymmetric under candidate swap, free of preference cycles, and fully transitive, matching or exceeding the internal consistency of the human panel. We publicly release the dataset, model weights, and code. LPIFM offers a practical instrument for human-aligned comparison and ranking of IVIF methods at scale.
Sources
- Rethinking the Evaluation of Visible and Infrared Image Fusion
- Bringing Multimodal Large Language Models to Infrared-Visible Image Fusion Quality Assessment
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models