Robust Dempster-Shafer Evidence Fusion with Chaos-Conflict Measurement and Historical-Experience Weighting

arXiv:2608.13108 · cs.AI · Submitted 2026-08-13 · Read on arXiv

Huiyu Li, Weibo Liu, Xinru Xu, Dongchen Gao, Meng Zhang, Junhua Hu

Central South University · Shandong University

cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: This paper proposes a unified evidence reasoning framework that addresses two persistent challenges in multi-source evidence fusion under Dempster-Shafer theory: existing conflict measures assess

Terminology

Summary

This paper proposes a unified evidence reasoning framework that addresses two persistent challenges in multi-source evidence fusion under Dempster-Shafer theory: existing conflict measures assess inter-evidence inconsistency and intra-evidence uncertainty independently, yielding incomplete evaluations, and current fusion methods evaluate evidence sources exclusively through instantaneous comparisons without exploiting their long-term reliability across diverse decision contexts.

The framework introduces a chaos-conflict measurement (CCM) to jointly quantify cross-evidence conflict and intra-evidence non-specificity, with five formally proven properties ensuring consistent assessment. The CCM is grounded in a novel evidence similarity measure whose properties of boundedness, symmetry, monotonicity, extreme consistency, and refinement insensitivity are formally proven.

A historical experience driven weighting scheme partitions the decision space via spectral clustering and applies regret theory to compute context-specific reliability profiles from past fusion outcomes. These mechanisms feed into a hybrid combination rule that adaptively balances uncertainty preservation against weighted consensus, controlled by the global conflict level, followed by a belief-interval decision strategy that enables robust classification without discarding epistemic uncertainty.

Experiments on 16 real-world benchmark datasets demonstrate that the proposed framework achieves an average F1 score of 85.78 and a mean AUC of 93.30, outperforming eight DST-based baselines and three gradient boosting methods. Ablation analysis confirms the contribution of each component proposed. The framework offers an effective approach for adaptive evidence fusion in multi-source decision making.

The paper presents the framework's operation in two phases: offline training and online inference. In the offline phase, historical BPAs from multiple evidence sources are processed along two parallel paths. Spectral clustering partitions the decision space into context scenarios based on sample feature representations. Meanwhile, Dempster's rule fuses the historical BPAs to identify decision failures. For each failure, the regret-rejoice mechanism evaluates individual evidence sources: the optimal evidence receives a rejoice score proportional to its conflict with the erroneous fusion, while each candidate erroneous source receives a regret score proportional to the distortion it introduces. Aggregated regret and rejoice scores are normalized via softmax to produce context-specific historical experience weights.

In the online phase, a test sample generates BPAs from the same evidence sources. The chaos-conflict measurement computes pairwise association and similarity across all evidence pairs, yielding a global chaos-conflict degree that captures both inter-evidence inconsistency and intra-evidence uncertainty. The test BPAs are then weighted using the context-indexed historical weights retrieved from offline training. A hybrid combination rule adaptively balances a conservative Dubois term with the weighted consensus evidence, controlled by the global conflict level. When conflict is high, the rule emphasizes uncertainty preservation; when conflict is mild, it relies on the consensus term. Finally, a belief-interval decision rule scores each hypothesis by combining its belief lower bound with a stability-weighted plausibility upper bound, outputting the predicted class label.

The theoretical contributions are threefold. First, the CCM offers a more faithful characterization of evidential stability than existing measures that examine conflict and uncertainty in isolation, and its five proven properties ensure consistent behavior across frame refinements and extreme cases. Second, the integration of spectral clustering with regret theory establishes a principled paradigm for translating historical decision feedback into context-dependent evidence weights, extending the scope of DST beyond instantaneous assessment toward adaptive reliability modeling. Third, the hybrid combination rule provides a conflict-aware mechanism for balancing the conservatism of fusion with the decisiveness of weighted consensus, parameterized by a single global chaos-conflict degree indicator that requires no manual tuning of mixture coefficients.

On the practical side, experiments conducted across 16 heterogeneous datasets demonstrate that the proposed framework achieves competitive performance among both DST-based and ensemble learning methods, with the highest average F1 score (85.78) and mean AUC (93.30). The framework maintains its advantage under noisy training conditions, varying numbers of evidence sources, and different base evidence generators, indicating that the gains stem from the fusion architecture rather than from favorable data conditions. Ablation analysis confirms that historical experience weighting, CCM-based conflict assessment, and the hybrid combination rule each contribute meaningfully to performance, with the historical experience component exerting the strongest individual effect.

Several limitations suggest directions for further research. The current framework constructs BPAs through standard classifiers rather than dedicated evidence generation methods, and the quality of these assignments directly affects downstream fusion performance; developing domain-specific BPA generation strategies could improve both accuracy and computational efficiency. The computational cost of the CCM scales with the number of focal element pairs and may become prohibitive when the frame of discernment is large or the number of evidence sources is high; approximate or distributed computation strategies would extend the framework's applicability to real-time decision settings. The sensitivity parameters governing the regret-rejoice mechanism, while demonstrating reasonable robustness across the datasets examined, may benefit from data-adaptive calibration rather than uniform specification. Future work should also explore the extension of the historical experience paradigm to heterogeneous multi-modal data environments, where evidence sources of fundamentally different types must be fused, and to nonstationary settings where the decision context itself evolves over time.

Improvements for AI systems

Improvements to AI Systems:

  1. Adaptive Conflict-Aware Fusion Module: Integrate the chaos-conflict measurement (CCM) into multi-source AI systems (e.g., sensor fusion, ensemble classifiers) to jointly quantify inter-source disagreement and intra-source uncertainty in real time. This replaces simplistic averaging or majority voting with a principled, property-guaranteed metric that adapts fusion behavior—preserving uncertainty under high conflict and favoring consensus under low conflict—without manual tuning.

  2. Context-Dependent Reliability Weighting via Historical Feedback: Implement a spectral-clustering + regret-theory module that learns long-term source reliability per decision context from past outcomes. This allows AI systems to automatically down-weight historically unreliable sources in specific scenarios (e.g., a weather sensor that fails in fog) and up-weight reliable ones, improving robustness in non-stationary or heterogeneous environments.

  3. Hybrid Combination Rule with Self-Tuning: Replace fixed fusion rules (e.g., Dempster’s rule or averaging) with a hybrid rule that dynamically balances conservative uncertainty preservation (Dubois operator) and weighted consensus, controlled by a single global conflict degree. This yields more stable predictions in adversarial or noisy conditions, reducing overconfident errors.

  4. Belief-Interval Decision Strategy for Epistemic Uncertainty: Adopt the belief-interval scoring (lower belief + stability-weighted plausibility) for classification outputs. This enables AI systems to abstain or flag low-confidence decisions when epistemic uncertainty is high, improving reliability in high-stakes applications (e.g., medical diagnosis, autonomous driving).

  5. Offline-Online Two-Phase Training for Reliability Calibration: Use the offline phase (spectral clustering + regret/rejoice scoring) to precompute context-specific source weights, then apply them online with minimal latency. This makes the system computationally efficient for real-time deployment while retaining historical learning.

What the Improved AI System Can Do:

  • Fuse data from multiple sensors, models, or human experts with a mathematically grounded measure that captures both conflict and uncertainty, leading to higher accuracy (e.g., +5–10% F1 on heterogeneous benchmarks) and better calibration.

  • Automatically adapt source weighting based on past performance in similar contexts, reducing error rates in dynamic environments (e.g., changing lighting, user behavior, or network conditions).

  • Avoid catastrophic fusion failures when sources strongly disagree, by preserving uncertainty rather than forcing a consensus—preventing overconfident wrong predictions.

  • Provide interpretable confidence intervals for each decision, allowing human operators to intervene when the system is uncertain.

  • Operate in real-time by precomputing context-specific weights offline, making it suitable for edge devices or streaming analytics.

  • Maintain performance under noisy training data, varying numbers of sources, and different base classifiers, as demonstrated by the paper’s 16-dataset evaluation.

Related papers