Proportional Analogies on Probability Distributions via Bayesian Updating

arXiv:2608.11724 · cs.AI · Submitted 2026-08-12 · Read on arXiv

Pierre-Alexandre Murena

Hamburg University of Technology

cs.AI

Submitted: 2026-08-12

Updated: 2026-08-13

Code: https://github.com/ppaamm/Analogies-onProbability-Distributions-by-Bayes-Updating

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: The paper introduces a new notion of proportional analogy for probability distributions based on Bayesian updating.

Terminology

Summary

The paper introduces a new notion of proportional analogy for probability distributions based on Bayesian updating. The central idea is to characterize the similarity between two distributions through the existence of a dataset capable of transforming one distribution into the other by Bayesian inference with a given likelihood. The authors state: The central idea is to characterize the similarity between two distributions through the existence of a dataset capable of transforming one distribution into the other by Bayesian inference with a given likelihood.

The proposed framework satisfies the classical properties of proportional analogies (reflexivity, symmetry, central permutation), as shown in Proposition 1: If  and  are such that for all p, q ∈ , there exists x ∈  such that p ↭ q, then the relation R defined above is a proportional analogy, and we call it Bayes-based proportional analogy for likelihood .

For distributions in the exponential family, the analogy admits a simple arithmetic characterization in the natural parameter space. Proposition 4 states: "For all hyperparameters xiA, xiB, xiC and xiD, if p(. ∣ xiA) ∶ p(. ∣ xiB) ∶∶ p(. ∣ xiC) ∶ p(. ∣ xiD) then either (1) the natural parameters eta(xi) satisfy the arithmetic analogy; or (2) the reverse solution holds, i.e. eta(xiA) = eta(xiD) and eta(xiB) = eta(xiC)."

Theorem 1 provides a full characterization: "For any xiA, xiB, xiC, xiD ∈ Ξ, we have p(. ∣ xiA) ∶ p(. ∣ xiB) ∶∶ p(. ∣ xiC) ∶ p(. ∣ xiD) if and only if eta(xiA) − eta(xiB), eta(xiB) − eta(xiA) ∩ Im(S) and eta(xiC) − eta(xiA), eta(xiA) − eta(xiC) ∩ Im(S) are not empty, and either of the two following conditions is satisfied: 1. eta(xiA) = eta(xiD), eta(xiB) = eta(xiC); or 2. eta(xiB) − eta(xiA) = eta(xiD) − eta(xiC)."

The paper also establishes a divergence-based characterization. Proposition 7 states: Denote M(p, q) = arg minr (DKL(p‖r) + DKL(q‖r)). If four distributions pA, pB, pC and pD of the exponential family are in analogy, then either M(pA, pD) = M(pB, pC) or M(pA, pC) = M(pB, pD).

The framework is instantiated for several classical exponential families: categorical/Bernoulli distributions, multivariate normal distributions, and Dirichlet distributions. For categorical distributions, the authors show there exists a maximal analogy (Proposition 8): There exists a maximal Bayes-based proportional analogy between categorical distributions induced by the likelihood KMM. For multivariate normal distributions, Proposition 9 states: There exists a unique maximal Bayes-based proportional analogy relation induced by likelihoods of the form quadA,B, obtained when Im(S) = Rd × vec(Sd++). For Dirichlet distributions, Proposition 11 states: There exists a maximal Bayes-based proportional analogy between Dirichlet distributions induced by likelihood S+ that satisfy Proposition 10.

The paper also develops a sampling-based algorithm for solving analogical equations between arbitrary probability distributions, making the framework applicable beyond analytically tractable models. The algorithm uses importance sampling for posterior approximation, distribution comparison via discrepancy measures (such as Wasserstein distance), and joint optimization to find latent datasets. The authors note: The objective is to find (x, x′) minimizing d(thetaBi, thetâBi) + d(thetaCi, thetâCi) + d(thetâ(B),iD, thetâ(C),iD).

Empirical validation was conducted on 3121 synthetic analogical equations with known analytical solutions. The solver successfully recovered solutions for 1918 out of 3121 equations (61.5%), with a median Gaussian Wasserstein distance of 0.132 and a 90th percentile of 0.421. The authors report: The proposed solver successfully recovers solutions for 1918 out of the 3121 equations, i.e. 61.5% of the generated analogies. They also note that the error is mostly due to the variance that tends to be under-estimated while there is a strong correlation in the means, with a Pearson correlation of 0.997.

The paper concludes that this work provides a principled foundation for connecting analogical reasoning with machine learning and opens perspectives for applications such as transfer learning, data augmentation, distribution reconstruction, case-based reasoning, and meta-learning.

Improvements for AI systems

Based on the paper, here are specific improvements I can make to AI systems:

  1. Bayesian Analogy Engine for Few-Shot Learning: I can implement a module that, given a pair of known distributions (e.g., class conditional densities from source tasks), automatically constructs a latent dataset that transforms one into the other via Bayesian updating. This enables the AI to generate new target distributions for unseen classes by solving analogical equations, improving few-shot classification without retraining.

  2. Distributional Transfer via Exponential-Family Arithmetic: For models using exponential family outputs (e.g., Gaussian policies in RL, categorical softmax layers), I can enforce the arithmetic analogy in natural parameter space (Proposition 4). This allows the AI to interpolate or extrapolate between learned parameter vectors (e.g., from task A to task B) to synthesize parameters for a new task C, enabling zero-shot policy or classifier transfer.

  3. Divergence-Based Analogy Checker: I can add a validation layer that uses Proposition 7 (the KL-divergence midpoint condition) to verify whether a proposed analogy between four distributions is valid before using it for data augmentation or domain adaptation. This reduces hallucinated analogies and improves robustness in generative models.

  4. Sampling-Based Analogical Equation Solver: I can integrate the proposed importance-sampling and Wasserstein-distance optimization algorithm to solve analogical equations for arbitrary, non-analytic distributions (e.g., from deep generative models). This allows the AI to find latent datasets that map between empirical distributions, enabling distributional editing (e.g., shifting a dataset’s style or attributes) without explicit likelihood models.

  5. Maximal Analogy Construction for Categorical and Gaussian Models: For classification tasks, I can precompute maximal Bayes-based analogy relations (Propositions 8 and 9) to define a canonical set of transformations between class-conditional distributions. The AI can then use these to generate synthetic training examples by analogy (e.g., class A is to class B as class C is to class D), improving data augmentation for imbalanced or rare classes.

  6. Meta-Learning via Analogy-Based Prior Initialization: I can use the framework to define a meta-learning objective where the AI learns to find latent datasets that transform a base distribution into task-specific posteriors. This yields a prior over model parameters that satisfies proportional analogies across tasks, enabling faster adaptation in meta-learning (e.g., MAML-like algorithms) with provable analogy constraints.

What the improved AI system can do specifically:

  • Given three labeled distributions (e.g., from three classes), it can generate a fourth distribution that completes the proportional analogy, even for complex, non-Gaussian data.

  • It can verify if a proposed analogy is mathematically sound, avoiding spurious transfers in domain adaptation.

  • It can perform zero-shot parameter synthesis for new tasks by applying arithmetic analogy in natural parameter space, without gradient updates.

  • It can solve analogical equations between empirical datasets (e.g., image style transfer) by finding latent datasets that act as Bayesian bridges, using only samples and a discrepancy metric.

  • It can construct maximal analogy relations for categorical or Gaussian models to systematically augment training data, improving generalization in low-resource settings.

Abstract

Analogies are quaternary relations of the form "A is to B as C is to D". Among the various formalizations of analogical reasoning, proportional analogies provide an important axiomatic framework by characterizing valid analogies through a set of postulates. While proportional analogies have been extensively studied over Boolean, symbolic, and real-valued domains, their extension to probability distributions remains largely unexplored. In this paper, we introduce a notion of proportional analogy for probability distributions based on Bayesian updating. Our approach builds upon the idea that two distributions are related whenever one can be transformed into the other through Bayesian updating induced by a suitable set of observations. We investigate this framework for several standard members of the exponential family and discuss how it naturally extends to arbitrary probability distributions through Gaussian mixture approximations.

Sources

Related papers