PAC-Bayes Beyond Parameter Space: Behavioral Equivalence, Z-Information, and Exact Complexity Decomposition

arXiv:2608.11465 · cs.LG, cs.AI · Submitted 2026-08-11 · Read on arXiv

Vasant G. Honavar, Satish Kumar Keshri, Neil Ashtekar, Zehao Liu

The Pennsylvania State University

cs.LG, cs.AI

Submitted: 2026-08-11

Updated: 2026-08-13

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 100/100

The gist: PAC-Bayes theory provides generalization guarantees by controlling the Kullback–Leibler (KL) divergence between posterior and prior distributions over a chosen hypothesis representation.

Terminology

Summary

PAC-Bayes theory provides generalization guarantees by controlling the Kullback–Leibler (KL) divergence between posterior and prior distributions over a chosen hypothesis representation. However, predictive risk depends only on the predictive behavior induced by a hypothesis, not on the particular internal realization that implements that behavior. In modern over-parameterized learning systems, many distinct configurations may induce identical predictive behavior, yet the classical PAC-Bayes KL divergence does not distinguish uncertainty over predictive behavior from variation among behaviorally equivalent realizations. We show that this distinction induces an exact structural decomposition of classical PAC-Bayes complexity. We formalize behavioral equivalence through a measurable behavior map and use measure disintegration to decompose probability measures on the configuration space into a distribution over predictive behaviors together with conditional distributions over behavioral fibers. This yields an exact decomposition of the classical PAC-Bayes KL divergence into a behavior-selection term and a realization-level term given by an expected conditional KL divergence within behavioral fibers.

We define PAC-Bayes Z-information as the negative of this realization-level contribution. Consequently, PAC-Bayes Z-information exactly quantifies the gap between the classical PAC-Bayes KL divergence and the irreducible complexity associated with uncertainty over predictive behavior. We further show that the behavior-selection term admits an exact variational characterization: it is the minimum classical PAC-Bayes KL divergence among all posteriors that induce the same distribution over predictive behaviors. Equivalently, every posterior admits a canonical fiber-symmetrized representative with identical predictive behavior and minimum classical PAC-Bayes complexity.

Finally, we show that symmetry, behavior-preserving directions, fiber geometry, and invariance under fiber-preserving perturbations arise naturally from the same behavior-map structure. Together, these results identify predictive behavior as the natural object of PAC-Bayes complexity, provide a unified measure-theoretic and geometric characterization of realization multiplicity, and reveal an exact structural decomposition that is implicit in classical PAC-Bayes theory.

Improvements for AI systems

Improvements to AI Systems:

  1. Behavior-Aware Generalization Bounds: Replace classical PAC-Bayes KL divergence with the new behavior-selection term (minimum KL over behaviorally equivalent posteriors). This yields tighter, more meaningful generalization guarantees for over-parameterized models (e.g., deep networks, transformers) by ignoring redundant parameterization noise.

  2. Canonical Fiber-Symmetrized Training: Add a post-training or regularization step that projects any learned posterior onto its canonical fiber-symmetrized representative. This reduces model complexity without changing predictive behavior, improving test-time robustness and compression (e.g., smaller effective model size for deployment).

  3. Uncertainty Calibration via Z-Information: Use PAC-Bayes Z-information as a new metric to quantify and minimize the gap between predictive uncertainty and parameter uncertainty. This enables better calibration of Bayesian neural networks, especially in safety-critical applications where overconfident or underconfident predictions are costly.

  4. Symmetry-Aware Optimization: Leverage the identified behavior-preserving directions and fiber geometry to design gradient updates that avoid redundant parameter changes. This accelerates convergence and reduces overfitting in models with known symmetries (e.g., permutation invariance in graph neural networks, weight-space symmetries in ReLU networks).

  5. Invariant Regularization for Perturbation Robustness: Use the invariance under fiber-preserving perturbations to construct adversarial training or data augmentation schemes that specifically target behavior-irrelevant parameter noise. This improves robustness to small parameter perturbations (e.g., weight quantization, pruning, or hardware noise) without sacrificing predictive accuracy.

  6. Efficient Model Selection and Ensembling: Apply the variational characterization to select among posteriors with identical predictive behavior the one with minimal complexity. This enables more efficient Bayesian model averaging and ensembling, reducing computational cost while preserving predictive diversity.

  7. Interpretable Complexity Decomposition: Provide practitioners with a breakdown of total PAC-Bayes complexity into behavior-selection and realization-level components. This allows debugging of over-parameterized systems—identifying whether high complexity stems from genuine predictive uncertainty or from redundant parameterization, guiding architectural choices.

What the Improved AI System Can Do:

  • Achieve tighter generalization guarantees and better test performance on over-parameterized models with fewer labeled samples.

  • Automatically simplify trained models (e.g., compress weights) without loss in accuracy, enabling deployment on edge devices.

  • Produce well-calibrated uncertainty estimates that reflect only predictive ambiguity, not parameterization artifacts.

  • Train faster and more stably by ignoring redundant parameter directions, especially in large-scale neural networks.

  • Maintain performance under weight perturbations, quantization, or pruning, making models more reliable in real-world noisy environments.

  • Perform principled model selection and Bayesian inference at lower computational cost.

  • Provide clear diagnostic reports on where model complexity originates, aiding in architecture design and hyperparameter tuning.

Sources

Related papers