A Compositional Theory of Curvature in Probabilistic Circuits
Hrithik Suresh, Sahil Sidheekh, Shelar Parth Vijay, Yasir Z, Sriraam Natarajan, Narayanan Chatapuram Krishnan
Indian Institute of Technology Palakkad · AT&T · University of Texas at Dallas
cs.LG, cs.AI
Submitted: 2026-08-13
Updated: 2026-08-14
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: This paper introduces a compositional theory of curvature in Probabilistic Circuits (PCs), showing that a sum node's contribution to the Hessian trace of the negative log-likelihood factorizes
Terminology
Summary
This paper introduces a compositional theory of curvature in Probabilistic Circuits (PCs), showing that a sum node's contribution to the Hessian trace of the negative log-likelihood factorizes exactly into two semantically distinct terms: contextual usage (circuit flow) and local curvature. The authors prove this decomposition theoretically and use it to develop an adaptive sharpness-aware regularizer that addresses underfitting issues in global trace regularization.
Probabilistic Circuits are generative models that support exact inference and, unlike deep neural networks, admit an exact and tractable measure of loss-surface curvature: the trace of the Hessian of the log-likelihood. Recent work (Suresh et al. 2026) regularizes this trace globally to bias learning toward flatter, better generalizing optima. The paper shows that treating sharpness as a global regularizer can be misspecified for PCs, whose curvature is inherently compositional.
The authors identify three empirical observations motivating their work:
-
A flatter optimum need not be a better one:
Global trace regularization helps reach flatter optima but can nevertheless attain lower test log-likelihood than the unregularized model when sufficient data are available.
The degradation is not explained by a widening train-test gap, as train log-likelihood also decreases, indicating underfitting. -
Global sharpness is concentrated in a small set of nodes: "A small minority (< 10%) of sum nodes accounts for nearly the entire trace (99.99%), while the contributions of other nodes are negligible."
-
Targeting the largest contributors makes underfitting worse:
The test log-likelihood degrades further as the penalty is restricted to progressively fewer top-(T̂n) nodes.
The paper proves that for any sum node n with strictly positive outgoing weights in a smooth, decomposable PC:
Tn(x) = Fn(x)2 · tn(x)
where:
-
Fn(x) is the circuit flow through node n, measuring
how strongly the node is used in explaining the input
-
tn(x) is a local curvature measure that depends only on node n
Specifically, tn(x) = Σc∈ch(n) (pc(x)/pn(x))2
The decomposition is exact and follows directly from the edge-flow factorization. Both factors are available from the same upward-downward computation used to evaluate circuit flows, so computing the decomposition adds no asymptotic cost.
The paper shows that the local Hessian of a sum node is rank one: "The gradient and Hessian of ln in the edge weights θn are ∇θn ln = −ρn and Hn(x) = ∇2θn ln = ρnρn⊤. Hence, Hn is positive semidefinite and, whenever ρn ≠ 0, has rank one, with unique nonzero eigenvalue λmax(Hn) = ∥ρn∥22 = tn."
Consequently, tn = Tr(Hn) = ∥Hn∥2 = ∥Hn∥F
— the local second-order geometry collapses to a single scalar that is simultaneously the local Hessian trace, maximum ambient curvature, and total Hessian magnitude.
The paper shows that flow is attenuated by routing responsibilities along sum edges: For a tree-structured PC, the node flow is the product of routing responsibilities along the unique root-to-n path.
This leads to a depth bias: "If every sum-edge responsibility on the root-to-n path satisfies re(x) ≤ ρ < 1, then Fn(x) ≤ ρ dΣ(n), and Tn(x) ≤ ρ(2dΣ(n)) tn(x), where dΣ(n) is the number of upstream sum edges."
This formalizes "a mechanism in which the contextual factor can suppress a locally sharp node n reached only through many attenuating routing decisions: its curvature is present, but its global trace contribution is discounted geometrically in the number of upstream sum edges."
The paper derives an exact condition for disagreement between rankings: "For any two sum nodes i, j, Ti > Tj ⇐⇒ ti/tj > (Fj/Fi)2. In particular, a locally less curved node i with ti √(tj/ti)."
This explains why selecting nodes by their global trace contribution can preferentially target highly used circuit components rather than intrinsically sharp mixtures.
The authors propose a gated regularizer that uses local curvature to determine where regularization should act:
Each sum node receives a gate derived from its local curvature: ωn = t̂n / maxm∈S t̂m
The regularizer becomes: Rω(θ) = Σn∈S ωn T̂n
The global trace regularizer is recovered when ωn = 1 for every node.
The gate is monotone increasing, so nodes with larger local curvature receive stronger regularization independently of their flow.
The update preserves the closed-form structure: "For fixed gates ωn n∈S, under the surrogate used by global trace-regularized EM, the update for each outgoing weight of sum node n satisfies λnθnc2 − Nncθnc − µωnNnc = 0, and its positive solution is given by θnc = (Nnc + √(Nnc2 + 4λnµωnNnc)) / (2λn)."
The update thus preserves the form and asymptotic complexity of the global method, changing only the effective node-wise strength, µ → µωn.
-
Global trace contributions concentrate in earlier circuit partitions (Figure 6)
-
Local curvature is substantially less concentrated than global contribution:
the top 10% of nodes account for more than 99.99% of T̂n, but it takes over 60% of nodes to contribute the same for t̂n
-
Concentration in the global trace is primarily caused by the contextual usage than local curvature
-
Selecting by global contribution (T̂n) degrades test log-likelihood on both baudio and ad
-
Selecting by local curvature (t̂n)
preserves the unregularized fit on baudio and improves it slightly on ad
-
Global contribution and local curvature provide materially different signals for allocating regularization
Table 2 shows results across all 20 DEBD datasets:
-
Global trace regularization improves over the unregularized model on only 2 datasets and degrades performance on the remaining 18, indicating underfitting
-
The proposed gated method outperforms global trace regularization and matches or improves upon the unregularized model on majority of the datasets
The paper concludes: "Our analysis showed that a node's global curvature contribution separates exactly into contextual usage and intrinsic local curvature, providing insights into why global trace regularization can flatten the wrong parts of the model and induce underfitting. Guided by this decomposition, we introduced a gated regularizer that preserves tractable learning, while allocating regularization effectively to prevent underfitting."
Future work includes developing richer gating functions that jointly account for local geometry and contextual influence
and applications to curvature-aware model compression, targeted robustness interventions, and adaptive circuit design.
Improvements for AI systems
Improvements to AI Systems:
- Compositional Curvature-Aware Regularization for Tractable Generative Models
-
Implement the gated regularizer (ωn = t̂n / max t̂m) in any Probabilistic Circuit (PC) trainer to replace global trace penalties.
-
The improved system automatically identifies and penalizes only intrinsically sharp sum nodes (high local curvature) while ignoring nodes that are merely heavily used (high flow).
-
Result: Training converges to flatter optima without underfitting, achieving higher test log-likelihood on data-rich regimes where global regularization fails.
- Exact Hessian Trace Decomposition for Model Diagnostics
-
Use Theorem 1 (Tn = Fn2 · tn) to build a diagnostic tool that reports, for every sum node, its contextual usage (flow) and intrinsic curvature separately.
-
The improved system can pinpoint whether a model’s sharpness stems from over-reliance on a few nodes (high flow) or from genuinely ill-conditioned local mixtures (high tn).
-
This enables targeted interventions: prune or re-parameterize nodes with high tn, or rebalance flow via architectural changes, rather than applying uniform smoothing.
- Rank-One Local Geometry Exploitation for Efficient Second-Order Optimization
-
Leverage Proposition 1 (local Hessian is rank-one, with trace = norm2 = Frobenius norm) to compute exact per-node curvature at no extra cost beyond flow evaluation.
-
The improved optimizer can perform per-node adaptive learning rates or preconditioning using tn as a scalar curvature estimate, without storing full Hessians.
-
This yields faster convergence and better-conditioned updates in PCs, especially for deep circuits where curvature varies widely across nodes.
- Flow-Aware Node Selection for Compression and Pruning
-
Use Corollary 1 (depth bias: Tn ≤ ρ(2dΣ) tn) to rank nodes by their potential global impact, not just current contribution.
-
The improved system can prune nodes that are locally sharp but deeply buried (low flow due to attenuating routing), preserving expressiveness while reducing model size.
-
Alternatively, it can identify shallow, high-flow nodes where even small local curvature dominates global loss, enabling targeted robustness fixes.
- Adaptive Gating for Mixed Generative-Discriminative Training
-
Extend the gating mechanism (ωn = t̂n / max t̂m) to any loss that involves PC log-likelihoods (e.g., hybrid classification-generation tasks).
-
The improved system dynamically adjusts regularization strength per node based on its local geometry, preventing over-smoothing of informative mixtures while still flattening pathological ones.
-
This leads to better generalization in semi-supervised or transfer learning settings where PCs are used as feature extractors.
- Real-Time Curvature Monitoring for Online Learning
-
Since both Fn and tn are computed from the same upward-downward pass, the improved system can track per-node curvature in streaming settings with O(1) overhead per update.
-
It can detect when a node’s local curvature spikes (e.g., due to distribution shift) and automatically increase regularization for that node, stabilizing online training without global penalty increases.
-
This enables robust continual learning in PCs for non-stationary data streams.
- Curvature-Guided Architecture Search
-
Use the global-local decomposition to guide PC structure learning (e.g., sum-product network topology search).
-
The improved system can propose new sum nodes in regions of high flow but low curvature (to add capacity where needed) and merge or simplify nodes with high curvature but low flow (to remove instability).
-
This yields more efficient and better-conditioned circuit architectures automatically, reducing manual design effort.
Abstract
Probabilistic Circuits (PCs) are generative models that support exact inference and, unlike deep neural networks, admit an exact and tractable measure of loss-surface curvature: the trace of the Hessian of the log-likelihood. Recent work regularizes this trace globally to bias learning toward flatter, better generalizing optima. We show that treating sharpness as a global regularizer can be misspecified for PCs, whose curvature is inherently compositional. We prove that each sum node's contribution to the Hessian trace factorizes exactly into its circuit flow, which measures how heavily the node is used, and a local sharpness term determined by its output distribution. This decomposition provides insights into why global sharpness regularization is depth biased and can lead to underfitting. Building on it, we introduce an adaptive sharpness aware regularizer that penalizes nodes based on intrinsic local curvature and preserves closed form EM updates. We also show that empirically, this targeted regularization recovers the generalization that global regularization sacrifices while retaining the robustness and benefits of sharpness aware learning.
Sources
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks