Thermodynamics of Learning: A Typed Four-Component Accounting of Memory, Fit, and Value

arXiv:2608.12791 · cond-mat.stat-mech, cs.IT, cs.LG, math.IT · Submitted 2026-08-13 · Read on arXiv

Akihito Sudo

ZeroStruct Inc.

cond-mat.stat-mech, cs.IT, cs.LG, math.IT

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: 36 pages, 3 figures, 5 tables

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

Terminology

Summary

Summary

This paper, Thermodynamics of Learning: A Typed Four-Component Accounting of Memory, Fit, and Value, develops a formal framework for finite-state learning devices that distinguishes between what a device has memorized and what will hold operational value for it on future tasks. The central thesis is that What a finite learning device has recorded and what will hold value for it on future tasks are not the same quantity. The paper establishes a typed accounting that separates four distinct component currencies: acquisition cost (C1), physical dissipation (C2), transport action (C3), and task value (C4). The work is explicitly positioned as exact statements about finite-state learning devices in a fixed operational framework, concerning finite-device value retention under task-distribution shift, and is not a theory of statistical generalization.

The paper's central definition is the capital value of a memory state, V(M; T, b), which is the work gap between an informed protocol class and a blind class obtained by deleting the memory-read port and re-optimizing from scratch. Formally, it is defined as:

"V (M; T, b) = Winformed (M; T, b) − Wblind (T, b): (1) the optimal expected work extractable over the future task distribution T when the memory-read port is available, minus the optimum of a blind agent from whom every read port on M has been deleted and who re-optimizes from scratch under the same tasks, the same access structure, and the same per-run budget b. The paper defines learning as an admissible memory-local update with ∆V > 0," a predicate that is distinct from merely increasing record-correlation (∆I(M; D) > 0) or improving a training-side fit functional.

The results are organized into three main theorem groups:

Main Theorem I (Separation) establishes the basic properties of V, including its well-posedness, non-negativity, affinity in the task distribution T, and invariance under bijective re-labeling of the memory. It also proves a strong separation result: For every n, there is a device family on which record correlation and world correlation grow by n ln 2 while the capital gain is exactly zero. This is demonstrated with the ledger-blocked correlation device LB, where the memory acquires correlation with an environment register Y that is independent of the task's working medium and is not manipulable, thus creating no value. The theorem also states that in the flat∗ regime, data-free updates never increase V.

Main Theorem II (Capitalization ledger) addresses the question of which part of the update account is capitalized into future value. It introduces the search ledger σM = ∆I(M; D) + Σtotal, a memory-side subsystem account. The central result is a bound on the capitalization efficiency ηcap = ∆V/(kT σM). The theorem states that for (F5′)-stable M-local updates under a no-discarded-record-correlation condition (f), the bound ηcap ≤ 1 holds. It provides necessary and sufficient conditions for equality, which include no forgetting, no waste, no Y-contamination, and a reversible implementation. The theorem also maps out regimes where this bound fails, including pure entropy-production denominators, finite-budget gate enablement, and recycling credit. A blank-start cumulative bound survives without (f).

Main Theorem III (Value retention) analyzes the impact of task-distribution shift on value. It defines the retention gap Lgen and retention ratio ρgen. The theorem establishes a two-layer alignment domain: first, an exact exchange rate between value and the side-information-adjusted record fit I(M′; D Y) without any independence assumption; second, an exchange rate with the raw record stock I(M′; D) under a joint side-information neutrality condition (M, D) ⊥ Y. The boundary of this condition is marked by an explicit one-time-pad witness. The theorem also presents a four-coordinate intervention theorem, showing that shift, budget, access, and record content each can reverse or restore the fit–value ranking.

The paper concludes with several corollaries, including a subject-difference identity for the retention gap and the two-way failure of 'overfitting equals low efficiency', demonstrating that ηcap and ρgen admit no functional or monotone relation. It also includes an application to a gate-family, contrasting per-bit linear retention with a threshold transition on a common redraw-depth axis. The paper is explicit about its scope and limitations, stating that it is not a theory of statistical generalization and that all results are exact statements about finite-state devices within a fixed operational framework.

Improvements for AI systems

Improvements to AI Systems Based on This Paper

  1. Value-Aware Memory Management
  • Improvement: Replace correlation-based memory retention (e.g., keeping data with high mutual information with training labels) with a capital-value metric V(M; T, b) that measures expected future work extractable under task-distribution shift.

  • What the improved AI can do: Automatically discard or compress memory states that have high record-correlation but zero capital gain (e.g., ledger-blocked correlation devices), preventing wasted storage and compute on task-irrelevant environmental noise.

  1. Typed Accounting for Learning Updates
  • Improvement: Implement a four-component ledger (acquisition cost C1, dissipation C2, transport action C3, task value C4) for every learning update, rather than a single loss or accuracy scalar.

  • What the improved AI can do: Distinguish between updates that increase training fit but decrease future value (e.g., overfitting to spurious correlations) and updates that genuinely capitalize into operational value, enabling selective gradient updates or replay prioritization based on ΔV > 0 rather than Δloss < 0.

  1. Capitalization Efficiency Bounding
  • Improvement: Use the bound ηcap = ΔV/(kT σM) ≤ 1 to design learning algorithms that enforce no-forgetting, no-waste, and no-Y-contamination conditions during optimization.

  • What the improved AI can do: Detect when an update is inefficient (ηcap < 1) and trigger corrective mechanisms—e.g., reversible weight updates, entropy-regularized objectives, or side-information filtering—to approach the reversible limit, improving sample efficiency and long-term retention under distribution shift.

  1. Task-Shift-Aware Retention Scheduling
  • Improvement: Apply the retention gap Lgen and retention ratio ρgen from Main Theorem III to schedule model fine-tuning or memory consolidation.

  • What the improved AI can do: Predict when a task-distribution shift will cause a drop in value despite stable fit, and proactively re-optimize from scratch (blind re-optimization) or adjust the memory-read port access structure, rather than blindly continuing gradient descent on stale data.

  1. Side-Information Neutrality Enforcement
  • Improvement: Use the joint side-information neutrality condition (M, D) ⊥ Y to regularize representation learning, ensuring that memory content does not become entangled with task-irrelevant environmental variables.

  • What the improved AI can do: Train models that are robust to one-time-pad style confounders—where raw record correlation is high but value is zero—by explicitly penalizing mutual information between memory and non-manipulable environment registers, leading to more transferable and causally grounded representations.

  1. Two-Layer Alignment for Fit–Value Ranking
  • Improvement: Replace single-objective training with a two-layer alignment: first align with side-information-adjusted fit I(M′; D Y), then with raw fit only under neutrality.

  • What the improved AI can do: In domains like continual learning or meta-learning, rank candidate models by their capital value rather than validation loss, and use the four-coordinate intervention theorem (shift, budget, access, record content) to dynamically re-rank when task priors change—preventing catastrophic value loss that pure fit-based selection would miss.

  1. Non-Monotone Efficiency Diagnostics
  • Improvement: Use the proven non-relation between ηcap and ρgen to build diagnostic tools that flag when a model is overfitting in a value sense but not in an efficiency sense, or vice versa.

  • What the improved AI can do: Provide explainable warnings during training—e.g., high retention but low capitalization or high capitalization but poor retention—enabling human operators to intervene with targeted strategies (e.g., budget reallocation, memory port deletion, or record re-sampling) rather than relying on a single overfitting metric.

  1. Finite-Device Budget-Aware Learning
  • Improvement: Incorporate per-run budget b explicitly into the learning objective, as in V(M; T, b), rather than treating compute as an external constraint.

  • What the improved AI can do: Optimize for value under fixed inference or training budgets, automatically trading off memory-read port access, transport action, and dissipation—useful for edge devices or real-time systems where energy and latency are critical, enabling graceful degradation under budget cuts.

Sources

Related papers