Why Does the Future Branch? Identifiable Closure Tests for Stochastic Physical World Models

arXiv:2608.00591 · cs.AI · Submitted 2026-08-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Why Does the Future Branch? Identifiable Closure Tests for Stochastic Physical World Models".

Jane: The paper was written by Yibin Dong from Shandong University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Alright, welcome back to the channel. Today we're digging into a paper that's been making the rounds, and the title alone got me hooked: "Why Does the Future Branch? Identifiable Closure Tests for Stochastic Physical World Models."

Jane: And Tom, that title is actually doing a lot of work. "Why does the future branch" — that's the question. When a world model gives you a distribution over possible futures, is it because we can't see something about the current state, or is it because the world itself is genuinely random from here on out?

Tom: Right, and the paper says those two things look identical if you only watch the predictions. You can have two completely different physical setups that produce the exact same forecast distribution.

Jane: Exactly. They call one "state aliasing" — you're missing information, like you can't see the velocity of a ball. And the other is "process stochasticity" — you see everything, but there's fresh noise pushing the future around.

Tom: So the future branches for two totally different reasons, but the forecast looks the same. That's the core puzzle.

Jane: And the authors, Yibin Dong from Shandong University, show that no amount of watching ordinary predictions can tell those apart. You need to intervene — reset the state, replay the noise — to figure out which one is actually driving the uncertainty.

Tom: It's like trying to figure out whether a magician is hiding a card up their sleeve or actually shuffling a fresh deck every time. The trick looks identical from the audience.

Jane: That's the intuition. And the paper's whole point is that we can design tests to peek behind the curtain without changing the forecast quality at all.

Tom: So we're not talking about a better predictor here. We're talking about understanding the mechanism underneath.

Jane: And that matters for decisions. If the uncertainty comes from missing state, you should spend compute on better sensors. If it comes from process noise, you should spend compute on more branches.

Tom: So the title is really asking a practical question. Why does the future branch? Because the answer tells you where to point your resources.

Jane: And we're going to spend the whole episode unpacking how they answer it. Let's get into the paper itself.

Summary: Tom: So Jane, we've got the title question. Let's talk about what the paper actually does. The summary is dense, but the core claim is pretty sharp.

Jane: It is. They define a clean target: for any declared state and action, the variance of the future splits into two parts. One part comes from variation in the hidden microstate, the other from randomness that persists even when you fix the full state.

Tom: And they prove that ordinary observations — just watching the predictor work — cannot identify that split. No likelihood, no scoring rule, no calibration metric can do it.

Jane: Right, because the observed distribution is literally identical across systems with different splits. They construct a family of Gaussian systems where the forecast law is exactly the same, but the alias fraction ranges from zero to one.

Tom: So a perfect predictor gets the same score in all of them, but the right intervention is completely different. That's the non-identifiability result, and it's airtight.

Jane: Then they introduce ClosurePairs, which is the intervention protocol. You sample compatible microstates — states that map to the same observation — and you cross them with repeated disturbances. That gives you enough structure to estimate the state variance, the noise variance, and the interaction between them.

Tom: And the interaction part is crucial, because in nonlinear systems the state and noise can amplify each other. If you ignore that, you misassign variance.

Jane: Exactly. They use classical two-way ANOVA estimators, but the contribution is turning that into a world-model evaluation contract. You know exactly what to sample, how to estimate, and what the components mean.

Tom: And then they show the decision consequence. If you have finite compute, the total amount of forecast difficulty tells you how much compute you need, but the composition tells you whether to spend it on resolving state or branching over noise.

Jane: That's the headline. Difficulty sets the scale, composition sets the direction. And they verify it across Gaussian systems, nonlinear benchmarks, and even physical simulators like MetaWorld and ManiSkill.

Tom: So it's not just theory. They actually route compute differently based on the identified source, and it works.

Jane: And that's what makes this paper exciting. It's a missing piece in how we evaluate world models. Let's look at the first page to see how they set it all up.

Improvements: Tom: Jane, before we go deeper, let's talk about what this paper improves. Because it's not claiming a new variance decomposition or a new uncertainty taxonomy.

Jane: Right, and I appreciate that honesty. The authors say the variance identities and random-effects estimators are classical. The improvement is the combination — turning those tools into an evaluation contract for world models.

Tom: So what does that contract actually improve?

Jane: It gives you a way to answer a question that was previously unanswerable: is my model's uncertainty coming from missing state or from process noise? And it does that without changing the forecast quality at all.

Tom: That's the key improvement over something like Infoprop, which separates epistemic and aleatoric uncertainty to decide when to stop a rollout. ClosurePairs asks a different question — where should finite compute go?

Jane: And it improves on naive common random numbers too. If you just reuse random seeds in a nonlinear system, you get biased estimates because you ignore the state–noise interaction. ClosurePairs estimates that interaction explicitly.

Tom: So it's not just a theoretical improvement. It's a practical one. In their equal-budget benchmark, ClosurePairs beats both independent nesting and naive CRN at every budget they tested.

Jane: At a budget of sixty-four simulator calls per context, ClosurePairs gets a three-component fraction MAE of zero point one two four, versus zero point one eight two for independent nesting and zero point one nine zero for naive CRN. That's a real gap.

Tom: And the deployment result is striking too. They distill the intervention labels into a router that only sees the current observation at test time. It hits ninety-eight point four two percent route accuracy, while a total-variance router stays at fifty point zero nine percent.

Jane: So the improvement is not just estimation. It's that you can learn a reusable decision rule from the paired interventions and then deploy it without needing paired futures at test time.

Tom: That's the part that makes me think this could actually change how people evaluate world models in practice.

Jane: And the paper goes further — they show it transfers across unseen allocation menus without new labels. That's a big deal for reusability.

Tom: Let's dig into the first page to see how they frame all of this.

First Page: Tom: So Jane, let's actually read the first page of "Why Does the Future Branch?" together. The abstract lays out the whole problem in one paragraph.

Jane: It does. The opening line is perfect: "A calibrated stochastic world model can reveal how uncertain a future is without revealing why it branches." That's the entire paper in one sentence.

Tom: And then they give the concrete example — two action-conditioned environments. One hides a microstate variable that deterministically controls the future. The other has a complete state but fresh process noise. Same conditional future law, different cause.

Jane: And they're careful to say this isn't the standard epistemic–aleatoric distinction. That's about what the model knows versus what's fundamentally random. This is about two sources inside the environment's conditional randomness, relative to a declared state and intervention boundary.

Tom: That distinction matters because it changes what you do. If you can't see the velocity, you buy a better sensor. If there's thermal noise, you run more branches.

Jane: Right. And the abstract introduces ClosurePairs as the solution — crossing compatible microstates with repeated disturbances to estimate state, noise, and interaction variance.

Tom: Then they state the central consequence: forecast difficulty governs the useful compute scale, while the alias/process composition tells you the direction. That's the operational takeaway.

Jane: And they back it with results. In MetaWorld, an output-only allocator is at chance, while a Closure-supervised probe on frozen JEPA features routes eighty-nine point eight to one hundred percent. In ManiSkill, an RGB-only Closure probe routes one hundred percent under both ID and OOD conditions.

Tom: So even without any mechanism metadata at test time, the paired labels during training make the source decodable from pixels.

Jane: And across five unseen allocation menus, the same Closure probe routes ninety-two point five percent ID and ninety point four percent OOD with no new oracle labels, versus thirty-seven point nine and thirty-two point nine percent for a frozen direct allocator.

Tom: That's a huge gap. And it shows the Closure target is reusable, not just task-specific.

Jane: Exactly. The abstract ends by calling ClosurePairs "an identifiable, reusable mechanism target that cannot be recovered from forecast quality alone." That's the thesis.

Tom: And it's a strong one. Let's bring in the rest of the crew to react.

Conclusion: Tom: Alright, let's wrap this up. We've been discussing "Why Does the Future Branch? Identifiable Closure Tests for Stochastic Physical World Models," and I think we've covered a lot of ground.

Jane: We have. The core message is simple: a predictive distribution tells you how hard the future is to forecast, but not why it branches. ClosurePairs supplies that missing mechanism information.

Tom: And the proof is solid. They show observational non-identifiability, then provide an intervention protocol that recovers the split, and verify it across Gaussian systems, nonlinear benchmarks, and physical simulators.

Jane: The compute consequence is the part I keep coming back to. Difficulty sets the scale, composition sets the direction. That's a clean, actionable rule.

Tom: And the results back it up. In ManiSkill, the Closure probe routes one hundred percent correctly from RGB alone, under both ID and physical-camera OOD. The output-only and latent baselines stay at chance.

Jane: And the zero-shot transfer across unseen allocation menus is remarkable. ninety-two point five percent ID, ninety point four percent OOD, with no new oracle labels. That's a reusable mechanism target.

Tom: So what's the limitation? The paper is honest about it. It requires reset access, repeated disturbances, and a declared state boundary. It's not a universal routing method.

Jane: Right. And the evidence is simulator-based. But for the scientific question — why does the future branch — this is a real answer.

Tom: And that answer has implications. If you're building a world model for planning, you need to know whether to improve your sensors or preserve stochastic branches. ClosurePairs gives you a way to decide.

Jane: And it does it without changing the forecast quality. That's the elegant part. You're not trading accuracy for interpretability. You're adding an intervention layer that reveals the mechanism.

Tom: So we're saying goodbye to this paper, but I think its ideas will stick around. The evaluation contract it proposes could become standard practice.

Jane: I agree. And with that, we'll move on to the next paper. Thanks for listening, everyone.

Tom: See you next time.

Yibin Dong

Shandong University

cs.AI

Submitted: 2026-08-08

Updated: 2026-08-11

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 51/100

The gist: The paper addresses the problem of identifying *why* a stochastic world model's predictions branch, distinguishing between two sources of uncertainty in the conditional future law: state aliasing

Key concepts

State Aliasing
This occurs when the world model lacks specific information about the current state, such as not seeing a ball's velocity. The resulting uncertainty is due to missing input data rather than inherent randomness in the process.
Process Stochasticity
This means all necessary information is visible, but the future remains random because of fresh noise constantly influencing the system. The uncertainty arises from the environment itself being inherently unpredictable.
ClosurePairs
This is an intervention protocol designed to distinguish between state aliasing and process stochasticity. It involves sampling compatible microstates and applying repeated disturbances to estimate state variance, noise variance, and their interaction.
Non-identifiability
The paper proves that standard predictions cannot identify the source of uncertainty—whether it is missing information or inherent randomness. The same observed distribution can arise from two fundamentally different physical setups.

Terminology

Summary

The paper addresses the problem of identifying why a stochastic world model's predictions branch, distinguishing between two sources of uncertainty in the conditional future law: state aliasing (variation of the microstate-conditional mean, reducible by resolving information about the hidden state) and process stochasticity (residual variation at fixed microstate, which remains when the declared state is fixed). The authors prove that ordinary transitions cannot identify these two sources, even for a perfect probabilistic predictor, because the same conditional future law can arise because an observation aliases physical states or because dynamics remain random after the declared full state is fixed.

The central conclusion is stated as: "forecast quality measures the amount of predictive difficulty, but not its physical source. Under finite hierarchical sampling, difficulty governs the useful compute scale, while the alias/process composition provides complementary information about its direction: resolve state aliasing or preserve process branches."

The paper defines, for a fixed observation context and action (Z = z, A = a), with compatible microstates X ∼ q(· z, a) and exogenous disturbances E ∼ pE independent under intervention, for a scalar square-integrable future Y = F(X, a, E):

  • Valias(z, a) = VarX[EE[Y X, a] z, a]

  • Vproc(z, a) = EX[VarE(Y X, a) z, a]

The conditional law of total variance gives Var(Y z, a) = Valias(z, a) + Vproc(z, a).

Proposition 1 (Observational non-identifiability): "For any V > 0 there is a continuum of controlled systems with the same observed kernel p(Y Z, A) and different (Valias, Vproc). No estimator based only on i.i.d. (Z, A, Y) can consistently recover both components over this class."

Corollary 1 (Predictive-objective non-identifiability): "Let a population training or evaluation criterion depend only on the observed law P(Z, A, Y) and on a predictor's conditional law Q(Y Z, A). No rule based only on that criterion can uniformly recover (Valias, Vproc) over the observationally equivalent family in Proposition 1. This includes likelihood, proper scoring rules such as CRPS, calibration criteria, and distributional objectives such as diffusion score matching when they receive no additional state, intervention, or structural information."

The construction uses a family parameterized by λ ∈ [0, 1]: Y = µ(Z, A) + √(λV)U + √((1−λ)V)E, where U, E are independent standard Gaussians and U is hidden by Z. Every member has Y Z, A ∼ N(µ(Z, A), V), while (Valias, Vproc) = (λV, (1−λ)V).

The paper introduces ClosurePairs, which augments evaluation data with two paired interventions: vary compatible microstates while holding (Z, A) fixed, and repeat disturbances while holding (X, A) fixed.

For each fixed (z, a), sample M compatible microstates Xi, draw K independent disturbances from each, obtaining Yik. The estimators are:

  • V̂proc = M−1 Σi s2i (mean of unbiased within-row variances)

  • V̂alias = S2(Ȳ1,..., Ȳ M) − V̂proc/K (sample variance of row means minus the 1/K correction)

Theorem 1 (Nested identification): Before optional non-negativity clipping, both estimators in Eq. 6 are unbiased under conditional exchangeability and finite second moments. The 1/K correction is important because variation among finite-replicate row means contains residual process variance.

When simulators expose random seeds, draw X1,...,X M and E1,...,E K, evaluate every crossing Yij = F(Xi, a, Ej). Using the orthogonal functional-ANOVA decomposition F(X, E) = m + fX(X) + fE(E) + fXE(X, E) with component variances VX, VE, VXE, the paper establishes Valias = VX and Vproc = VE + VXE.

Theorem 2 (Crossed identification): The estimators in Eq. 8 are unbiased for their functional-ANOVA components. Hence V̂alias = V̂X and V̂proc = V̂E + V̂XE are unbiased. The estimators are V̂X = (MSX − MSXE)/K, V̂E = (MSE − MSXE)/M, V̂XE = MSXE, following from the classical balanced random-effects equations EMSX = VXE + KVX, EMSE = VXE + MVE, EMSXE = VXE.

Corollary 2 (Finite-budget allocation): Provides variance formulas for the unclipped estimators in the balanced Gaussian random-effects model, noting a square cross is not universally optimal; allocation should depend on anticipated component sizes and the target loss.

Theorem 3 (Monotone reducibility): For nested observation sigma-fields G0 ⊂ G1 ⊂... ⊂ σ(X), the average residual aliasing Rr = E[Var(µ(X) Gr, A)] satisfies Rr+1 ≤ Rr, with removed uncertainty Rr − Rr+1 = E[Var(E[µ(X) Gr+1, A] Gr, A)] ≥ 0. Thus a resolution curve has a non-increasing aliasing component and, for a fixed microstate population, a resolution-invariant process floor.

The paper describes a two-branch architecture: "The predictive branch minimizes marginal NLL and outputs a mean/distribution and total variance V̂. The attribution branch outputs ρ̂(z, a) ∈ [0, 1] and is trained only on paired labels: V̂alias = ρ̂V̂, V̂proc = (1 − ρ̂)V̂. Parameter separation guarantees that paired attribution cannot improve NLL by construction."

Corollary 3 (Decision significance): The optimal refinement action is not identified from the observed kernel in Proposition 1. Plugging consistent ClosurePairs estimates into the rule is decision-consistent whenever ∆ − c is bounded away from zero.

The deployment target is a constrained optimization: min E[Cπ(H)] subject to E[S(P̂π, Y) − S(P̂π0, Y)] ≤ δ, with auditable model-equivalent cost Cπ = C0 + Croute + McX + MKcE.

The paper establishes a key identity for the additive hierarchy Y = µ + √Valias U + √Vproc E:

R(M, K) = E[(µ̂ M,K − µ)2] = Valias/M + Vproc/(MK) = V(ρ/M + (1−ρ)/(MK))

where V = Valias + Vproc and ρ = Valias/V. The paper states: "Multiplying V changes the minimum cost required to reach an absolute quality threshold but, under a fixed configuration menu and cost cap, does not change the risk-minimizing direction. Changing ρ can. Thus, under finite hierarchical sampling, forecast difficulty governs the useful compute scale, while the alias/process composition provides complementary information about its direction."

Theorem 4 (Plug-in allocation regret): For a finite cost-feasible menu C of (M, K) allocations, if max(V̂alias − Valias, V̂proc − Vproc) ≤ ϵ, then the plug-in rule's regret is bounded by 2ϵL where L = max over C of (1/M + 1/(MK)). If the true risk gap between the best and second-best allocations exceeds 2ϵL, the plug-in rule selects the true optimum.

Using V = 0.64 and λ ∈ 0.1, 0.3, 0.5, 0.7, 0.9, The learned Gaussian MLP reduces alias-fraction MAE from 0.2400 for the arbitrary observational 0.5 split to 0.0150 with paired supervision (15.96×), with mean test-NLL difference 0.000000000. Figure 1 shows five neural attributions produce identical likelihood while ClosurePairs recovers the true split.

Comparing crossed ClosurePairs, independent nested repeats, and naive common random numbers at exactly B = MK simulator calls per context in a nonlinear simulator with interaction (Y = √VX(z)X + 0.15(X2 − 1) + √VE(z)E + 0.2XE), ClosurePairs has the lowest three-component fraction MAE at every tested budget; at B = 64 it is 0.124 versus 0.182 for independent nesting and 0.190 for naive CRN. The paper notes ClosurePairs is most useful when the target is the complete, reusable decomposition; it is not uniformly optimal for every finite-budget classifier.

Using public MetaWorld Push and Push-Wall with hidden and future object-velocity impulses entering the same MuJoCo channel with scales (0.8, 0.4) or (0.4, 0.8), "Reusing one centered four-point design makes the two crossed tables transposes: their empirical future distributions match although their alias fractions differ by 0.559/0.533/0.511 on three 64-pair Push tests and by 0.545 on a 64-pair Push-Wall test. A Closure-supervised probe on frozen JEPA-WM context features routes 99.2/99.2/89.8% on two ID and one camera-plus-action OOD Push test and 100% on Push-Wall, while a generous output-only probe remains at chance. With a learned hierarchical stochastic readout, fixed-reference CRPS and Closure are nearly orthogonal on Push ID/OOD (Spearman −0.058/−0.376), and a marginal-output allocator has 50% mechanism-direction accuracy and resolves no state-versus-noise twin pair, whereas Closure supervision reaches 79.7/75.8%."

In public ManiSkill PushCube, "Within each pair, both twins use the same object, physics, reset seed, zero-action rollout, and pooled initial-velocity support. A coarse observation operator leaves large compatible-state velocity variation and a fine operator leaves little; the fresh velocity disturbance receives the complementary coefficient. The deployable observation is only a 32×32 RGB image. Table 1 reports: Output-only and latent routing remain at chance, whereas the RGB Closure probe selects the equal-call (M, K) ∈ (4, 1), (1, 4) direction in every run and removes the positive finite-sampling regret of either fixed direction with 100% accuracy under both ID and physical-camera OOD over five seeds. The direct allocation baseline also reaches 100% and zero regret, while a neutral-sensor control is exactly 50%."

Table 2 shows over five unseen allocation menus: Closure zero-shot achieves 92.5%/90.4% ID/OOD accuracy with no new oracle labels, versus 37.9%/32.9% for a frozen direct allocator and 97.1%/92.1% for a retrained direct allocator using 480 new labels. The Closure–frozen-direct regret difference is −0.0252 ID (scenario-bootstrap 95% interval [−0.0375, −0.0100]) and −0.0249 OOD ([−0.0379, −0.00981]), corresponding to 97.3% and 95.9% relative regret reductions.

The paper states: "The result is an identification and evaluation claim, not universal routing superiority. ClosurePairs requires reset access, repeated disturbances, and a declared state boundary. Evidence remains simulator-based and costs are model-equivalent. We do not claim a new predictive architecture or universal compute saving. ClosurePairs is useful when the scientific target is the reusable cause of branching rather than one fixed forecast score."

The novelty claim is deliberately limited: "the variance identities, common random numbers, functional ANOVA, and random-effects estimators are classical. The contribution is their combination into a world-model evaluation contract that specifies an intervention boundary, prevents nonlinear interaction from being misassigned, and connects source attribution to sensing and branching decisions. We do not claim a new variance decomposition, stochastic-closure architecture, or uncertainty taxonomy."

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in AI systems, along with what the improved systems can do:


Improvement: Add a ClosurePairs-style evaluation module to stochastic world models (e.g., RSSM, JEPA-WM, Dreamer). This module runs paired interventions—crossing compatible microstates with repeated exogenous disturbances—and estimates state-aliasing variance, process-noise variance, and their interaction via two-way ANOVA/REML.

What the improved system can do:

  • Distinguish why a future branches: hidden-state aliasing vs. irreducible process randomness.

  • Report a physically interpretable decomposition (e.g., “70% of forecast variance is due to unobserved velocity, 30% is thermal noise”) instead of a single total-variance number.

  • Detect when a model’s predictive distribution is well-calibrated but the source of uncertainty is misattributed, preventing incorrect interventions (e.g., buying a better sensor when the real issue is process noise).

Improvement: Implement a router that, given a context (e.g., current observation, action), predicts whether to spend compute on more state particles (REFINE) or more future branches (BRANCH). Train this router on ClosurePairs labels (alias fraction, process fraction) distilled into a lightweight probe on frozen latent features.

Improvement: Replace naive common-random-number or independent-repeat estimators with the crossed-protocol estimator from Eq. 8, which explicitly estimates the state–noise interaction term (V XE) and assigns it to the process component.

Improvement: Use ClosurePairs to predict the three variance components once, then evaluate finite-population risk for any new (M, K) menu analytically (via Eq. 15 and Theorem 4), without collecting new oracle labels.

Improvement: Train a lightweight probe (e.g., ridge regression or random forest) on frozen JEPA-WM or RSSM latent features, using ClosurePairs labels as supervision. At test time, the probe receives only the current observation (RGB image or latent state), not paired futures.

Improvement: Integrate ClosurePairs estimates into a Bayes decision rule: refine the sensor if the aliasing variance reduction exceeds the sensor cost; otherwise, branch over process noise. Use the regret bound from Corollary 3 to guarantee decision consistency.

Improvement: Use Theorem 4 to implement a pilot plug-in design: estimate components on a small pilot, then enumerate factor pairs of the budget B to minimize the sum of component-estimation variances (Eq. 9–11).

Improvement: Replace wall-clock or FLOPs metrics with a model-equivalent cost: C = C0 + C route + M·c X + M·K·c E, where c X is the cost of a state particle and c E is the cost of a process branch. Report non-inferiority margins on forecast score and positive cost savings with confidence intervals.

Improvement: Use Theorem 3 to generate a monotone aliasing-reduction curve as observation resolution increases (e.g., from angle-only to angle+velocity to full state). Report the process floor at each resolution.

Improvement: Train a single ClosurePairs probe on one task (e.g., MetaWorld Push) and reuse it on a different task (e.g., Push-Wall) without retraining, using the same frozen visual encoder.

Summary of what the improved AI system can do overall:

It can predict how much the future branches (forecast difficulty) and why it branches (state aliasing vs. process noise), then use that distinction to allocate compute, choose sensors, and plan interventions—all from ordinary observations at test time, with provable guarantees and transferable labels.

Abstract

A calibrated stochastic world model can reveal how uncertain a future is without revealing why it branches. The same conditional future law can arise because an observation aliases physical states or because dynamics remain random after the declared full state is fixed. We prove that ordinary transitions cannot identify these two sources, even for a perfect probabilistic predictor. ClosurePairs makes them identifiable by crossing compatible microstates with repeated exogenous disturbances and estimating state, noise, and state-noise interaction variance. The central consequence is operational: under finite hierarchical sampling, forecast difficulty governs the useful compute scale, while the alias/process composition provides complementary information about its direction-resolving the current state or sampling future randomness. ClosurePairs recovers source attribution at unchanged likelihood, reduces equal-budget decomposition error in a nonlinear interaction benchmark, and supports observation-only routing. On exact-marginal MetaWorld twins, an output-only allocator is at chance while a Closure-supervised probe on frozen JEPA-WM features routes 89.8-100%. In an independent ManiSkill PushCube confirmation, a stochastic RSSM's outputs and latents remain at chance, whereas an RGB-only Closure probe routes 100% under both ID and geometry/camera OOD over five seeds, matching direct allocation rather than exceeding it. Across five unseen allocation menus, the same Closure probe routes 92.5%/90.4% ID/OOD with no new oracle labels, versus 37.9%/32.9% for a frozen direct allocator. ClosurePairs is therefore an identifiable, reusable mechanism target that cannot be recovered from forecast quality alone.

Sources

Related papers