AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses

arXiv:2608.12307 · cs.LG, cs.AI, cs.CL · Submitted 2026-08-12 · Read on arXiv

Cheng Qian, Wenting Zhao, Liangwei Yang, Heng Wang, Jielin Qiu, Heng Ji, Silvio Savarese, Huan Wang, Shelby Heinecke

Salesforce AI Research · University of Illinois Urbana-Champaign

cs.LG, cs.AI, cs.CL

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: 23 Pages, 12 Figures, 6 Tables

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 95/100

The gist: This paper, from Salesforce AI Research and the University of Illinois Urbana-Champaign, investigates whether capability transfer from strong to weak language models can occur at test time rather

Terminology

Summary

This paper, from Salesforce AI Research and the University of Illinois Urbana-Champaign, investigates whether capability transfer from strong to weak language models can occur at test time rather than through conventional training-time distillation. The authors formalize a setting called strong-to-weak scaffolding: whether a stronger builder model can construct inference-time harnesses that help a weaker target model solve tasks more reliably without any parameter updates.

The setup involves a strong builder model that constructs an inference-time scaffold for a fixed weaker target model. The builder has access only to a 5% validation split of each benchmark, while the remaining 95% is held out as a hidden test set. The builder iteratively refines its harness over multiple rounds using validation feedback, after which the finalized scaffold is evaluated on the full test set. The builder "may implement any inference-time procedure, including prompt templates, benchmark routing, deterministic pre- or post-processing, answer-format enforcement, verification passes, few-shot retrieval, or direct symbolic solvers."

The study uses four Theory-of-Mind (ToM) benchmarks aggregated into a 3,900-item hidden test set:

  • BigToM: 1,200 data points; binary belief/goal/action questions hinging on whether an agent observed a world change

  • Hi-ToM: 1,200 data points; nested belief questions of recursion order 0–4, with deception and multi-room object tracking

  • MMToM-QA: 600 data points; binary Bayesian goal/belief inference from action traces

  • MuMA-Tom: 900 data points; 3-choice multi-agent belief/social-goal/belief-of-goal questions

The experimental design controls four hyper-parameters: platform (Cursor, Claude Code, GPT Codex), builder model (11 configurations including Opus-4.7 at four reasoning effort tiers, Sonnet-4.6, GPT-5.5, GPT-5.4-mini, Codex-5.3, Gemini-3.1-Pro, Gemini-3.5-flash, and Grok-0.1), target model (GPT-5.4-mini and Gemini-3.5-flash), and repeats (3 per setting), yielding 72 total experiment runs.

Baselines include: (1) Vanilla: direct target model calls with naive prompts, yielding macro-average accuracy of 0.488 for GPT-5.4-mini and 0.761 for Gemini-3.5-flash; (2) Human-Inspired Harness (UserHarness): a human-designed harness framework for ToM problems, yielding 0.939 for GPT-5.4-mini and 0.941 for Gemini-3.5-flash.

Aspect 0: Main Results. Strong-to-weak scaffolding produces a large and uniformly positive effect. Across all 57 scaffolded GPT-5.4-mini runs, the mean macro-average accuracy is 0.763, an uplift of +0.275 over the 0.488 baseline, with 100% of runs exceeding baseline. The best individual run, produced by GPT-5.5 on GPT Codex, reaches 0.912, an uplift of +0.423 (86.7% relative improvement). The best scaffold outperforms both vanilla baselines on all four benchmarks, exceeding raw GPT-5.4 and GPT-OSS-120B performance. On BigToM, the automated scaffold reaches near-ceiling performance (1.00 vs. 0.95 for UserHarness), but gaps remain on Hi-ToM (0.80 vs. 0.87), MMToM-QA (0.84 vs. 0.98), and MuMA-Tom (0.88 vs. 0.96).

Aspect 1: Run-to-Run Stability. The scaffold building procedure is fairly stable, with mean standard deviation of 0.036 across repeats—roughly an order of magnitude smaller than the +0.275 mean uplift. However, the widest setting has a repeat range of 0.201. The largest spreads occur in settings where builders pursue deterministic-solver strategies: a single logic error in a benchmark-specific rule can shift accuracy by tens of points.

Aspect 2: Refinement on Validation. Builders use validation evaluations sparingly (mean 4.9, median 5, range 2–15). Refinement is productive: mean validation accuracy increases by 0.216 from first to best logged iteration. The best validation score is a strong proxy for held-out performance, with Pearson r = 0.96 and only a small optimism gap (mean 0.021). However, the number of validation iterations is essentially uncorrelated with final full-set accuracy (Pearson r = 0.17). The authors conclude: the limiting factor is not the amount of feedback, but the quality of the builder's hypotheses.

Aspect 3: Scaffolding Techniques. Analysis of scaffold code reveals a twelve-technique taxonomy. Two techniques are nearly universal: robust format enforcement (100% of runs) and greedy decoding with temperature 0 (98%). Other common techniques include benchmark routing (95%), forced chain-of-thought (79%), polarity/negation logic (79%), and token-budget tuning (75%). More complex strategies are less common: deterministic solvers (54%), self-consistency voting (5%), and verification/arbiter passes (12%). Technique choice is strongly task-dependent: BigToM is often solved deterministically or with hybrid rules... whereas MuMA-ToM is almost always handled by the model itself.

Aspect 4: Platform Effects. Platform effects are second-order and conditional rather than universal. In matched native-vs-Cursor comparisons, moving from Cursor to native platform changes macro accuracy by only +0.013 (paired permutation test p = 0.484). However, there is a platform×effort interaction: for Opus-4.7, Claude Code trails Cursor at low effort (−0.034) but leads at higher efforts (medium +0.045, high +0.038, extra-high +0.032). The authors conclude: builder identity dominates platform identity.

Aspect 5: Target Model Analysis. The weaker target benefits much more from scaffolding: GPT-5.4-mini improves by +0.262 (0.488 → 0.750) versus Gemini-3.5-flash's +0.110 (0.761 → 0.871). The authors identify a headroom law: realized uplift is strongly predicted by the target's available headroom (1 − baseline accuracy), with Pearson r = 0.75. For Gemini-3.5-flash, gains concentrate almost entirely on BigToM (96% of macro uplift). Builders adapt strategies: the share of model-only benchmark handling rises on every task when scaffolding the stronger target. Notably, on a strong target, scaffolding can backfire—every builder regresses on at least one benchmark for Gemini-3.5-flash (9/20 cases), particularly on Hi-ToM (−0.04 average) and near-saturated MuMA-ToM (−0.02 average).

Aspect 6: Builder Reasoning Effort. For Opus-4.7 with GPT-5.4-mini target, macro accuracy increases monotonically with effort: 0.711 (low) → 0.793 (medium) → 0.807 (high) → 0.856 (extra-high), with Spearman ρ = 0.77. The extra-high tier significantly outperforms both high (p = 0.013) and low (p = 0.002) tiers. The largest gain occurs from low to medium effort, with higher tiers providing smaller but positive improvements. Scaffold size also grows with effort (∼510–650 LOC at low vs. ∼1000–1300 at extra-high).

Aspect 7: Attribution of Improvement. The strongest positive associations with accuracy come from techniques that compile task structure: polarity/negation logic (+0.09), structured extraction (+0.06), hybrid fallback (+0.04), and deterministic solving. The best scaffold produces a large, item-level reliable improvement under paired McNemar testing (χ2 ≫ 104, p < 10−4), fixing 1,717 baseline errors while breaking only 105 previously correct items. Self-scaffolding (GPT-5.4-mini building for itself) improves performance by +0.17 to +0.22, showing even a weak target can use validation feedback and task structure to engineer a useful harness, but stronger builders achieve substantially larger gains. Different strong scaffolds are complementary: the union of items fixed across top scaffolds covers 97% of all baseline errors, exceeding any individual scaffold's coverage.

Aspect 8: Cognitive-Load Reduction. Deterministic offloading is strongly associated with scaffold quality (Pearson r = 0.72): runs with higher determinism fractions (share of items answered entirely by code or structured rules) achieve higher accuracy. This relationship is benchmark-dependent: BigToM is almost fully offloadable (mean determinism ≈ 0.94), Hi-ToM partially (≈ 0.51), MMToM-QA similar but slightly lower (≈ 0.44), and MuMA-ToM least reducible (≈ 0.36). Scaffold code size is only weakly related to accuracy (r ≈ 0.22), indicating what matters is not simply writing more code, but writing code that removes the right cognitive load.

Aspect 9: Remaining Error Analysis. Top scaffolds are broadly corrective rather than merely redistributive: on average they repair 83% of baseline-wrong items and break only 7% of baseline-correct items. Residual errors concentrate in three hard regions: (1) Hi-ToM accuracy declines with recursion depth (0.999 at order 0 to 0.700 at order 4, with deception further reducing performance); (2) MMToM-QA errors concentrate in Bayesian goal-inference subtypes, especially type-2 which container/goal questions (accuracy 0.680); (3) MuMA-ToM remains difficult on social-goal (0.872) and belief-of-goal (0.880) labels, though simpler belief questions are nearly solved (0.985).

The paper makes three main contributions:

  1. Formalizing strong-to-weak scaffolding as a distinct inference-time capability transfer setting

  2. Providing a systematic empirical analysis of its effect size, stability, validation efficiency, platform and target dependence, mechanisms, and cognitive-load reduction

  3. Identifying actionable design principles: successful scaffolds often rely on deterministic offloading, benchmark-aware routing, format control, and targeted decomposition than brute-force validation search

The central takeaway is that "strong-to-weak scaffolding works because a capable builder can act as a compiler of task competence. It spends a one-time reasoning budget to identify structure in the task and encode that structure into an inference-time scaffold. Once compiled, a weaker and cheaper target model can execute the task at a level closer to that of much stronger models. The authors emphasize that scaffolding does not replace raw reasoning capability; it reallocates it."

The practical recipe suggested is: "use the strongest available builder, allocate high reasoning effort during scaffold construction, spend only a modest number of validation evaluations, prioritize cognitive offloading for provable sub-tasks, and, when budget permits, build several independent scaffolds and select or ensemble them to capture complementary repairs."

The paper also discusses implications for harness self-evolution, proposing that the setting could become a standard benchmark for builder models, and frames the work within the broader context of two complementary routes to stronger systems: improving internal model capability versus making tasks easier for models to execute.

Improvements for AI systems

Improvements to AI Systems:

  1. Implement a Test-Time Scaffolding Module: Add a system component where a stronger builder model constructs inference-time harnesses (prompt templates, routing logic, deterministic solvers, format enforcement) for a weaker target model, enabling capability transfer without retraining or parameter updates.

  2. Add Benchmark-Aware Routing: Build a router that classifies incoming queries by task type and directs them to specialized scaffolds—deterministic solvers for rule-based tasks (e.g., BigToM), model-based reasoning for complex social inference (e.g., MuMA-ToM)—improving accuracy by up to +0.423 over vanilla baselines.

  3. Integrate Deterministic Offloading for Provable Subtasks: Automatically detect sub-problems with verifiable logic (e.g., polarity/negation checks, structured extraction) and compile them into code, reducing cognitive load on the target model. This correlates strongly with quality (r = 0.72) and yields near-ceiling performance on fully offloadable benchmarks.

  4. Enable Iterative Validation-Based Refinement: Allow the builder to evaluate its scaffold on a small validation split (mean 4.9 iterations) and refine based on feedback, achieving a validation-to-test correlation of r = 0.96, ensuring robust generalization without overfitting.

  5. Add Effort-Adaptive Scaffold Construction: Dynamically allocate builder reasoning effort based on task difficulty—higher effort for complex benchmarks (e.g., Hi-ToM recursion depth) yields monotonic accuracy gains (0.711 → 0.856), with the largest jump from low to medium effort.

  6. Implement Complementary Scaffold Ensembling: Build multiple independent scaffolds and combine their outputs, covering 97% of baseline errors (vs. any single scaffold’s coverage), reducing residual errors by repairing 83% of wrong items while breaking only 7% of correct ones.

  7. Add Headroom-Aware Uplift Prediction: Before scaffolding, estimate the target model’s headroom (1 − baseline accuracy) to predict realized uplift (r = 0.75), enabling cost-benefit decisions on whether to scaffold a given model or task.

  8. Prevent Scaffolding Backfire on Strong Targets: For high-capability targets, monitor per-benchmark regression (e.g., Hi-ToM −0.04 average) and selectively apply scaffolding only where it improves, avoiding harmful over-engineering on near-saturated tasks.

  9. Incorporate Recursion-Depth-Aware Handling: For nested belief tasks (Hi-ToM), add specialized decomposition that breaks down recursion order >2 into sub-steps, mitigating accuracy decline from 0.999 (order 0) to 0.700 (order 4).

  10. Add Self-Scaffolding Capability: Enable weaker models to build their own harnesses using validation feedback, yielding +0.17 to +0.22 improvement—useful when stronger builders are unavailable or costly.

What the Improved AI System Can Do:

  • Achieve up to 86.7% relative accuracy improvement on Theory-of-Mind tasks (0.488 → 0.912) using a weaker target model with no retraining.

  • Match or exceed human-designed harness performance on simple belief tasks (1.00 vs. 0.95) and approach it on complex social reasoning (0.88 vs. 0.96).

  • Generalize across diverse benchmarks (binary, multi-choice, nested recursion, Bayesian inference) via automated routing and technique selection.

  • Operate cost-effectively: use cheap target models at inference time while offloading reasoning to a one-time builder effort.

  • Provide stable, reproducible performance (mean std. dev. 0.036 across runs) with efficient validation use (median 5 iterations).

Sources

Related papers