AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses
Cheng Qian, Wenting Zhao, Liangwei Yang, Heng Wang, Jielin Qiu, Heng Ji, Silvio Savarese, Huan Wang, Shelby Heinecke
Salesforce AI Research · University of Illinois Urbana-Champaign
cs.LG, cs.AI, cs.CL
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: 23 Pages, 12 Figures, 6 Tables
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 95/100
The gist: This paper, from Salesforce AI Research and the University of Illinois Urbana-Champaign, investigates whether capability transfer from strong to weak language models can occur at test time rather
Terminology
Summary
This paper, from Salesforce AI Research and the University of Illinois Urbana-Champaign, investigates whether capability transfer from strong to weak language models can occur at test time rather than through conventional training-time distillation. The authors formalize a setting called strong-to-weak scaffolding: whether a stronger builder model can construct inference-time harnesses that help a weaker target model solve tasks more reliably without any parameter updates.
The setup involves a strong builder
model that constructs an inference-time scaffold for a fixed weaker target
model. The builder has access only to a 5% validation split of each benchmark, while the remaining 95% is held out as a hidden test set. The builder iteratively refines its harness over multiple rounds using validation feedback, after which the finalized scaffold is evaluated on the full test set. The builder "may implement any inference-time procedure, including prompt templates, benchmark routing, deterministic pre- or post-processing, answer-format enforcement, verification passes, few-shot retrieval, or direct symbolic solvers."
The study uses four Theory-of-Mind (ToM) benchmarks aggregated into a 3,900-item hidden test set:
-
BigToM: 1,200 data points; binary belief/goal/action questions hinging on whether an agent observed a world change
-
Hi-ToM: 1,200 data points; nested belief questions of recursion order 0–4, with deception and multi-room object tracking
-
MMToM-QA: 600 data points; binary Bayesian goal/belief inference from action traces
-
MuMA-Tom: 900 data points; 3-choice multi-agent belief/social-goal/belief-of-goal questions
The experimental design controls four hyper-parameters: platform (Cursor, Claude Code, GPT Codex), builder model (11 configurations including Opus-4.7 at four reasoning effort tiers, Sonnet-4.6, GPT-5.5, GPT-5.4-mini, Codex-5.3, Gemini-3.1-Pro, Gemini-3.5-flash, and Grok-0.1), target model (GPT-5.4-mini and Gemini-3.5-flash), and repeats (3 per setting), yielding 72 total experiment runs.
Baselines include: (1) Vanilla: direct target model calls with naive prompts, yielding macro-average accuracy of 0.488 for GPT-5.4-mini and 0.761 for Gemini-3.5-flash; (2) Human-Inspired Harness (UserHarness): a human-designed harness framework for ToM problems, yielding 0.939 for GPT-5.4-mini and 0.941 for Gemini-3.5-flash.
Aspect 0: Main Results. Strong-to-weak scaffolding produces a large and uniformly positive effect.
Across all 57 scaffolded GPT-5.4-mini runs, the mean macro-average accuracy is 0.763, an uplift of +0.275 over the 0.488 baseline, with 100% of runs exceeding baseline. The best individual run, produced by GPT-5.5 on GPT Codex, reaches 0.912, an uplift of +0.423 (86.7% relative improvement). The best scaffold outperforms both vanilla baselines on all four benchmarks,
exceeding raw GPT-5.4 and GPT-OSS-120B performance. On BigToM, the automated scaffold reaches near-ceiling performance (1.00 vs. 0.95 for UserHarness), but gaps remain on Hi-ToM (0.80 vs. 0.87), MMToM-QA (0.84 vs. 0.98), and MuMA-Tom (0.88 vs. 0.96).
Aspect 1: Run-to-Run Stability. The scaffold building procedure is fairly stable,
with mean standard deviation of 0.036 across repeats—roughly an order of magnitude smaller than the +0.275 mean uplift.
However, the widest setting has a repeat range of 0.201. The largest spreads occur in settings where builders pursue deterministic-solver strategies: a single logic error in a benchmark-specific rule can shift accuracy by tens of points.
Aspect 2: Refinement on Validation. Builders use validation evaluations sparingly (mean 4.9, median 5, range 2–15). Refinement is productive: mean validation accuracy increases by 0.216 from first to best logged iteration. The best validation score is a strong proxy for held-out performance,
with Pearson r = 0.96 and only a small optimism gap (mean 0.021). However, the number of validation iterations is essentially uncorrelated with final full-set accuracy
(Pearson r = 0.17). The authors conclude: the limiting factor is not the amount of feedback, but the quality of the builder's hypotheses.
Aspect 3: Scaffolding Techniques. Analysis of scaffold code reveals a twelve-technique taxonomy. Two techniques are nearly universal: robust format enforcement
(100% of runs) and greedy decoding
with temperature 0 (98%). Other common techniques include benchmark routing (95%), forced chain-of-thought (79%), polarity/negation logic (79%), and token-budget tuning (75%). More complex strategies are less common: deterministic solvers (54%), self-consistency voting (5%), and verification/arbiter passes (12%). Technique choice is strongly task-dependent: BigToM is often solved deterministically or with hybrid rules... whereas MuMA-ToM is almost always handled by the model itself.
Aspect 4: Platform Effects. Platform effects are second-order and conditional rather than universal.
In matched native-vs-Cursor comparisons, moving from Cursor to native platform changes macro accuracy by only +0.013 (paired permutation test p = 0.484). However, there is a platform×effort interaction: for Opus-4.7, Claude Code trails Cursor at low effort (−0.034) but leads at higher efforts (medium +0.045, high +0.038, extra-high +0.032). The authors conclude: builder identity dominates platform identity.
Aspect 5: Target Model Analysis. The weaker target benefits much more from scaffolding: GPT-5.4-mini improves by +0.262 (0.488 → 0.750) versus Gemini-3.5-flash's +0.110 (0.761 → 0.871). The authors identify a headroom law
: realized uplift is strongly predicted by the target's available headroom (1 − baseline accuracy), with Pearson r = 0.75. For Gemini-3.5-flash, gains concentrate almost entirely on BigToM (96% of macro uplift). Builders adapt strategies: the share of model-only benchmark handling rises on every task
when scaffolding the stronger target. Notably, on a strong target, scaffolding can backfire
—every builder regresses on at least one benchmark for Gemini-3.5-flash (9/20 cases), particularly on Hi-ToM (−0.04 average) and near-saturated MuMA-ToM (−0.02 average).
Aspect 6: Builder Reasoning Effort. For Opus-4.7 with GPT-5.4-mini target, macro accuracy increases monotonically with effort: 0.711 (low) → 0.793 (medium) → 0.807 (high) → 0.856 (extra-high), with Spearman ρ = 0.77. The extra-high tier significantly outperforms both high (p = 0.013) and low (p = 0.002) tiers. The largest gain occurs from low to medium effort, with higher tiers providing smaller but positive improvements. Scaffold size also grows with effort (∼510–650 LOC at low vs. ∼1000–1300 at extra-high).
Aspect 7: Attribution of Improvement. The strongest positive associations with accuracy come from techniques that compile task structure: polarity/negation logic (+0.09), structured extraction (+0.06), hybrid fallback (+0.04), and deterministic solving. The best scaffold produces a large, item-level reliable improvement
under paired McNemar testing (χ2 ≫ 104, p < 10−4), fixing 1,717 baseline errors while breaking only 105 previously correct items. Self-scaffolding (GPT-5.4-mini building for itself) improves performance by +0.17 to +0.22, showing even a weak target can use validation feedback and task structure to engineer a useful harness,
but stronger builders achieve substantially larger gains. Different strong scaffolds are complementary: the union of items fixed across top scaffolds covers 97% of all baseline errors, exceeding any individual scaffold's coverage.
Aspect 8: Cognitive-Load Reduction. Deterministic offloading is strongly associated with scaffold quality
(Pearson r = 0.72): runs with higher determinism fractions (share of items answered entirely by code or structured rules) achieve higher accuracy. This relationship is benchmark-dependent: BigToM is almost fully offloadable (mean determinism ≈ 0.94), Hi-ToM partially (≈ 0.51), MMToM-QA similar but slightly lower (≈ 0.44), and MuMA-ToM least reducible (≈ 0.36). Scaffold code size is only weakly related to accuracy (r ≈ 0.22), indicating what matters is not simply writing more code, but writing code that removes the right cognitive load.
Aspect 9: Remaining Error Analysis. Top scaffolds are broadly corrective rather than merely redistributive
: on average they repair 83% of baseline-wrong items and break only 7% of baseline-correct items. Residual errors concentrate in three hard regions: (1) Hi-ToM accuracy declines with recursion depth (0.999 at order 0 to 0.700 at order 4, with deception further reducing performance); (2) MMToM-QA errors concentrate in Bayesian goal-inference subtypes, especially type-2 which container/goal
questions (accuracy 0.680); (3) MuMA-ToM remains difficult on social-goal (0.872) and belief-of-goal (0.880) labels, though simpler belief questions are nearly solved (0.985).
The paper makes three main contributions:
-
Formalizing strong-to-weak scaffolding as
a distinct inference-time capability transfer setting
-
Providing
a systematic empirical analysis of its effect size, stability, validation efficiency, platform and target dependence, mechanisms, and cognitive-load reduction
-
Identifying actionable design principles:
successful scaffolds often rely on deterministic offloading, benchmark-aware routing, format control, and targeted decomposition than brute-force validation search
The central takeaway is that "strong-to-weak scaffolding works because a capable builder can act as a compiler of task competence. It spends a one-time reasoning budget to identify structure in the task and encode that structure into an inference-time scaffold. Once compiled, a weaker and cheaper target model can execute the task at a level closer to that of much stronger models. The authors emphasize that
scaffolding does not replace raw reasoning capability; it reallocates it."
The practical recipe suggested is: "use the strongest available builder, allocate high reasoning effort during scaffold construction, spend only a modest number of validation evaluations, prioritize cognitive offloading for provable sub-tasks, and, when budget permits, build several independent scaffolds and select or ensemble them to capture complementary repairs."
The paper also discusses implications for harness self-evolution, proposing that the setting could become a standard benchmark for builder models, and frames the work within the broader context of two complementary routes to stronger systems: improving internal model capability versus making tasks easier for models to execute.
Improvements for AI systems
Improvements to AI Systems:
-
Implement a Test-Time Scaffolding Module: Add a system component where a stronger builder model constructs inference-time harnesses (prompt templates, routing logic, deterministic solvers, format enforcement) for a weaker target model, enabling capability transfer without retraining or parameter updates.
-
Add Benchmark-Aware Routing: Build a router that classifies incoming queries by task type and directs them to specialized scaffolds—deterministic solvers for rule-based tasks (e.g., BigToM), model-based reasoning for complex social inference (e.g., MuMA-ToM)—improving accuracy by up to +0.423 over vanilla baselines.
-
Integrate Deterministic Offloading for Provable Subtasks: Automatically detect sub-problems with verifiable logic (e.g., polarity/negation checks, structured extraction) and compile them into code, reducing cognitive load on the target model. This correlates strongly with quality (r = 0.72) and yields near-ceiling performance on fully offloadable benchmarks.
-
Enable Iterative Validation-Based Refinement: Allow the builder to evaluate its scaffold on a small validation split (mean 4.9 iterations) and refine based on feedback, achieving a validation-to-test correlation of r = 0.96, ensuring robust generalization without overfitting.
-
Add Effort-Adaptive Scaffold Construction: Dynamically allocate builder reasoning effort based on task difficulty—higher effort for complex benchmarks (e.g., Hi-ToM recursion depth) yields monotonic accuracy gains (0.711 → 0.856), with the largest jump from low to medium effort.
-
Implement Complementary Scaffold Ensembling: Build multiple independent scaffolds and combine their outputs, covering 97% of baseline errors (vs. any single scaffold’s coverage), reducing residual errors by repairing 83% of wrong items while breaking only 7% of correct ones.
-
Add Headroom-Aware Uplift Prediction: Before scaffolding, estimate the target model’s headroom (1 − baseline accuracy) to predict realized uplift (r = 0.75), enabling cost-benefit decisions on whether to scaffold a given model or task.
-
Prevent Scaffolding Backfire on Strong Targets: For high-capability targets, monitor per-benchmark regression (e.g., Hi-ToM −0.04 average) and selectively apply scaffolding only where it improves, avoiding harmful over-engineering on near-saturated tasks.
-
Incorporate Recursion-Depth-Aware Handling: For nested belief tasks (Hi-ToM), add specialized decomposition that breaks down recursion order >2 into sub-steps, mitigating accuracy decline from 0.999 (order 0) to 0.700 (order 4).
-
Add Self-Scaffolding Capability: Enable weaker models to build their own harnesses using validation feedback, yielding +0.17 to +0.22 improvement—useful when stronger builders are unavailable or costly.
What the Improved AI System Can Do:
-
Achieve up to 86.7% relative accuracy improvement on Theory-of-Mind tasks (0.488 → 0.912) using a weaker target model with no retraining.
-
Match or exceed human-designed harness performance on simple belief tasks (1.00 vs. 0.95) and approach it on complex social reasoning (0.88 vs. 0.96).
-
Generalize across diverse benchmarks (binary, multi-choice, nested recursion, Bayesian inference) via automated routing and technique selection.
-
Operate cost-effectively: use cheap target models at inference time while offloading reasoning to a one-time builder effort.
-
Provide stable, reproducible performance (mean std. dev. 0.036 across runs) with efficient validation use (median 5 iterations).
Sources
- Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision
- Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks
- Distilling the Knowledge in a Neural Network
- A Survey of Frontiers in LLM Reasoning: Inference Scaling, Learning to Reason, and Agentic Systems
- DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines
- Decomposed Prompting: A Modular Approach for Solving Complex Tasks
- Code as Agent Harness
- UserHarness: Harnessing User Minds for Stronger Agent Theory-of-Mind
- Position: Theory of Mind Benchmarks are Broken for Large Language Models
- Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- ReAct: Synergizing Reasoning and Acting in Language Models
- Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows
- From Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents
- Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks