Can Bayesian Optimization Efficiently Find a Strong Single Expert in Neural Thickets?

arXiv:2608.10867 · cs.LG · Submitted 2026-08-11 · Read on arXiv

Nigel Bastian Cendra, Abdelhamid Ezzerg, Fernando Julio Cendra, Jeremias Knoblauch, Jakob Zeitler

University College London · Institut Polytechnique de Paris · University of Oxford

cs.LG

Submitted: 2026-08-11

Updated: 2026-08-12

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: The paper investigates whether Bayesian optimization (BO) can efficiently find a strong single expert model through gradient-free post-training of large language models (LLMs), under a modest

Terminology

Summary

The paper investigates whether Bayesian optimization (BO) can efficiently find a strong single expert model through gradient-free post-training of large language models (LLMs), under a modest evaluation budget. The authors propose a method that applies BO within a random linear embedding of weight space, using a Gaussian process (GP) surrogate to guide candidate evaluations. The method requires no backpropagation.

The research question is: Given a fixed and modest evaluation budget and limited selection data, can a structured gradient-free method find a stronger single expert than random sampling?

The method works as follows: Instead of sampling perturbations isotropically in the full weight space (as RandOpt does), the authors construct a random linear embedding A = [e1,..., ed] ∈ R(D×d) of d Gaussian directions, spanning a d-dimensional subspace. A candidate is parameterized by a coefficient vector α ∈ R d and a scale σ ∈ R>0, with the perturbed model given by θ′ = θ0 + σ·√D·(Aα/Aα). The normalization ensures the perturbation has norm σ√D, matching the expected norm of a RandOpt perturbation. The search space is (d+1)-dimensional, with d=64 and σ searched on a log scale. Each candidate is scored by its reward on a selection set D sel (accuracy under the task's verifier). The GP surrogate is fitted to observed scores and proposes the next (α, σ) by maximizing an acquisition function.

Key results from preliminary experiments on Countdown, GSM8k, and MATH500 with Qwen2.5-Instruct models (0.5B, 1.5B, 3B):

  1. Sample efficiency at K=1: With a 5× smaller budget (200 vs 1000 evaluations), BO matches or beats RandOpt-1000 at K=1 in most settings. For example: Countdown-0.5B (9.34 vs. 6.81), Countdown-1.5B (15.05 vs. 14.60), Countdown-3B (19.71 vs. 18.89), GSM8k-1.5B (67.78 vs. 66.13), and MATH500-1.5B (40.43 vs. 39.80). BO is at parity on GSM8k-0.5B (43.92 vs. 43.99). The advantage tracks the headroom available: where the base model is weakest and good perturbations are rare (Countdown-0.5B), BO gains most (+2.53 over a headroom of 6.46); where little is available (GSM8k-0.5B, MATH500), it gains nothing.

  2. MATH500 anomaly: BO attains the highest selection reward at both model sizes (44.55 and 63.75), but test accuracy moves in the opposite direction. At 0.5B it is the lowest of any method (28.63), and at both sizes BO falls below the base model (30.67 and 41.00). The authors note: "Where the neighbourhood contains no genuinely better model, selection reward can still be raised by several points, but those gains are fit to D sel rather than to the task, and the more effective the search, the more it overfits."

  3. Same subspace, different search: Within a fixed basis, BO reaches higher selection reward in all seven settings (by 1.6 to 9.97 points) and higher test accuracy at K=1 in five of seven settings. The gains are attributable to the surrogate rather than dimensionality reduction: at matched budget, random search in the basis does not outperform full-dimensional RandOpt sampling, so the low-dimensional parametrization is a concession to tractability rather than a better place to sample.

  4. Selection reward saturation: Evaluating all 200 candidates from a GSM8k-1.5B run shows that selection, validation, and test accuracy rise together over the lower three quarters of the range, but separate above roughly rank 75 (at a selection reward near 80%). Selection reward continues climbing to 88% while validation and test accuracy stay flat at approximately 79% and 68%. The final eight points of selection reward buy no measurable improvement on either held out set. The plateau reflects the selection set rather than noise, and the growing gap between selection and validation toward the selected extreme indicates selection bias.

  5. Majority voting ensembles: At K=20, BO matches alternatives at matched budget on every setting, but shows a clear deficit against RandOpt-1000 (which uses five times the search budget), suggesting ensemble quality benefits from a broader pool of candidates.

  6. Evidence of a higher ceiling: In one basis and BO seed combination at Qwen2.5-Instruct-3B, BO located a region with top selection rewards of 42.18%, well above the 32.16% reached by the next best run, and its best candidate reached 34.45% test accuracy at K=1 against 21.25% for that run and 19.71% ± 4.24% across all ten. RandOpt never produced a candidate of this quality. However, this region was not recovered in later runs, indicating run-to-run variance is dominated by which region a run settles in.

The paper also discusses planned experiments: (i) scaling to Qwen2.5-Instruct-3B with a larger budget (N=500 rather than 200), and (ii) a selection-data sweep (D sel = 200 to 1,000) with the test set fixed, to test whether the saturation threshold moves with resolution.

The main conclusion is that surrogate-guided search can substantially reduce the evaluation cost of gradient-free post-training while producing stronger deployable single experts, though the gains transfer imperfectly due to selection set overfitting near the top of the reward range.

Improvements for AI systems

Improvements to AI Systems:

  1. Efficient Single-Model Fine-Tuning Without Backpropagation
  • The improved system can perform post-training of LLMs using Bayesian optimization over a low-dimensional random subspace of weights, requiring only forward passes and reward evaluations.

  • It can match or exceed the performance of random perturbation search (RandOpt) with 5× fewer evaluations (e.g., 200 vs. 1000), making it viable for compute-constrained environments or tasks where gradient computation is infeasible (e.g., black-box APIs, proprietary models).

  1. Adaptive Search with Overfitting-Aware Stopping
  • The system can monitor the divergence between selection-set reward and held-out validation/test accuracy during optimization.

  • It can automatically halt or switch to a validation-based criterion once selection reward saturates (e.g., beyond rank 75 in the paper), preventing wasted evaluations on overfit candidates and ensuring the final model generalizes better.

  1. Region-Diversity-Aware Exploration
  • Given the paper’s finding that run-to-run variance is dominated by which weight-space region a run settles in, the system can incorporate multiple independent BO runs with different random bases and seeds.

  • It can then aggregate or select the best candidate across runs, reducing the risk of missing high-quality regions (e.g., the 42.18% selection reward region found only once in the 3B model).

  1. Budget-Aware Ensemble Construction
  • For tasks where a single expert is insufficient, the system can use BO to build a diverse set of top-K candidates (e.g., K=20) from a single run, but with a clear understanding that ensemble quality degrades when the search budget is too small.

  • The improved system can dynamically allocate more budget to ensemble scenarios, or alternatively, use BO to select a smaller, higher-quality ensemble that matches the performance of a larger random-search ensemble at lower cost.

  1. Task-Headroom-Aware Resource Allocation
  • The system can estimate the “headroom” (difference between base model performance and best achievable performance) early in the search.

  • If headroom is small (e.g., GSM8k-0.5B), it can reduce the search budget and avoid overfitting; if headroom is large (e.g., Countdown-0.5B), it can allocate more evaluations to exploit the potential gains, maximizing the return on compute.

  1. Selection-Set Resolution Tuning
  • The system can adjust the size of the selection set (e.g., from 200 to 1,000 examples) based on the observed saturation point.

  • It can detect when selection reward no longer correlates with validation improvement and then increase selection-set size or switch to a more robust verifier, ensuring that the final model’s gains are real and transferable.

  1. Gradient-Free Domain Adaptation
  • The system can be applied to any LLM (including those with frozen weights or inaccessible gradients) for tasks like arithmetic reasoning, math problem solving, or instruction following, where a verifiable reward exists.

  • It can produce a stronger single expert model for deployment, with test accuracy improvements of up to +2.53 points over random search at a fraction of the cost, and in some cases, discover candidates that random search never finds (e.g., 34.45% vs. 19.71% test accuracy on Countdown-3B).

Abstract

Gradient-free post-training has emerged as a compelling alternative to gradient-based optimization for large language models (LLMs), but existing approaches remain costly. We ask whether structured search can identify a strong single expert under a modest evaluation budget. Motivated by evidence that useful weight updates lie in low-dimensional subspaces, we apply Bayesian optimization within a random linear embedding of weight space. Our method requires no backpropagation and uses a Gaussian process surrogate to guide candidate evaluations efficiently. Across several reasoning benchmarks with Qwen2.5-Instruct models from 0.5B to 3B parameters, Bayesian optimization using five times less candidate evaluations matches or exceeds RandOpt. These results show that surrogate-guided search can substantially reduce the evaluation cost of gradient-free post-training while producing stronger deployable single experts.

Sources

Related papers