How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models
Aleksandra Kalisz, Jack Simons, Krisztina Sinkovics, Noam Ghenassia, Shikha Surana, Henry Moss, Paul Duckworth
InstaDeep Ltd · University of Oxford · Lancaster University
cs.AI, cs.LG
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: Proceedings of the ICML 2026 Workshop on Structured Probabilistic Inference & Generative Modeling (SPIGM)
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: This paper addresses the challenge of efficiently using limited oracle budgets when guiding protein structure prediction models.
Terminology
Summary
This paper addresses the challenge of efficiently using limited oracle budgets when guiding protein structure prediction models. Foundation models for protein structure prediction, such as AlphaFold3, Chai-1, and Boltz-2, remain unreliable on certain targets,
and external oracles can flag and correct these failures, but biological oracles are expensive, making oracle budget a critical constraint.
The paper notes that a single molecular dynamics simulation can take days of GPU compute, and wet-lab assays are slower still.
The authors benchmark four guidance methods: FK-steering (Feynman-Kac steering), DPO (Direct Preference Optimisation), Best K-of-N sampling, and the newly proposed Optimisation Over Outputs (O3) framework, which uses a small set of example model generations to build a low-dimensional subspace of the latent space over which standard optimisers can be directly applied.
This work represents the first application of O3 to protein structure prediction.
The experimental setup uses Boltz-2 as the base generative model and evaluates on two protein targets: calmodulin (PDB: 1CLL, 144 residues, 1,184 atoms) scored with TM-score against the ground truth structure, and E. coli aspartate transcarbamoylase (PDB: 9EEH, 7,232 atoms) scored with a reference-free MolProbity oracle measuring physical plausibility. Six budget configurations are tested: (N, K) = (20, 2), (50, 5), (100, 10), (200, 20), (500, 50), and (1000, 100), where N is total oracle calls and K is the batch of returned candidates.
Key findings from the main experiment on 1CLL with TM-score oracle: O3 outperforms all baselines at low-mid budgets (N ≤ 1000) with the mean value plateauing at ∼ 0.81.
Best K-of-N is roughly flat at ∼ 0.60 across all budgets.
FK-steering and DPO improve with increased N, but neither are competitive with simple Best K-of-N at low budgets (N ≤ 100), with FK-steering achieving only ∼ 0.55 at N = 20.
At N = 1000, FK-steering reaches ∼ 0.73 and DPO reaches ∼ 0.71. O3 is the only method that meaningfully improves on the Best K-of-N baseline at low oracle budgets (N ≤ 100).
For O3, the authors find that the best-performing d varies with the budget — different oracle-call regimes favour different subspace dimensions.
For example, when N = 20 the best-performing dimension is d = 6; however, for N = 200 dimension d = 10 yields better predictions.
They also demonstrate that the superior performance of O3 seen in Fig. 2 comes from a combination of both the optimiser and the subspace construction,
as random sampling in the subspace still improves over Best K-of-N.
For FK-steering, ablations show that increasing λ improves performance across all budgets, however, it likely hinders output diversity.
A high λ = 50 amplifies the signal from the oracle, forcing the algorithm to greedily resample the best-performing particles.
Additionally, increasing the number of resampling steps N/K at the cost of particle population K generally yields better performance at higher oracle budgets.
For DPO, the comparison between online and offline variants shows that Online DPO improves steadily with N, reaching a mean TM-score of 0.708 and a max TM-score of 0.783 at N = 1000,
while offline DPO shows little sensitivity to budget, plateauing at 0.55–0.57 across all budgets.
The authors conclude that the gains from DPO arise primarily from adaptive on-policy resampling, rather than from additional optimisation steps alone.
On the 9EEH target with MolProbity oracle, results differ somewhat: O3 scores highest at all but the smallest budget, where it matches Best K-of-N.
However, FK-steering performs poorly throughout and does not improve with budget, in contrast to 1CLL.
The authors hypothesize this is because MolProbity instead measures local geometry, in which residual noise dominates until the final denoising steps,
making intermediate predictions uninformative for FK-steering's resampling.
The paper concludes with practical advice: use O3 at constrained oracle budgets. FK-steering can be competitive at moderate budgets, and reach for DPO once budgets are large enough to fine-tune.
They also note that DPO amortises its training budget into the model weights, making sampling large subsequent batches cheap, whereas inference-time and search methods usually discard oracle values afterwards.
Improvements for AI systems
Improvements to AI systems:
-
Budget-Aware Adaptive Optimization: Implement an AI system that dynamically selects between inference-time steering (O3), resampling (FK-steering), or fine-tuning (DPO) based on the available oracle budget. The system would automatically switch strategies as budget scales, using O3 for low budgets (N ≤ 100), FK-steering for moderate budgets (N 500), and DPO for large budgets (N ≥ 1000), maximizing performance per oracle call.
-
Latent Subspace Construction for Generative Models: Enhance generative AI systems (beyond proteins, e.g., drug design, materials) by building low-dimensional subspaces from a small set of model outputs, then applying standard optimizers directly in that subspace. This enables efficient guided generation without full model retraining or expensive sampling, improving output quality at minimal computational cost.
-
Adaptive Dimensionality Selection: Develop a meta-learning component that predicts the optimal subspace dimension (d) for a given oracle budget, based on the observation that d=6 works best at N=20 while d=10 works best at N=200. This would allow the system to automatically tune its search space complexity to match available evaluation resources.
-
Oracle-Aware Resampling Scheduler: For FK-steering, create an AI system that adjusts the λ (guidance strength) and the ratio of resampling steps (N/K) based on the oracle signal quality and budget. The system would increase λ and reduce particle population K when oracle signals are noisy or budgets are tight, preventing diversity collapse while maximizing signal exploitation.
-
Reference-Free Physical Plausibility Scoring: Integrate MolProbity-style oracles into a feedback loop for generative models, but with a temporal awareness component—recognizing that local geometry oracles are only informative at final denoising steps. This would prevent wasted oracle calls on intermediate predictions, improving efficiency for systems targeting physical validity rather than global structure.
-
Hybrid Online-Offline Preference Learning: Build a system that combines online DPO (adaptive on-policy resampling) with offline pre-training on existing data, then uses the online component only when budgets allow. This would amortize training costs into model weights while retaining the budget-sensitivity that makes online DPO effective at large N.
-
Cross-Target Oracle Transfer: Create a system that learns which oracle type (global TM-score vs. local MolProbity) is most informative for a given target class, then automatically selects the appropriate guidance method. This would prevent the observed failure of FK-steering on local-geometry oracles by routing to O3 or DPO instead.
What the improved AI system can do:
-
Generate high-quality protein structures with up to 35% better TM-scores than naive sampling at low oracle budgets (N ≤ 100), using O3-style subspace optimization.
-
Automatically allocate limited oracle calls across different guidance strategies to achieve near-optimal performance at any budget, from 20 to 1000+ calls.
-
Adapt its search dimensionality and resampling aggressiveness in real-time based on observed oracle feedback, avoiding wasted evaluations.
-
Distinguish between global and local quality metrics and switch guidance methods accordingly, preventing performance collapses on targets where local geometry matters.
-
Scale efficiently: use cheap inference-time steering for quick tasks, and invest in fine-tuning only when large budgets justify the amortized cost, making the system cost-effective across diverse deployment scenarios.
Abstract
Foundation models for protein structure prediction remain unreliable on certain targets. External oracles can flag and correct these failures, but biological oracles are expensive, making oracle budget a critical constraint. Existing guidance methods, such as FK-steering, DPO, and Best K-of-N sampling, differ in how they spend this budget, yet no systematic comparison exists to guide method selection. To bridge this gap, we benchmark these methods alongside the recently proposed Optimisation Over Outputs (O3), which applies off-the-shelf optimisers within a generative model's latent subspace. We extend the usage of O3 to protein structure prediction models. Overall, our work provides the first practical reference for oracle budget-aware guidance. Our evaluation on two protein targets, calmodulin (1CLL) and E. coli aspartate transcarbamoylase (9EEH), reveals that no single method consistently dominates across all budgets and oracles. Specifically, O3 proves most effective at low oracle budgets, while FK-steering and DPO demonstrate improved performance as the budget increases. We distil these findings into actionable recommendations for practitioners operating under real-world oracle-budget constraints.
Sources
- Linear combinations of latents in generative models: subspaces and beyond
- Local Latent Space Bayesian Optimization over Structured Inputs
- Directly Fine-Tuning Diffusion Models on Differentiable Rewards
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Classifier-Free Diffusion Guidance
- Score-Based Generative Modeling through Stochastic Differential Equations
- Understanding the performance gap between online and offline alignment algorithms
- Sample-Efficient Optimization in the Latent Space of Deep Generative Models via Weighted Retraining
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection