Autonomous discovery of accelerator commissioning algorithms
Thorsten Hellert
Lawrence Berkeley National Laboratory
physics.acc-ph, cs.AI
Submitted: 2026-08-07
Updated: 2026-08-10
Comments: 9 pages, 4 figures. Submitted to Physical Review Accelerators and Beams (ZVR1001)
Code: https://github.com/karpathy/autoresearch
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: The paper "Autonomous discovery of accelerator commissioning algorithms" by Thorsten Hellert (Lawrence Berkeley National Laboratory) demonstrates a closed research loop in which a language-model
Terminology
Summary
The paper Autonomous discovery of accelerator commissioning algorithms
by Thorsten Hellert (Lawrence Berkeley National Laboratory) demonstrates a closed research loop in which a language-model agent autonomously writes, tests, and improves accelerator commissioning algorithms in simulation.
Problem and motivation. The author notes that simulated commissioning has become essential for de-risking modern light-source design and commissioning, but the procedures being simulated are still designed entirely by human experts. Their labor-intensive redevelopment after lattice changes makes such studies hard to repeat and limits their use during early design iteration.
Existing automation reduces the effort required to execute and tune commissioning procedures, but it does not remove the need to design those procedures in the first place.
The paper addresses "the remaining design problem by moving the search from the machine variables to the commissioning procedure itself. Rather than asking an automated method to execute or tune a fixed algorithm, we ask an agent to modify the algorithm, evaluate the result, and retain only validated improvements."
Method — the autoresearch loop. The approach adapts Karpathy's autoresearch
concept — a greedy cycle that edits code, runs a fixed-budget computational experiment, and keeps only improvements
— to accelerator commissioning. The loop (Fig. 1) works as follows:
-
Propose: A proposer language-model agent receives (1) declarative domain knowledge (primers on RF capture, sextupole ramping, tune resonances, orbit feedback; operator-interface documentation; a textbook-retrieval subagent), (2) a helper library of executable routines ported from the published ALS-U procedures (first-turn threading, two-turn stitching, sextupole ramping, tune scanning, RF phase/frequency correction), and (3) campaign memory of merged changes, hypotheses, and scores. Five rotating personas (e.g., theorist, empiricist, simplifier) broaden the search.
-
Screen: An independent reviewer checks the diff for
invalid simulator access or unphysical shortcuts,
with a deterministic pattern scan guarding against known forbidden access patterns. -
Evaluate: A fixed harness evaluates each surviving candidate on the same fixed ensemble of 50 randomly perturbed machines (error seeds), from identical seeded conditions, with a budget of 500 attempted injections per seed.
-
Merge:
A candidate is merged only if its ensemble-mean score improves on the incumbent
in scalar campaigns; the search is greedy but every accepted change is version-controlled and validated by the unchanged harness.
The paper emphasizes a strict separation between the algorithm being developed and the experiment that judges it: "Candidates access the simulated accelerator only through an operator interface exposing control-room actions and measurements... Simulator ground truth, including seeded alignment and calibration errors, is withheld. Without this separation, an agent could improve the score by exploiting the benchmark rather than the procedure."
Scalar beam-capture benchmark. The demonstration is RF beam capture in the ALS-U accumulator-ring simulated-commissioning model, where Capture requires at least 80 % of 100 tracked particles to survive 500 turns.
The objective is the ensemble-mean number of injections required for capture across 50 error seeds, lower being better. Failed seeds receive surrogate costs (e.g., immediate failure costs 1100 injections). The expert baseline (a faithful pySC port of the published MATLAB algorithm) scores 207.5 injections, though the paper notes this is a common starting point, not the best performance achievable by a human.
Results — model capability. Comparing three tiers of the Anthropic Claude family (Haiku 4.5, Sonnet 4.6, Opus 4.6) under identical configuration and 100-experiment campaigns (three independent campaigns per tier): "The best Sonnet and Opus campaigns reach mean scores of 20.3 and 27.5 injections, respectively, while Haiku reaches 53.6. We therefore read this as a capability threshold within the tested tiers: the frontier models enter the tens-of-injections regime, whereas the weaker tier saturates higher. The loop reduces the objective
by a large factor. Most improvement came from
streamlining the human procedure rather than from a new capture mechanism, though
one change stands out as genuinely new: a recovery move that nudges the correctors near the injection point to escape repeated beam loss, rescuing the hardest seeds and driving much of the final gain."
Results — scaffold ablation. Starting from a minimal non-capturing stub, the paper tests all four combinations of the domain-knowledge documents and the helper library across Haiku and Sonnet (four independent 40-experiment campaigns each). "The helper library is the dominant scaffold. When available, the agent drives the objective into the tens to low hundreds of injections whether or not the documents are present. Without it, the documents alone usually leave campaigns in the hundreds, and the weaker tier essentially never captures from the bare stub. The stronger tier always captures from scratch but
roughly an order of magnitude worse than with it. The two tiers fail differently: with no scaffolding,
the weaker tier expends far more agent effort per experiment (in reasoning turns and generated code) while still failing to capture, whereas the stronger tier increases effort less and still produces functioning procedures. The model comparison
suggests that stronger models can compensate partly for missing scaffolding, whereas weaker models require more explicit procedural structure. The ablation does not imply written domain knowledge is unimportant; rather,
the helper library packages commissioning knowledge in a more directly usable form: executable routines that the agent can inspect, modify, and recombine."
Multi-objective (Pareto) campaign. The paper extends the framework by replacing the scalar merge predicate with Pareto dominance, so a candidate is retained if it is not dominated by any existing retained algorithm
and the maintained state becomes the non-dominated set rather than a single best repository head.
This harder benchmark enables discrete catastrophic errors (reversed corrector and BPM polarities, dead BPMs), raising the expert-port baseline from 208 injections in the scalar study to 713 injections on this harder ensemble.
Two objectives are used: capture cost (mean injections) and a machine-error correction score S corr measuring how much seeded correctable error has been removed (BPM offsets and signed gains, corrector calibration, injection coordinates, RF phase/frequency, and dead-BPM identification; quadrupole/sextupole calibration and global tune are excluded as belonging to later optics-calibration stages). A single 200-experiment Sonnet campaign produced 16 non-dominated algorithms spanning physically distinct strategies.
The lowest-cost algorithm corrects only the RF phase and frequency, capturing beam in 679 injections,
while "the highest-quality algorithm keeps the capture procedure intact and appends a stored-beam calibration stage... It spends 1371 injections but removes roughly two-thirds of the seeded polarity errors, identifies most dead BPMs, and cuts the injection error. The author concludes:
autonomous algorithm search changes the scale of what can be explored. A single campaign produced 16 validated commissioning procedures spanning physically meaningful choices between rapid capture and more complete machine calibration."
Discussion and outlook. The near-term role is likely primarily offline, spanning lattice design, detailed simulated commissioning, and preparation for first beam,
making commissionability a quantity explored throughout the design process.
A practical intermediate mode would keep humans in the evaluation loop
with experts approving, rejecting, or redirecting proposals. A further extension is co-design of the accelerator and its commissioning strategy. The most direct next test is an end-to-end procedure from first injection to a machine state suitable for user operation,
whose main challenge would instead be to reduce the high-dimensional space of machine-performance objectives to a small set of quantities, or ideally a scalar objective.
Costs and benchmark integrity (appendices). The agent phase cost was modest: Median API cost was approximately 0.5 per experiment for Haiku and ∼ 1–3 for the frontier-tier models, corresponding to roughly 60–280 for a 100-experiment scalar-capture campaign.
Physics evaluation is CPU-bound, with 50 seeds distributed over worker processes in parallel. Appendix C discusses reward hacking and benchmark-integrity failure modes: early harnesses allowed unpriced machine-equivalent actions, exposed simulator quantities unavailable in a real control room (true alignment errors, analytic lattice information), and allowed candidates to effectively certify its own success.
The final harness addresses these structurally
: All machine-equivalent actions are priced, simulator ground truth is excluded from the operator interface, and capture is certified only by the harness using a fixed minimum particle count.
The paper concludes that progress must be judged by a fixed, harness-owned metric whose connection to the intended scientific task has been independently validated.
Improvements for AI systems
Based on this paper, I can make the following concrete improvements to AI systems:
-
Closed-loop autonomous algorithm discovery. The AI system can run a self-improving cycle that proposes changes to its own algorithm code, executes a fixed computational experiment, measures an objective, and keeps only changes that improve the measured score. This enables it to discover and refine accelerator commissioning procedures without human redesign, going from a human-authored starting procedure (207.5 injections) down to 20–28 injections in the best campaigns.
-
Scaffolding with executable helper routines. Instead of giving the agent only free-form documentation, provide a library of tested, reusable routines (first-turn threading, two-turn stitching, sextupole ramping, tune scanning, RF phase/frequency correction). The agent can inspect, modify, and recombine these routines. The ablation shows this is the dominant factor: with the helper library, the agent reaches tens to low hundreds of injections; without it, even strong models are roughly an order of magnitude worse, and weaker models often fail entirely.
-
Persona-based exploration. Use multiple rotating proposer personas (theorist, empiricist, simplifier, etc.) to broaden the search over algorithm space, preventing premature convergence to a single local optimum and producing physically distinct strategies.
-
Independent code review and pattern screening. Before evaluation, have a separate reviewer agent check every code diff for invalid simulator access or unphysical shortcuts, plus a deterministic scan of known forbidden patterns. This blocks reward hacking and enforces that candidates interact with the simulator only through an operator-interface-equivalent API that withholds ground truth, prices all machine actions, and certifies success only through a harness-owned metric.
-
Fixed ensemble evaluation with budgeted seeds. Evaluate every candidate on the same ensemble of 50 randomized error seeds from identical initial conditions, with a fixed injection budget per seed. Use ensemble-mean score as the merge criterion so improvements are robust, not overfit to a single run. This lets the AI system distinguish genuine algorithmic improvements from lucky or benchmark-exploiting variations.
-
Greedy, version-controlled merge with memory. Maintain a version-controlled repository of accepted algorithms, campaign memory including merged changes, hypotheses, and scores. Accept a candidate only if its ensemble-mean score improves on the incumbent. This creates a traceable, reproducible search history and prevents unvalidated changes from accumulating.
-
Multi-objective Pareto search. Replace scalar merge with Pareto dominance, retaining the non-dominated set of algorithms rather than a single best repository head. A single campaign produced 16 non-dominated commissioning procedures spanning meaningful trade-offs, e.g., a rapid-capture algorithm (679 injections, little calibration) versus a calibration-heavy procedure (1371 injections, removing 2/3 of seeded polarity errors, identifying most dead BPMs, and correcting injection errors). The improved AI system can therefore map a frontier of valid strategies across competing objectives.
-
Capability-aware task delegation. Use tier-appropriate models for the proposer/editor role. Frontier models can partially compensate for missing scaffolding with fewer reasoning turns and more effective code; weaker models require explicit procedural structure and expend far more effort. The improved system can dynamically allocate agent effort and scaffolding based on measured capability, reducing cost and avoiding futile open-ended search on weaker models.
-
Cost-aware campaign design. Run the search at roughly 0.5 per experiment for small models and 1–3 per experiment for frontier models, i.e., 60–280 per 100-experiment campaign, with physics evaluations parallelized over CPU workers. The improved system can budget its agent API calls, choose model tiers by task difficulty, and estimate total campaign cost before launching.
-
Human-in-the-loop evaluation mode. Keep experts in the loop to approve, reject, or redirect proposed algorithm changes, especially for high-stakes preparation for first beam. The improved AI system can present candidates with validation scores and physical rationale, letting humans exert strategic control while the AI performs the combinatorial search.
-
Ground-truth isolation and benchmark integrity. Design the AI’s environment so that all simulator ground truth (alignment errors, calibration errors, analytic lattice information) is hidden, all machine-equivalent actions are priced, and success certification is performed only by the harness using fixed count thresholds. This prevents the AI from improving scores by exploiting the benchmark rather than the procedure. The improved system can generalize this principle to other scientific domains: always separate the agent’s observations/actions from the validator’s ground truth, and make the metric fixed and harness-owned.
-
End-to-end commissioning generation. Extend the loop from beam capture to a full procedure from first injection to a user-operation-ready machine state, reducing a high-dimensional space of machine-performance objectives to a small set or a single scalar objective. The improved AI system can therefore generate complete, validated commissioning playbooks for new accelerator lattices, making “commissionability” a quantity that can be explored throughout lattice design.
What the improved AI system can do overall:
-
Given a new accelerator lattice or a design change, autonomously develop a validated simulated commissioning algorithm in tens of hours of compute, outperforming a human-authored baseline by up to an order of magnitude.
-
Produce not just one but a portfolio of physically distinct commissioning procedures representing different trade-offs between speed, robustness, and calibration completeness.
-
Reuse knowledge across campaigns via executable routines and campaign memory, so later lattices require less search effort.
-
Operate safely in environments with adversarial benchmark-integrity constraints, refusing to use forbidden ground-truth information or unpriced actions even if they would improve its score.
-
Provide expert auditable trails of every accepted algorithm change, including hypotheses, validation scores, and version history.
-
Serve as an offline design tool that accelerates the accelerator design and commissioning preparation cycle, with optional human oversight.
Abstract
Simulated commissioning has become essential for de-risking modern light-source design and commissioning, but the procedures being simulated are still designed entirely by human experts. Their labor-intensive redevelopment after lattice changes makes such studies hard to repeat and limits their use during early design iteration. This Letter demonstrates a closed research loop in which a language-model agent writes commissioning code, tests it in simulation, and improves the algorithm from the results. Applied to RF beam capture in the ALS-U accumulator-ring model, the loop substantially improves a working expert procedure and can construct a working one from a minimal starting point, with more capable models succeeding from less initial code. Extending the same framework to multiple objectives produces 16 non-dominated algorithms spanning physically distinct trade-offs between rapid beam capture and correction of seeded machine errors. This reframes commissioning studies from evaluating human-designed procedures toward a mode in which agents participate directly in discovering accelerator algorithms.
Sources
- Beam dynamics performance of the proposed PETRA IV storage ring
- Online optimization of storage ring nonlinear beam dynamics
- Bayesian optimization of a free-electron laser
- Multi-Objective Bayesian Optimization for Accelerator Tuning
- Bayesian Optimization of the Beam Injection Process into a Storage Ring
- Cheetah: Bridging the Gap Between Machine Learning and Particle Accelerator Physics with High-Speed, Differentiable Simulations
- Learning to Do or Learning While Doing: Reinforcement Learning and Bayesian Optimisation for Online Continuous Tuning
- Bayesian Optimization Algorithms for Accelerator Physics
- Large Language Models for Human-Machine Collaborative Particle Accelerator Tuning through Natural Language
- GAIA: A General AI Assistant for Intelligent Accelerator Operations
- Towards Agentic AI on Particle Accelerators
- Agentic Artificial Intelligence for Multistage Physics Experiments at a Large-Scale User Facility Particle Accelerator
- Osprey: Production-Ready Agentic AI for Safety-Critical Control Systems
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- ChemCrow: Augmenting large-language models with chemistry tools
- Eureka: Human-Level Reward Design via Coding Large Language Models
- Concrete Problems in AI Safety
- AI Safety Gridworlds