FUSION: a skill-based research agent for publicly obtainable nuclear-physics codes

arXiv:2609.04742 · nucl-th, cs.LG, physics.comp-ph · Submitted 2026-09-04 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "FUSION: a skill-based research agent for publicly obtainable nuclear-physics codes".

Jane: A skill-based research agent, FUSION, addresses the difficulty of running and interpreting unfamiliar nuclear-physics codes by embedding code-specific skills that ensure correct input conventions and validate outputs against benchmarks.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, we’re talking about this paper called "FUSION: a skill-based research agent for publicly obtainable nuclear-physics codes." It sounds like they’ve built something specifically to handle the headache of running these complex physics programs.

Jane: It seems like the core idea is that just having the right code isn't enough; you still have to know how to set it up correctly and interpret what it spits out, which this agent aims to solve.

Lu: The focus on "skill-based" suggests they aren't just trying to be a general coding assistant; they are building specialized tools that understand the specific conventions of nuclear physics software.

Meng: I wonder how deep this specialization goes; is it just about following instructions, or does it actually understand the underlying physics constraints?

Lalam: This paper introduces a layered architecture, which suggests a very thoughtful approach to keeping sensitive user information separate from the public code execution environment.

The paper's summary: Tom: What’s really interesting about this is how FUSION tackles the problem of getting accurate results when using unfamiliar nuclear physics codes by embedding code-specific skills that handle everything from setup to validation.

Jane: They emphasize that a general assistant can write syntactically correct input, but it often fails because it uses the wrong physical conventions, leading to results that look plausible but are fundamentally wrong.

Lu: The paper describes how a single skill handles fetching the code from its public source, starting with verified inputs, running the calculation in a clean space, parsing the output according to specific rules, and crucially, checking that result against a benchmark before reporting anything.

Meng: That systematic approach of starting from known good inputs instead of writing everything from scratch sounds like a massive practical help for researchers who are just trying to get their experiments moving quickly.

Lalam: They detail the three layers: an engine, the nuclear-physics layer with skills and knowledge bases, and a private layer for user materials, which really shows they’re thinking about security from the start.

The paper's improvements: Tom: One of the key improvements FUSION proposes is this rigorous verification process built into every skill. They mandate that each skill must reproduce a stated benchmark to a specific tolerance before it can report a result, which is much stronger than just running the code.

Jane: They also focus heavily on documentation for known failure modes, which means the agent doesn't just fail silently; it’s supposed to record practical issues like exit codes or locale problems so you know what might go wrong.

Lu: The verification loop they suggest is quite thorough, involving building the code on different operating systems and having a second language model check for defects written into the skill itself.

Meng: From an engineering standpoint, this kind of self-checking mechanism addresses the exact issue of silent errors that plague complex software execution environments.

Lalam: This emphasis on reproducing results with stated tolerances provides a very high bar for soundness, moving beyond simple syntactic correctness to actual physical accuracy verification.

Conclusion: Tom: So, to wrap things up on "FUSION: a skill-based research agent for publicly obtainable nuclear-physics codes," the big idea is creating these highly specialized tools that ensure accuracy through rigorous skill execution and benchmark checks.

Jane: It really boils down to moving away from general agents that might make convention errors and toward systems that are explicitly designed to handle the unique idiosyncrasies of specific physics software.

Lu: The architecture, separating the engine from the nuclear-physics layer and having dedicated skills for twenty different codes, shows a very deliberate plan for comprehensive coverage within a structured framework.

Meng: I think what this paper offers is a tool that actually gets researchers past the initial setup hurdle and into meaningful calculation much faster because it handles those tedious checks automatically.

Lalam: This work sets a precedent for how specialized AI agents can be built to handle highly technical, domain-specific tasks while maintaining strict separation for user data, which is really important for future research tools.

School of Physics Science and Engineering, Tongji University · Southern Center for Nuclear-Science Theory (SCNT), Institute of Modern Physics, Chinese Academy of Sciences

nucl-th, cs.LG, physics.comp-ph

Submitted: 2026-09-04

Updated: 2026-09-04

Comments: 11 pages, 10 figures. Platform at https://github.com/jinleiphys/FUSION (MIT), documentation at https://vibeinscience.com

Code: https://github.com/jinleiphys/FUSION

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 91/100

The gist: A skill-based research agent, FUSION, addresses the difficulty of running and interpreting unfamiliar nuclear-physics codes by embedding code-specific skills that ensure correct input conventions and

Key concepts

Skill Mechanism
A skill is a self-contained set of instructions for a specific code. It includes steps for installation, using verified inputs to avoid errors, execution wrappers, output parsing rules, and documented failure modes. Crucially, each skill includes a benchmark to reproduce known results within specified tolerances.
Verification Process
Skills undergo rigorous testing before release using mechanical checks and adversarial review loops. This involves building the code on different operating systems (like macOS ARM and Linux x86-64) and ensuring that guard conditions are set up so they fail exactly as expected, confirming the skill's reliability.
Benchmark Certification
A benchmark confirms that a compiled code build reproduces a known reference result. The level of certification varies; Tier 1 requires byte-for-byte matches or agreement to six significant figures for some codes, while Tier 2 involves reproducing invariants or analytic solutions.
Separation Layers
FUSION uses three distinct layers: an engine (the core agent), a nuclear-physics layer (holding the code skills and literature), and a private user layer. This architecture ensures that public software remains isolated from private user data, preventing exposure during cloning or operation.

Terminology

Summary

A skill-based research agent, FUSION, addresses the difficulty of running and interpreting unfamiliar nuclear-physics codes by embedding code-specific skills that ensure correct input conventions and validate outputs against benchmarks. This system mitigates common failure modes where general-purpose agents produce syntactically valid but physically incorrect results, making it a valuable tool for researchers working with complex computational physics software.

The Gist

FUSION addresses the problem of running unfamiliar nuclear-physics codes by embedding code-specific skills that obtain the code from its public source, start from a verified input, run and parse the calculation, record known failure modes, and must reproduce a stated benchmark to a stated tolerance before reporting a result.

FUSION Architecture

FUSION is structured in three layers designed to keep public software separate from user material. The engine is a fork of the open-source opencode agent, kept by policy as a brand-only patch and rebased weekly, connecting to any language model reachable via API. The nuclear-physics layer contains 26 skills operating individual physics codes and a literature collection. A private layer holds the user’s literature notes, research profile, and credentials, which FUSION defines where they may be attached but does not create or distribute. This separation ensures that the public repository works without access to private user data, and cloning cannot expose it.

The Skill Mechanism

A skill is defined as a directory of instructions, examples, and small scripts that contains no language model itself. When a request names a supported code or calculation, the agent reads the corresponding skill before acting. Each skill contains six components:

  1. Installation from upstream: fetching the code at a pinned version and building it.

  2. Input from verified examples: carrying working inputs that reproduce known results to start from the closest one, defending against syntactically valid nonsense.

  3. Execution: a wrapper runs the code in a clean directory, capturing output and enforcing conventions.

  4. Output parsing: determining which file holds which observable, units, and sign convention.

  5. Failure modes: a written record of how the code fails, containing practical knowledge like exit code issues or locale-dependent makefile sources.

  6. A benchmark with a stated tolerance: reproducing a known result and reporting agreement in significant figures.

Verification Process

Skills are rigorously verified before being released, using mechanical checks and adversarial review loops. The verification process involves:

  1. Building the code on both macOS on ARM and Linux on x86-64 for most skills, with exceptions noted.

  2. Reproducing a benchmark and stating the agreement in significant figures.

  3. Disabling each guard to confirm that exactly its own test fails.

  4. A second language model, different from the one that drafted the skill, runs the code to find defects written into the skill and shipped with it.

Coverage and Knowledge Base

FUSION covers twenty codes spanning various nuclear physics areas, including optical models (e.g., FRESCO), nuclear structure (e.g., GSM), fission and statistical models (e.g., TALYS), astrophysics (e.g., AZURE2), and heavy-ion transport (e.g., SMASH). Coverage follows an access rule, not a ranking, meaning a code is included only if its source can be obtained without registration, can be built on a supported platform, and is described in a published paper. FUSION also ships a searchable literature collection containing 61,059 arXiv nucl-th papers and 703,430 citation links. This collection allows the agent to read all of it with ordinary text search guided by the kb-search skill.

Benchmark Certification

A benchmark certifies that a build reproduces a known result and does not certify that your calculation is right. The evidence for benchmarks differs by code:

: tier 1 when the code itself documents a reference value and the skill reproduces it to a stated precision (fourteen skills). This includes matches byte for byte or agreement to about six significant figures.

Tier 2 involves invariants, analytic solutions, measured data, or cross-platform identity. For instance, CGMF reproduces the 252Cf spontaneous-fission histories bit for bit, while AZURE2 reproduces the R-matrix output with a specific ratio against data. The skill records the agreement stated in significant figures and ships with verification documentation like commit hash or archive checksum.

Real Run Example

A recorded session demonstrated FUSION's capability by requesting a calculation (compute n+90Zr elastic scattering at 50 MeV) and comparing the result against EXFOR. FUSION caught a silent trap where the input convention differed, causing all radii to be 22% too large, but it reported this discrepancy.

Improvements for AI systems

Here are specific improvements to AI systems derived from the principles outlined in the FUSION paper, and what those improved systems can achieve:


The core improvement is shifting general-purpose language models (LLMs) from being mere code/text generators to becoming specialized, verifiable scientific agents capable of executing complex, domain-specific computational tasks with rigorous self-verification.

Here are the specific improvements and capabilities:


  1. Automated Scientific Code Verification and Silent Failure Detection:

A general LLM agent can no longer be trusted to produce correct numerical results simply because the syntax is valid (the syntactically valid nonsense failure mode).

  • The improved system (FUSION skill) must incorporate a mandatory numerical check against a stated tolerance before reporting any result.

  • It must enforce knowledge of physical conventions specific to the code being run (e.g., radius conventions, sign conventions).

  • This allows the AI to catch subtle, physically incorrect errors in input parameters or code interpretation that lead to plausible but wrong outputs (like the 22% radius convention error described in Section I).


  1. Verified Input Generation from Verified Examples (Skill-Based Learning):

Instead of writing an input file from memory, the AI must operate on a principle of starting from a verified, known-good example deck.

  • The system is improved to use verified examples as its primary starting point, instructing the agent to modify these existing inputs rather than inventing new ones.

  • This drastically reduces the risk of introducing fundamental errors related to input structure or convention that are common in domain-specific programming.


  1. Systemic Failure Mode Documentation and Knowledge Extraction:

The AI must be equipped with explicit knowledge of how specific physics codes fail, rather than just running them blindly.

  • Each skill includes a dedicated failure modes record detailing known pitfalls (e.g., exit code ambiguity in TALYS, locale-dependent makefile globbing).

  • The improved system can proactively identify potential failure modes during execution or input drafting, allowing it to avoid known traps.


  1. Adversarial Self-Review and Iterative Debugging:

The verification process must be cyclical and rigorous, involving checks against the code's own test scripts.

  • The system should perform an adversarial review loop where a second language model attempts to run the code, checking guard conditions sequentially until it confirms its own test fails precisely as expected.

  • This iterative process ensures that fixes made in one round do not introduce new defects of the same class, leading to high confidence in the skill's reliability.


  1. Layered Architectural Separation (Decoupling Private State):

The AI system architecture must strictly separate its functional logic from sensitive user data (literature notes, credentials).

  • The system should define a clear private layer where user materials reside, but the core computational engine and skills operate independently of this layer.

  • This ensures that cloning or running the public repository does not inadvertently expose private research data, maintaining security and isolation.


  1. Domain-Specific Knowledge Base Integration (Contextual Retrieval):

The system must utilize a highly structured, searchable knowledge base derived from scientific literature (61k papers).

  • The AI can perform deep contextual queries by reading the digest and citation edges of papers to synthesize answers based on how related concepts are cited, contrasted, or applied.

  • This allows the AI to move beyond simple keyword matching to understanding complex relationships within the nuclear physics literature.


  1. Rigorous Benchmark Reproduction for Soundness:

The system must be able to reproduce known results with verifiable precision across different computing platforms (e.g., macOS/ARM vs. Linux/x86-64).

  • A skill is only considered sound if it reproduces a stated benchmark within the required significant figures, and this agreement must be explicitly reported alongside the result.

  • This provides a measure of installation soundness (the code runs correctly) independent of whether the specific physical model is correct for the user's problem.

The improved AI system can perform:

  • Compute complex nuclear physics calculations (e.g., elastic scattering cross sections, fission histories) with a high degree of confidence in the execution integrity.

  • Draft and validate complex input files for specialized scientific codes, ensuring they adhere to specific physical conventions.

  • Act as an automated research assistant capable of synthesizing information from vast, unstructured scientific literature (arXiv papers) into structured knowledge accessible via semantic search.

  • Perform automated quality assurance checks on custom or public computational routines, identifying silent errors and verifying numerical accuracy against known benchmarks before the user ever needs to perform a manual check.

Abstract

Running an unfamiliar nuclear-physics code is rarely difficult because of the physics alone. One must find and build the program, learn its input conventions, and decide whether a plausible output is actually correct. A general-purpose coding agent helps with the first two tasks but may make the last one harder: it can write an input file that runs with the wrong physical convention. FUSION addresses this problem with code-specific skills. A skill obtains the code from its public source, starts from a verified input, runs and parses the calculation, records known failure modes, and must reproduce a stated benchmark to a stated tolerance before reporting a result. The current release covers twenty codes, spanning optical models and reactions, nuclear structure, fission and statistical models, astrophysics and R-matrix analysis, and heavy-ion transport. It also includes an offline, searchable collection of 61 167 pages derived from the nucl-th literature. User notes and credentials remain outside the public repository. FUSION is available under the MIT license at https://github.com/jinleiphys/FUSION; documentation is at https://vibeinscience.com. Here I describe the design, the checks behind the current release, and one complete calculation from input to comparison with measured data.

Sources

Related papers