ChemWorld: Programmable Chemical Worlds for Controlled and Replayable Agent Experimentation
Jiangjie Qiu, Yijun Li, Xiaonan Wang
Tsinghua University
cs.AI, cs.LG
Submitted: 2026-08-11
Updated: 2026-08-12
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 100/100
The gist: ChemWorld is a programmable chemical environment in which reusable process and observation components are compiled into executable worlds.
Terminology
Summary
ChemWorld is a programmable chemical environment in which reusable process and observation components are compiled into executable worlds. ChemWorld separates the public experimental contract available to an agent from evaluator-owned chemical and material laws. Researchers can therefore vary world composition and operating conditions, or change a single hidden law while holding the public task and interaction conditions fixed. Transactional execution records operations, failures, resource changes, and state transitions, allowing complete environment–action trajectories to be replayed exactly and audited. Full-census qualification covered the reference registry, 52 generated compositions, and module, interface, compilation, and invalid-action tests. Eight deterministic experimental cases demonstrated shared lifecycle semantics, failure recovery, and exact replay, while six parent–child world-fork pairs isolated the effects of single private-law interventions under matched public conditions. An independent agent also completed a full lifecycle in a non-reference world through the same public interface. Within the declared component and model domain, ChemWorld provides a controlled and replayable substrate for studying experimentation across systematically varied chemical worlds, complementary to physical-laboratory evidence and calibration.
The paper makes four contributions: (1) Programmable chemical-world construction—reusable process and observation components and a compatibility compiler construct executable worlds whose topology, operating conditions, instruments and private chemical or material laws can be varied systematically. (2) A common experimental interface across worlds—compatible worlds share one agent-facing contract for operations, instruments, observations, resources, failure handling, termination and evaluation. (3) Complete, replayable and attributable experimental records—each operation executes transactionally and records success, failure, rollback and resource changes, providing process-complete evidence for replay and intervention attribution. Complete environment–action traces can be reconstructed exactly, while single-private-law forks isolate the effect of a registered world change. (4) Agent experimentation through the same environment—deterministic workflows and an independent agent use the same public interface while the evaluator retains the complete experimental process record without exposing private state.
ChemWorld represents reaction, thermal, phase, separation, crystallization, distillation, continuous-flow, electrochemical and observation processes as reusable components. Each component declares its parameter domains, dependencies, owned state and public interfaces. A complete world is represented as W = (Wpub, θ), where Wpub contains the public component topology, public parameter domains and interfaces, whereas θ contains evaluator-owned constitutive laws, material properties, hidden parameters and private initialization. The public task contract T joins the public world description and initial-state projection with the allowed actions, instruments, observations, resources, termination rule and evaluation rule. Three identity levels are distinguished: world-spec ID = id(Wpub, θ), scenario ID = id(Wpub, θ, ζinit, ζdyn, ζobs), and task–world unit = (task ID, scenario ID).
The compatibility compiler normalizes the declaration and checks component dependencies, state ownership, unit agreement, supported parameter domains, resource feasibility, instrument availability, operation exposure and lifecycle closure. Only declarations that satisfy all checks are compiled into executable worlds. An invalid declaration fails closed before environment construction and returns structured diagnostics. A successful compilation returns both the executable world and one public experimental contract that exposes the allowed typed operations, instruments, observations, resources, failure semantics, termination and evaluation, while evaluator-owned mechanisms and hidden state remain private.
Every submitted action follows the same transactional sequence: preflight admission, runtime-precondition evaluation, candidate execution, post-execution validation and either commit or rollback. A schema, compatibility and resource predicate P(st, at, Rt) first determines whether an action may enter the runtime. An admitted action is then tested against its context-dependent runtime preconditions. If those preconditions pass, the bound mechanism θ, committed state st, resource ledger Rt, action at and recorded random variates ξt generate a candidate transition: (s̃t+1, R̃t+1, ẽt+1) = Fθ(st, Rt, at, ξt). Candidate state is not installed immediately. A runtime commit-gate predicate C ∈ 0, 1 covers both a runtime-precondition rejection before candidate generation and post-execution checks of state integrity, solver status, runtime invariants and the observation path. The transaction commits only when admission and all applicable runtime checks pass: P = 1, C = 1 =⇒ (st+1, Rt+1, e⋆t+1) = (s̃t+1, R̃t+1, eacc t+1). The two non-commit branches remain distinct in the record. If P = 0, the runtime emits a preflight-rejection event and receipt without runtime execution. If P = 1 but C = 0, it emits a runtime-rollback event and receipt. For b ∈ pre, roll, st+1 = st, Rt+1 = Gb(Rt, at, ebt+1), e⋆t+1 = ebt+1. The branch-specific ledger function installs only the protocol-declared attempt cost or penalty. Candidate physical, observation and uncommitted resource effects are discarded. The committed runtime state also binds the observation-RNG state ρt; a non-commit branch restores ρt, preventing an unsuccessful attempt from changing future observation noise. Public and evaluator records are separate projections of the realized branch: ot+1 = πpub(st+1, e⋆t+1), ôt+1 = πeval(st+1, e⋆t+1).
Exact replay binds the normalized contract, runtime, mechanism and scoring identities, together with the seeds and intervention record. It reconstructs the compiled world and resubmits the full submitted action/transaction trace—including committed actions, preflight rejections and runtime rollbacks—and compares public observations, transaction outcomes, affected-ledger declarations, world events, rewards and terminal flags at zero numerical tolerance. Resource deltas and rollback receipts are reconciled separately.
Qualification was conducted at four connected levels: construction coverage, complete-world execution, module and interface semantics, and failure, resource and replay behavior. The public capability map contains 15 registered reference tasks, 28 typed operation kinds and five synthetic instrument contracts. The reference qualification set contains 64 task–world units and 1,786 boundary and categorical recipes. An execution protocol fixed eight component patterns and generated 52 additional compositions before authoritative qualification. Eighteen compositions use three topologies absent from the reference registry: phase–observation, phase–separation–observation and reaction–thermal–continuous-flow–observation. Eight reaction–thermal–distillation–observation compositions reuse a registered topology but have zero exact task–world identity overlap with the frozen registry. The remaining 26 rows provide additional coverage within registered topologies.
All 64/64 reference units passed, and the boundary and categorical set produced 1,786/1,786 complete executions. All 52/52 generated compositions also passed, including all 8/8 protocol-frozen non-reference reaction–distillation worlds. There were no missing receipts, failure classes or undeclared private-field-exposure findings. Seven deliberately invalid declarations covered missing dependencies, conflicting state ownership, unit mismatches, invalid parameter domains, resource impossibility and lifecycle gaps. All 7/7 failed closed before environment construction with the registered diagnostic class. Thirty-two module probes exercise zero input, declared parameter boundaries, monotonic directions, conservation and model-specific invariants. Seven cross-module paths test whether material amount, unit, identity and state meaning survive transfer between components. All 32/32 module probes and 7/7 interface paths passed.
The frozen campaign contained 192 negative probes: 64 invalid-schema/unknown-operation probes, 64 campaign-resource-exhaustion probes and 64 runtime-precondition probes. Every probe produced its registered outcome and preserved committed physical state. The evidence therefore qualifies 128 P = 0 admission rejections and 64 P = 1, C = 0 runtime-precondition rollbacks. The resource ledger independently tracks material, sample, instrument use, process time, operation count and terminal assay. Failed attempts retain only their declared costs or penalties, and observation checks require public packets to contain only task-declared fields.
Eight protocol-frozen use cases span reaction-to-crystallization, resource-limited equilibrium characterization, planned failure and recovery, continuous flow, electrochemistry, distillation, partition and a second crystallization world. Across the eight cases, all 89 submitted actions have complete schema, transaction, state-integrity, event, resource and public-observation receipts. Eighty-eight actions committed and one protocol-frozen action rolled back. Every case completed one final assay, closed its lifecycle, reconciled resources and replayed exactly with zero numerical error. The failure–recovery case places one deliberate invalid operation inside an otherwise complete experiment. Its first action passed schema, compatibility and campaign-resource admission, but a runtime-precondition check found that no separable phase had yet formed. The transaction consequently entered the recorded P = 1, C = 0 branch before candidate physical state was generated. It preserved committed physical state and observation-RNG state, created no ghost state, and reconciled the declared failed-attempt consequences. The next 18 actions continued from the last committed state and completed the recovery path, final assay, resource reconciliation and exact replay of all 19 submitted actions.
The qualification contains six parent–child pairs: two intervention classes evaluated over three seeds. Every pair preserves nine versioned public-contract components—task, actions, instruments, observations, resources, failures, scoring, material catalogue and the contracted invariant/safety surface—as well as the fixed typed-action sequence and bound randomness. Parent and child consequently have distinct complete-world identities while presenting the same public experiment. The admissible change is restricted to exactly one protocol-frozen private constitutive or material law. The partition intervention changes the hidden response from K 1.00 to K 1.75; the registered terminal organic-product amount and public product in organic assay both increase. The electrochemical intervention keeps public material labels fixed while reassigning hidden electrolyte-response profiles; the registered selective-product amount and public ohmic efficiency both decrease. Repeating both variants produced 24 deterministic traces. All six pairs passed lineage, exactly-one-private-target, public-contract-invariance, same-sequence-executability, expected-state-divergence, expected-observation-divergence and exact-replay gates.
ChemWorld separates agent-facing information from evaluator-side evidence. The agent receives the public task card, typed operations, instruments, public observations, available resources and termination and final-assay interfaces. It cannot directly inspect hidden material properties, private constitutive-law parameters, private initialization or complete simulator state. The evaluator retains the complete state transitions, transaction outcomes, failure and resource consequences and environment/action trajectory. ChemWorld groups 19 registered process coordinates into terminal commitment, evidence acquisition, evidence-conditioned action, resource deployment and outcome trajectory. They remain separate coordinates rather than a composite score or a set of uniformly higher-is-better metrics.
World qualification and agent execution were kept as independent experimental units so that agent success could not serve as evidence for world correctness. The selected protocol-frozen non-reference world combines reaction, thermal, distillation and observation components. A deterministic 12-action path first qualified its construction, lifecycle closure and exact replay. Only after world qualification was complete did an independent agent enter the same public instrument contract and complete a separate 15-action lifecycle. The agent received no private world fields and issued every experimental operation, including termination and final assay. In one uninterrupted experiment, it submitted 15 actions; all 15 committed and closed the lifecycle, with no rollback, right-censoring or undeclared private-field exposure. The environment used 8,158.454 of 10,440 simulated process seconds, four of four instrument uses and 0.00085 of 0.001 L sample. The entire 15-step submitted-action trace replayed with zero numerical mismatch.
The conclusions of this study are limited to the declared component vocabulary, compatibility rules and authored model domains. Qualification establishes consistent executable semantics and internally coherent behavior for these software models within their stated scope; it does not establish that ChemWorld fully reproduces real chemical systems. The current synthetic instruments are controlled software observation models rather than digital twins calibrated to particular devices. The coverage design systematically samples prespecified parameter and composition spaces but does not exhaust chemical space or all higher-order process interactions. Module fixtures and directionality oracles primarily test consistency with the authored models rather than an independent reference implementation or empirical ground truth. Exact replay is likewise restricted to the bound software identities, world identity and environment–action trace; it does not imply policy re-execution, cross-platform numerical identity or cross-version archival replay. ChemWorld and physical self-driving laboratories are complementary experimental regimes rather than substitutes. A natural workflow is to use software experiments to identify conditions, mechanisms and agent behaviors that merit closer study, then return questions requiring real chemical evidence to the physical laboratory.
A frozen, versioned public release associated with this study is openly available under the MIT License in the ChemWorld public repository; the frozen code-and-evidence snapshot referenced here is tagged v0.1.0. The release contains the executable code, publication configurations and protocols, processed machine- and human-readable qualification evidence, figure source files, replayable simulator environment–action trajectories, the arXiv PDF and source bundle, and an offline manifest verifier for the reported hashes and denominators.
Improvements for AI systems
Based on this paper, I can improve AI systems in the following specific ways:
1. Transactional Action Execution with Rollback
-
Implement a preflight admission check, runtime-precondition evaluation, candidate execution, post-execution validation, and commit-or-rollback sequence for every AI action.
-
The improved AI can reject invalid actions before execution, roll back failed attempts without corrupting state, and preserve randomness/noise state across failures.
2. Exact Replay and Auditability
-
Bind all seeds, intervention records, and environment identities to enable zero-tolerance replay of full action traces, including failed attempts.
-
The improved AI can reproduce any past experiment exactly, reconcile resource deltas, and provide attributable evidence for every decision.
3. Separated Public Contract vs. Private State
-
Expose only a task-declared public interface (actions, observations, resources, termination) while keeping hidden parameters, constitutive laws, and internal state private.
-
The improved AI can operate under incomplete information, learn from public observations only, and cannot exploit or leak private evaluator state.
4. Programmable World Composition with Compatibility Checking
-
Build environments from reusable components with declared dependencies, state ownership, unit agreements, and lifecycle closure; fail closed on invalid declarations.
-
The improved AI can be tested across systematically varied worlds (topologies, operating conditions, hidden laws) while holding the public task fixed, enabling controlled intervention studies.
5. Single-Private-Law Forking for Causal Isolation
-
Create parent–child world pairs that differ in exactly one hidden law while preserving the public contract, action sequence, and randomness.
-
The improved AI can attribute observed behavioral changes to specific private interventions, supporting causal analysis of agent policies.
6. Branch-Specific Resource Ledgering
-
Track material, sample, instrument use, process time, operation count, and terminal assays separately; failed attempts only incur declared costs or penalties.
-
The improved AI can manage scarce resources realistically, plan around failed attempts, and reconcile resource usage without ghost state.
7. Observation-RNG State Preservation on Failure
-
Restore the observation-noise random state on rollback so unsuccessful attempts do not alter future observation noise.
-
The improved AI can rely on consistent observation distributions across retries, improving learning stability.
8. Multi-Level Qualification Before Agent Deployment
-
Validate world construction, module semantics, interface paths, failure modes, and replay fidelity independently of agent performance.
-
The improved AI can be deployed only after the environment is proven correct, avoiding conflating agent success with environment correctness.
9. Coordinated Evaluation Metrics
-
Group process coordinates into terminal commitment, evidence acquisition, evidence-conditioned action, resource deployment, and outcome trajectory—rather than a single composite score.
-
The improved AI can be evaluated on multiple orthogonal dimensions, enabling nuanced assessment of exploration, efficiency, and final performance.
10. Deterministic Workflow and Independent Agent Parity
-
Use the same public interface for scripted deterministic workflows and autonomous agents, with the evaluator retaining full process records.
-
The improved AI can be benchmarked against optimal or scripted baselines under identical conditions, and its full action trace can be audited post-hoc.
Abstract
Autonomous chemistry increasingly depends on environments in which agents can repeatedly act, observe, and adapt.Physical laboratories provide essential real-material evidence but are costly to repeat and difficult to use for tightly matched interventions, whereas most digital environments keep the underlying experimental world largely fixed. We introduce ChemWorld, a programmable chemical environment in which reusable process and observation components are compiled into executable worlds. ChemWorld separates the public experimental contract available to an agent from evaluator-owned chemical and material laws. Researchers can therefore vary world composition and operating conditions, or change a single hidden law while holding the public task and interaction conditions fixed. Transactional execution records operations, failures, resource changes, and state transitions, allowing complete environment-action trajectories to be replayed exactly and audited. Full-census qualification covered the reference registry, 52 generated compositions, and module, interface, compilation, and invalid-action tests. Eight deterministic experimental cases demonstrated shared lifecycle semantics, failure recovery, and exact replay, while six parent-child world-fork pairs isolated the effects of single private-law interventions under matched public conditions. An independent agent also completed a full lifecycle in a non-reference world through the same public interface. Within the declared component and model domain, ChemWorld provides a controlled and replayable substrate for studying experimentation across systematically varied chemical worlds, complementary to physical-laboratory evidence and calibration.
Sources
- BoxingGym: Benchmarking Progress in Automated Experimental Design and Model Discovery
- Measuring Scientific Capabilities of Language Models with a Systems Biology Dry Lab
- Scaling Scientific Discovery Environments for Turn-Level Agentic RL
- PC-Gym: Benchmark Environments For Process Control Problems
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection