Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence
Brian Wang, Bin Feng, Xiaoman Pan, Chenyang An, Felix Liu, Tangqi Fang, Gongbo Sun, Lingfeng Shen, Ning Wang, Handuo Zhang, Feng Chen, Fuchao Yang, Xiang Wang, Jiacheng Lin, Siting Li, Zixuan Liu, Chi Han, Zhenhailong Wang, Kunlun Zhu, Lawrence Zhao, Yueqi Guo, Kailong Wen, Feng Xing, Yiling Guo, Lidong Bing, David Tan, Bo An, Heng Ji, Sheng Wang
cs.AI
Submitted: 2026-08-11
Updated: 2026-08-13
Comments: 85 pages, 9 figures, 38 tables
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 95/100
The gist: " Summary This paper introduces Apodex Discovery, a comprehensive framework for building and evaluating "discoverative AI"—AI systems directed at making genuine discoveries rather than solving
Terminology
Summary
"
Summary
This paper introduces Apodex Discovery, a comprehensive framework for building and evaluating discoverative AI
—AI systems directed at making genuine discoveries rather than solving predefined problems. The authors argue that frontier AI models are increasingly capable of solving difficult tasks once the problem, tools, and success criteria are specified, but the most consequential real-world challenges are rarely presented in such an executable or verifiable form. The paper's central thesis is that the next critical challenge for AI lies in constructing the infrastructure that translates model capabilities into solutions for open-ended, high-value problems.
To address this, the paper presents three principal contributions:
-
A systematic process for identifying high-value, real-world problems.
-
A formalized framework for creating executable, reality-based environments for these problems.
-
A novel process-verification metric, HDS6, to evaluate the quality of the investigation itself.
The paper also introduces TRACES, a reality benchmark that instantiates this framework, and reports on initial results demonstrating the framework's utility.
1. The Heavy-Duty Solver and Core Infrastructure
The operational unit of discoverative AI is defined as a heavy-duty solver
(HDS). This is not just a foundation model but a complete system comprising the model, a harness, tools, memory, and control policies. The paper identifies four essential components of infrastructure required to translate an open-ended ambition into a tractable problem for an HDS:
-
Problem Manifest: A static, declarative contract that defines the problem's question and success criteria, its decomposition into tasks, the target outcome and hidden verifier, the data and tools available, and the resource budgets.
-
Reality-Based Environment: A stateful, interactive substrate that provides the solver with data, tools, and feedback. Its defining property is that it returns fresh observations in response to the solver's actions, unlike a static prompt. The environment is isolated to prevent leakage of hidden ground truth.
-
Verification Mechanisms: Two complementary forms of verification are used. Outcome verification scores the final submission against hidden ground truth. Process verification evaluates the recorded trajectory of the solver's actions to assess the quality of the investigation.
-
Repair Loops: A mechanism that returns the results of verification to the solver, allowing it to localize failures, correct errors, and improve its solution in a measure–diagnose–improve loop.
2. Scouting High-Value Real-World Problems
The paper details a rigorous, two-month process by a ten-person team of STEM Ph.D. researchers to identify suitable problems. This process involved:
-
Surveying 561 industries across 16 sectors.
-
Assessing each industry on technical fit, value, and commercial access.
-
Selecting 10 domains for further exploration.
-
Expanding these domains into 423 high-value candidate problems.
-
Screening these candidates against strict requirements for verifiability, impact, and a clear beneficiary, resulting in a registry of 20 problems for the initial TRACES release.
The initial benchmark focuses on five problem families: AAV capsid design, drug repurposing and reformulation, clinical-trial analysis, data for large language models, and training/systems for large language models. The problems are designed to be reality-facing,
meaning they cannot be solved by simple retrieval or memorization. They are divided into retrospective tasks (where a known outcome is withheld) and prospective discovery challenges
(where no authoritative answer exists and evaluation happens as real-world outcomes emerge).
3. The HDS6 Process Verification Metric
A key innovation is the HDS6 metric, which evaluates the quality of a solver's process by analyzing its recorded trajectory. It scores six capabilities, denoted by the mnemonic TRACES:
-
Tools: Selecting, invoking, and correctly interpreting external tools.
-
Repair: Responding to verification feedback or observed failure, correcting the underlying error, and confirming the fix.
-
Alternatives: Making competing hypotheses explicit and adjudicating among them as evidence accumulates.
-
Coherence: Maintaining state, constraints, and logical consistency over an extended horizon.
-
Evidence: Grounding claims in observations, tool outputs, data, experiments, or references.
-
Scope: Identifying the conditions under which a conclusion holds and the boundaries beyond which it should not be applied.
The HDS6 process verifier operates blind,
meaning it cannot access the verified outcome, the solver's private chain-of-thought, or the solver's identity. It scores the trajectory using two complementary methods: a fixed subrubric-based checklist and a load-bearing step-based analysis. The process verifier also generates a structured repair note
that can be returned to the solver to guide a new attempt, closing the repair loop.
4. Results on High-Value Problems
The paper reports initial results that demonstrate the framework's effectiveness:
-
AAV Capsid Design: In this four-task pipeline (viability, tropism, structure prediction, and generative design), the Apodex system surpassed the task-level published state of the art by 7% across all four tasks. For example, on viability prediction, apodex-1.1 achieved an out-of-distribution AUROC of 0.904, above the CAP-PLM capsid language model's 0.878. The paper also shows that a domain-specific environment improves outcomes; for instance, claude-opus-4-8 scored 0.741 in the Apodex environment versus 0.716 with a generic harness.
-
Drug Repurposing and Reformulation: Adding a task-specific biomedical environment improved the mean normalized prediction score of GPT-5.5 and GPT-5.6-sol by 2.5 and 7.6 points, respectively, over the same closed-book backbone. The environment also led to significant improvements in HDS6 process scores, particularly in tool use, evidence fidelity, and long-horizon coherence.
-
LLM Problems and Ablations: The paper presents a detailed analysis of solver configurations across multiple harnesses and models. It demonstrates that the fixed TRACES episode interface enables performance differences to be attributed to specific components of the solver system, such as the foundation model, harness, initial skill guidance, and terminal verifier-guided repair. For example, skill guidance improved the OpenHands score on the LLM-pretrain-data-probe task from 0.442 to 0.557 with Opus-4.8 and from 0.534 to 0.730 with GPT-5.6-sol.
-
Verification-Driven Repair: The paper shows that the repair loop is effective. Across 434 deficient trajectories (HDS6 < 3), the repaired re-runs scored 0.155 higher on average, with 204 trajectories improving against 86 declining. Two detailed case studies illustrate how the process verifier identifies specific flaws and guides the solver to correct them, leading to improved outcomes.
5. Conclusion and Limitations
The paper concludes that Apodex Discovery shifts evaluation from asking whether a model produced the right answer to asking whether a solver conducted a reliable, evidence-grounded, and self-correcting investigation. The authors acknowledge limitations, noting that the current TRACES benchmark instantiates only 17 of the planned environments and that the process-verification results, while promising, leave room for further development to better align with outcome scores. The benchmark is available at discovery.apodex.com.
Improvements for AI systems
Improvements to AI systems based on this paper:
-
Add a
Problem Manifest
module to AI agents. The AI system can now accept open-ended, high-level goals (e.g.,design a better gene therapy vector
) and automatically decompose them into executable tasks with explicit success criteria, available tools, data sources, and resource budgets—without human specification. This enables the system to tackle ambiguous, real-world problems rather than only well-defined benchmarks. -
Implement a
Reality-Based Environment
interface. The AI system gains a stateful, interactive substrate that returns fresh observations in response to its actions (e.g., running a simulation, querying a database, executing a tool) rather than relying on static prompts. This allows the system to actively probe the world, gather new data, and adapt its strategy based on live feedback—critical for discovery tasks where the answer is not in training data. -
Add a
Process Verifier
(HDS6) that scores the AI's investigation, not just the final answer. The system can now self-evaluate its own trajectory on six dimensions: Tool use, Repair, Alternatives, Coherence, Evidence, and Scope. It can identify weaknesses in its reasoning process (e.g., failing to consider competing hypotheses, overgeneralizing conclusions, or not grounding claims in evidence) and generate a structuredrepair note
to guide its next attempt. This enables the AI to improve its methodology iteratively, not just its output. -
Integrate a
Repair Loop
for self-correction. After each attempt, the AI receives process-verification feedback (e.g.,You ignored a key alternative hypothesis
orYour conclusion lacks evidence for condition X
) and automatically re-runs its investigation with targeted corrections. The system can now localize failures, fix underlying errors, and confirm improvements—demonstrating a measure–diagnose–improve cycle that raises both process quality and outcome scores (e.g., +0.155 average improvement on deficient trajectories). -
Add
Skill Guidance
injection for domain-specific environments. The AI system can now load pre-built, task-specific guidance (e.g., biomedical tool usage patterns, capsid-design heuristics) into its harness, improving performance on specialized tasks (e.g., +0.196 improvement on LLM-pretrain-data-probe with GPT-5.6-sol). This allows the system to leverage domain expertise without retraining, making it more effective in niche, high-value fields. -
Enable
Blind Process Evaluation
for unbiased self-assessment. The AI can now evaluate its own process without access to the ground-truth outcome, its private chain-of-thought, or its identity—preventing overfitting to known answers. This makes the system's self-critique more honest and transferable to novel problems where no authoritative solution exists. -
Support
Prospective Discovery Challenges.
The AI system can now operate on problems with no known answer, where evaluation happens as real-world outcomes emerge (e.g., proposing a new drug formulation and waiting for clinical trial results). This extends the system's utility from retrospective problem-solving to genuine forward-looking discovery.
What the improved AI system can do:
-
Take a vague, high-impact ambition (e.g.,
find a new use for an existing drug
) and autonomously convert it into a structured, executable investigation with clear success metrics. -
Actively interact with external tools and data sources, gathering fresh evidence and adjusting its approach in real time.
-
Self-critique its own investigation process (not just its answer) across six quality dimensions, and automatically repair flaws in its reasoning before resubmitting.
-
Outperform current state-of-the-art on domain-specific tasks (e.g., 7% better than published SOTA on AAV capsid design across four tasks) and show measurable gains from environment and guidance enhancements.
-
Operate reliably on open-ended, unverifiable problems by focusing on process quality (evidence-grounded, coherent, scope-aware) rather than chasing a single correct answer.
-
Continuously improve its own methodology through iterative repair loops, making it more trustworthy for high-stakes, real-world applications like drug discovery, clinical trial analysis, and LLM training-data curation.
Sources
- Accelerating scientific discovery with Co-Scientist
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
- Humanity's Last Exam
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- AAVGen: Precision Engineering of Adeno-associated Viral Capsids for Renal Selective Targeting
- AAVDiff: Experimental Validation of Enhanced Viability and Diversity in Recombinant Adeno-Associated Virus (AAV) Capsids through Diffusion Generation
- Position: Agentic Evolution is the Path to Evolving LLMs
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- OpenAI Gym
- Evaluating Large Language Models Trained on Code
- Training Software Engineering Agents and Verifiers with SWE-Gym
- Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
- MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
- RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
- PaperBench: Evaluating AI's Ability to Replicate AI Research
- ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
- DiscoveryBench: Towards Data-Driven Discovery with Large Language Models
- LAB-Bench: Measuring Capabilities of Language Models for Biology Research
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection