SimulRAG: Simulator-based RAG for Grounding LLMs in Long-form Scientific QA
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "SimulRAG: Simulator-based RAG for Grounding LLMs in Long-form Scientific QA".
Jane: The paper was written by Haozhou Xu, Dongxia Wu, Matteo Chinazzi, Ruijia Niu, Rose Yu et al. from University of California San Diego and Stanford University and Northeastern University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Alright, welcome back to the show, everybody. Today we’re cracking open a paper that’s got a mouthful of a title: “SimulRAG: Simulator-based RAG for Grounding LLMs in Long-form Scientific QA.” Jane, I’ll be honest, I needed to read that title three times before it started making sense.
Jane: Ha, you and me both, Tom. But once you unpack it, it’s actually a really elegant idea. So you’ve got LLMs, the big language models, and they’re great at writing essays, but they’re also great at making things up. That’s the hallucination problem. And RAG, which stands for Retrieval-Augmented Generation, is the trick where you give the model some outside information to ground it, so it’s not just pulling stuff out of thin air.
Tom: Right, and usually that outside information is like a Wikipedia page or a database of documents. But this paper from UC San Diego and Stanford, they’re saying, what if the thing you retrieve from isn’t a document at all? What if it’s a scientific simulator? Like, a piece of software that actually models the climate or the spread of a disease.
Jane: Exactly. And that’s why the title is so clever. It’s not just RAG, it’s SimulRAG. They’re swapping out the library of books for a laboratory. The authors, Haozhou Xu, Dongxia Wu, and the whole team, they’re basically asking the LLM to not just recall facts, but to run experiments to check its own answers.
Tom: Which is a wild leap. I mean, think about the difference. If I ask a model about the weather in two thousand fifty a normal RAG system might pull a static report. But SimulRAG would actually fire up a climate emulator, plug in the CO2 levels, and see what the temperature comes out as. That’s not retrieval, that’s computation.
Jane: And that’s the part that gets me excited. It’s turning the LLM from a student who memorizes the textbook into a researcher who runs the test to see if the hypothesis holds up. The implications for trust are huge, because you’re not just hoping the model is right, you’re giving it a way to prove it.
Tom: So we’ve got the big idea, but I’m dying to know how they actually built this thing. How do you get a language model to talk to a physics engine? That sounds messy.
Jane: Oh, it’s definitely messy, and that’s exactly what we’re going to dig into next. They had to build a whole translation layer, and I want to see how they pulled that off.
Summary: Tom: So, Jane, we left off with the big question: how do you bridge the gap between words and numbers? The paper’s summary lays it out pretty clearly. They built a “simulator retrieval interface.” It’s basically a translator that takes a question like “what happens if CO2 goes up thirty percent?” and turns it into the exact parameters the simulator needs.
Jane: Right, and I love that they use the simulator’s handbook to do it. So the LLM reads the manual, figures out what knobs it can turn, and then turns them. It’s like giving the model a remote control for the simulator and telling it to press the right buttons. Then it runs the simulation, gets the raw numbers, and translates those back into plain English sentences.
Tom: But here’s the part that really blew my mind. They don’t just generate one answer and check it. They generate five different answers, break each one down into tiny atomic claims, and then only check the ones that are both uncertain and actually verifiable by the simulator. They call it UE+SBA, uncertainty estimation plus simulator boundary assessment.
Jane: And that boundary part is so smart. Because not every claim can be checked by a climate model. If the model says “this policy will be unpopular,” the simulator can’t help with that. So they have a separate step that filters out those claims before wasting any compute on them. They’re being really deliberate about where they spend their verification budget.
Tom: And the results speak for themselves. On their benchmarks, they improved informativeness by over thirty percent compared to the strongest RAG baselines, and factuality by over sixteen percent. That’s not a small bump. That’s a real leap in quality.
Jane: It is. And it makes sense when you think about it. By checking the uncertain claims against the simulator, they’re catching the hallucinations before they make it into the final answer. But they’re also not rewriting the whole answer, which would mess up the parts that were already right. It’s surgical.
Tom: Surgical is a good word for it. But I’m curious about the actual mechanics of the update. When the simulator says “no, you’re wrong, it’s actually one point zero eight degrees, not zero point five,” how does the model handle that? Does it just swap the number in?
Jane: That’s the million-dollar question, and it’s actually the next part of the paper. They have a whole process for how the claim gets updated, and it’s not just a find-and-replace. Let’s get into the improvements they suggest, because that’s where the real engineering lives.
Improvements: Tom: Okay, so we’ve covered the framework, and now we’re looking at the specific improvements the paper suggests. Jane, you were just about to tell me how the claim update actually works when the simulator contradicts the model.
Jane: Right. So the paper describes a three-way decision. When a claim gets checked against the simulation output, it’s either aligned, contradicted, or indeterminate. If it’s aligned, they keep it as is. If it’s contradicted, they rewrite it based on the simulation evidence. And if the simulator just can’t tell, they leave it alone. It’s a really clean way to handle the uncertainty.
Tom: And that’s where the “improvements” part of the paper comes in. They’re not just proposing a new framework, they’re showing how to make it efficient. Because checking every single claim against a simulator would be way too slow and expensive. So they use that uncertainty score to only check the claims that are actually in doubt.
Jane: Exactly. And they tested this really thoroughly. They compared their UE+SBA method against just picking claims randomly, and against asking the LLM to say how confident it is. Their method won across the board, on every benchmark and every model they tried. The improvement was consistent.
Tom: I remember the numbers. They got up to a six percent absolute improvement in AUPR, which is a measure of how well they’re ranking which claims to check. And the really cool part is that they showed the simulator boundary assessment is doing real work. It filters out a huge chunk of claims that the simulator just can’t verify, like over fifty percent in the urban planning domain.
Jane: That’s a massive efficiency gain. You’re not wasting time trying to verify things that are outside the simulator’s scope. And they also showed that their method works with different uncertainty estimators, not just one specific formula. So it’s robust.
Tom: Lu, you’ve been quiet. What’s your take on this? You’re the researcher, does this hold up?
Lu: I think the most impressive part is the benchmark they built. They created a dataset for climate, epidemiology, and urban planning, and they had human annotators verify the ground truth. That’s a lot of work, and it makes the results much more credible. It’s not just a toy example.
Meng: Yeah, but from an engineering standpoint, I want to know about the cost. Running a simulator for every question, even selectively, has to be expensive. Did they talk about that at all?
Jane: They did, actually. They showed that with just a fifteen percent verification budget, they get most of the benefit. And going from forty-five percent to one hundred percent only adds a tiny bit of quality. So you can get most of the factuality boost without running the simulator on every single claim. That’s a practical win.
Tom: So it’s not just a lab experiment, it’s something that could actually be deployed. That’s a great segue into the bigger picture. What does this mean for the future of eye in science?
Conclusion: Tom: Alright, we’ve been deep in the weeds on “SimulRAG: Simulator-based RAG for Grounding LLMs in Long-form Scientific QA,” and I think it’s time to zoom out and say goodbye to this one. Jane, what’s the big takeaway for our listeners?
Jane: The big takeaway is that we’re moving from eye that talks about science to eye that does science. SimulRAG gives language models a way to actually test their claims against physical reality, or at least against our best computational models of it. That’s a huge step for trust.
Tom: And it’s not just for climate. They showed it works for epidemiology, for urban planning. Anywhere you have a simulator, you can plug it into this framework. The authors even mention that future work could focus on automatically detecting when a question is even relevant to a simulator, so you don’t waste time on mismatches.
Lu: I think the most exciting implication is for hypothesis generation. If an eye can propose a hypothesis and then immediately test it against a simulator, that accelerates the whole scientific loop. It’s not just answering questions, it’s exploring possibilities.
Meng: And from a practical side, the efficiency gains are what make it viable. They showed you don’t need to verify everything to get most of the benefit. That’s what makes this deployable, not just a research curiosity.
Lalam: If I may add, the cultural impact here is about democratizing expertise. A tool like this could help a city planner without a PhD in climate science ask “what if we change this road?” and get an answer that’s grounded in actual traffic simulation. It lowers the barrier to using complex scientific tools.
Tom: That’s a beautiful way to put it, Lalam. So, to wrap it up, “SimulRAG” is a framework that lets LLMs query scientific simulators to verify their claims, and it does it efficiently and effectively. It’s a big win for factuality and informativeness in long-form scientific answers.
Jane: And it’s a sign of where we’re headed. eye that checks its own homework against the real world. We’re excited to see where this goes, and we’re ready to move on to the next paper. Thanks for listening, everyone.
Tom: See you on the next one.
Haozhou Xu, Dongxia Wu, Matteo Chinazzi, Ruijia Niu, Rose Yu, Yi-An Ma
University of California San Diego · Stanford University · Northeastern University
cs.CL, cs.LG
Submitted: 2026-08-18
Updated: 2026-08-19
Comments: Haozhou Xu and Dongxia Wu are co-first authors
Code: https://github.com/HZEmpire/SimulRAG
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 83/100
The gist: The paper introduces SimulRAG, a simulator-based Retrieval-Augmented Generation (RAG) framework for long-form scientific question answering.
Key concepts
- RAG
- Retrieval-Augmented Generation (RAG) is a technique that provides large language models with outside information, such as databases or documents, to ground their answers and prevent them from hallucinating or making up information.
- SimulRAG
- SimulRAG is a framework that swaps traditional text retrieval for scientific simulators. It allows language models to act like researchers by running computational experiments, such as climate emulators, to verify their hypotheses against simulated physical reality.
- UE+SBA
- UE+SBA (uncertainty estimation plus simulator boundary assessment) is a method to improve efficiency. It identifies uncertain claims and filters out those the simulator cannot verify, ensuring the system only spends computational resources on claims that can be checked.
Terminology
Summary
The paper introduces SimulRAG, a simulator-based Retrieval-Augmented Generation (RAG) framework for long-form scientific question answering. The authors identify that while Large Language Models (LLMs) show promise in generating long-form scientific explanations, they often hallucinate in long-form scientific QA tasks. Traditional RAG approaches that ground generation in external textual sources are insufficient for scientific domains where quantitative hypotheses and evolving dynamics must be validated. The paper states: "Retrieval-Augmented Generation (RAG) improves trustworthiness by grounding generation in external sources; scientific simulators are valuable because they can validate quantitative hypotheses and capture evolving dynamics."
The framework addresses two fundamental challenges: "First, the discrepancy between textual space and the numerical space where simulation parameters and outputs reside creates difficulties for querying scientific simulators. Second, existing RAG generators cannot effectively update long-form answers with new context because they lack fine-grained control."
Simulator Retrieval Interface: The paper proposes a generalized simulator retrieval interface to transform between textual and numerical modalities, enabling seamless integration of scientific simulators into RAG systems.
The interface uses the question and simulator handbook to guide the LLM in extracting parameter settings, executes the simulator with those parameters, and converts outputs to textual context via templates. The authors note: Adapting a new simulator only requires a lightweight adapter specifying its functions, parameter schema, and outputs; the retrieval and claim-verification pipeline remains unchanged.
Claim-Level Generation: The framework decomposes long-form answers into atomic claims and verifies or updates each claim.
Multiple diverse answers are generated, decomposed into atomic claims, and merged into a deduplicated set. The authors state this "serves two critical purposes: (1) enabling targeted verification and updates of individual claims rather than holistic response modification, providing more precise and flexible long-form answer refinement; (2) simplifying the verification task by focusing on atomic factual statements."
UE+SBA Method: To improve efficiency, the paper introduces uncertainty estimation scores and simulator boundary assessment (UE+SBA) to selectively verify and update claims only when necessary.
Uncertainty estimation uses graph centrality metrics (specifically closeness centrality) on bipartite entailment graphs between answers and claims. Simulator boundary assessment uses an LLM judge with the simulator handbook to determine whether claims fall within simulator operational boundaries. Claims undergo verification only when they meet two criteria: "uncertainty: conf(ci) < τ and
boundary compatibility: bound(ci, h) = 1."
Benchmark Construction: The authors construct a benchmark dataset for long-form scientific QA using simulators as retrieval tools, covering climate modeling, epidemiology, and urban planning domains.
The climate dataset uses a climate emulator trained on CMIP6 simulations with 1000 questions; the epidemiology dataset uses GLEAM-AI for influenza transmission with 1000 questions; the urban planning dataset uses the SUMO simulator with 200 questions. Ground truth answers are verified by both scientific simulators and human annotators.
Key Results: Experiments show SimulRAG improves informativeness by 30.4% and factuality by 16.3% over the strongest adapted RAG baselines.
The UE+SBA method consistently outperforms all baseline methods across all metrics, models, benchmarks, and budgets,
achieving up to 6.2% absolute improvements
in AUPR and up to 6.3% absolute improvements
in AUROC over the best baseline. Ablation studies show that SimulRAG with 15% RAG significantly outperforms no-RAG across the three benchmarks and both models,
with F1 improving by 5.2, 8.5, and 14.1 percentage points at verification budgets of 15%, 25%, and 45% respectively.
Contributions: The paper summarizes six contributions: introducing SimulRAG, proposing the generalized simulator retrieval interface, presenting claim-level generation, utilizing UE+SBA for efficient verification, constructing the long-form scientific QA benchmark, and conducting extensive experiments verifying effectiveness.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in an AI system:
-
Implementation: Add a module that converts natural language queries into structured simulator parameters (JSON) using LLM prompting with a simulator handbook, executes the simulator, and verbalizes outputs back into text via templates.
-
What it enables: The AI can now query scientific simulators (climate models, epidemic models, traffic simulators) directly from open-ended questions, without needing fine-tuning or predefined question templates.
-
Implementation: Decompose long-form answers into atomic, independently verifiable claims; merge claims across multiple sampled answers via semantic equivalence detection; generate final answers from high-confidence claims only.
-
What it enables: The AI produces answers with finer-grained control—each claim can be individually verified, updated, or filtered, leading to more precise and trustworthy long-form responses.
-
Implementation:
-
Compute claim-level confidence via closeness centrality on an entailment graph between answers and claims.
-
Use an LLM judge with the simulator handbook to filter claims that fall outside simulator verification boundaries.
-
Only verify/update claims that are both uncertain (confidence < threshold) and simulator-verifiable.
-
What it enables: The AI selectively allocates verification effort—reducing computational cost by 55%+ while maintaining answer quality (within 3.6% of full verification). It avoids wasting resources on unverifiable or already-confident claims.
-
Implementation: For each selected claim, compare against simulation context: (a) if aligned, keep as-is; (b) if contradicted, update with simulation evidence; (c) if indeterminate, preserve original with unchanged confidence.
-
What it enables: The AI can correct factual errors in its own outputs using quantitative simulation evidence, without over-editing or introducing new unsupported content.
-
Implementation: Sample multiple diverse initial answers (m=5) at high temperature, decompose each, merge claims, then ground only uncertain claims with simulator evidence.
-
What it enables: The AI achieves 30.4% higher informativeness (more unique true claims) and 16.3% higher factuality than standard RAG baselines, because it first maximizes coverage then selectively corrects, rather than conditioning the whole response on retrieved context.
-
Implementation: Use the paper's three-domain benchmark (climate, epidemiology, urban planning) with ground truth verified by simulators and human annotators; evaluate at claim level using F1, AUPR, AUROC across verification budgets.
-
What it enables: The AI can be systematically tested for long-form scientific QA quality, with clear metrics for trade-offs between verification cost and answer accuracy.
-
Answer open-ended scientific questions (e.g.,
What is the projected temperature change for City X under SSP585 with +47% CO2 and +36% CH4?
) by querying the appropriate simulator, extracting parameters, and grounding claims in quantitative simulation outputs. -
Self-correct its own long-form answers at the claim level—identifying which claims are uncertain and verifiable, then updating only those with simulation evidence, without rewriting the entire response.
-
Handle multiple scientific domains (climate, epidemiology, urban planning) through a generalized interface that only requires a lightweight adapter (function schema, parameter ranges, output templates).
-
Operate efficiently—verifying only 15-45% of claims while achieving near-full-verification quality, making it practical for real-time or cost-sensitive applications.
-
Provide more trustworthy answers—reducing hallucinations by grounding quantitative claims in simulator outputs, while maintaining answer coverage and coherence.
-
Adapt to new simulators quickly—just provide a handbook and parameter schema; the retrieval and verification pipeline remains unchanged.
Sources
- Linguistic Calibration of Long-Form Generations
- ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
- Shifting Attention to Relevance: Towards the Predictive Uncertainty Quantification of Free-Form Large Language Models
- Retrieval-Augmented Generation for Large Language Models: A Survey
- Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering
- Generalization through Memorization: Nearest Neighbor Language Models
- Teaching Models to Express Their Uncertainty in Words
- Adapting While Learning: Grounding LLMs for Scientific Problems with Intelligent Tool Usage Adaptation
- LLM and Simulation as Bilevel Optimizers: A New Paradigm to Advance Physical Scientific Discovery
- DiscoveryBench: Towards Data-Driven Discovery with Large Language Models
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models
- ClimaQA: An Automated Evaluation Framework for Climate Question Answering Models
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation
- Language Models with Conformal Factuality Guarantees
- Multi-Fidelity Residual Neural Processes for Scalable Surrogate Modeling
- Towards Scientific Intelligence: A Survey of LLM-based Scientific Agents
- ClimateGPT: Towards AI Synthesizing Interdisciplinary Research on Climate Change
- Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback
- SciQAG: A Framework for Auto-Generated Science Question Answering Dataset with Fine-grained Evaluation
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering