SimulCost: A Cost-Aware Benchmark and Toolkit for Automating Physics Simulations with LLMs
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "SimulCost: A Cost-Aware Benchmark and Toolkit for Automating Physics Simulations with LLMs".
Jane: The paper was written by Yadi Cao, Sicheng Lai, Jiahe Huang, Yang Zhang, Zach Lawrence et al. from University of California San Diego and The Chinese University of Hong Kong, Shenzhen and Peking University and University of California, Los Angeles and ETH Zurich and California Institute of Technology.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we’re digging into a fresh arXiv paper called *SimulCost: A Cost-Aware Benchmark for Automating Physics Simulations with LLMs*. Jane, I gotta say, the title alone got me excited—it’s tackling something we don’t hear enough about in the AI-for-science space.
Jane: Absolutely, Tom. And I think the key word there is "cost-aware." Most benchmarks for AI agents just ask, "Did you get the right answer?" They ignore how much compute or time it took to get there. This paper says, hey, in physics simulations, the cost of running the simulation is often way bigger than the cost of the AI thinking.
Tom: Right, and that’s a huge blind spot. You could have an AI that nails the answer but only after burning a supercomputer’s worth of time. That’s not useful in the real world. So SimulCost is trying to measure both success and efficiency.
Jane: Exactly. They built a benchmark with twelve different physics simulators—covering fluid dynamics, solid mechanics, and plasma physics. And for eleven of them, they can actually count the cost analytically, like a formula for how many operations the simulation runs. That’s clever because it’s platform-independent.
Tom: So you’re saying they don’t just time it on one machine? Because that would be hard to reproduce.
Jane: Precisely. They count the FLOPs, the core operations, which is a much more stable measure. The twelfth simulator, a production plasma code called EPOCH, is too complex for that, so they report it separately using wall-clock time.
Tom: Got it. And the authors—this is a big team from UCSD, CUHK-Shenzhen, Peking University, UCLA, ETH Zurich, and Caltech. Led by Yadi Cao and Rose Yu. They’ve got a lot of domain experts on board, which shows in the detail.
Jane: It really does. They have twelve experts from physics and mechanical engineering helping design the tasks. That’s not something you can fake.
Tom: So what’s the big picture here? Why should listeners care about a benchmark for tuning simulation parameters?
Jane: Because simulations are everywhere—designing airplanes, predicting weather, modeling fusion reactors. If we want AI to help scientists and engineers, we need to know if it can do the job without wasting resources. This paper is the first real attempt to measure that.
Tom: And I bet the results are eye-opening. Let’s get into what they actually found in the next segment.
Jane: You bet. Stay tuned.
Summary: Tom: We’re back with *SimulCost* and, Jane, the results are honestly a bit sobering. They tested five frontier LLMs—GPT-five Claude-three point seven-Sonnet, Llama-three-70B, Qwen3-32B, and GPT-OSS-120B—on over two thousand six hundred single-round tasks and two thousand three hundred multi-round tasks.
Jane: And in single-round mode, where the AI has to just guess the right parameter value on the first try, the best model only hit about sixty-two percent success. That means even the best AI is wrong nearly four times out of ten.
Tom: Yeah, and it gets worse when you demand high accuracy. Success rates drop to between thirty-four percent and fifty percent. So if you need a precise simulation, the AI’s initial guess is basically a coin flip.
Jane: That’s a strong finding. It tells us that LLMs don’t have reliable physics intuition for picking parameters like grid resolution or time step size. They might know the general ballpark, but not the sweet spot that balances accuracy and cost.
Tom: But here’s the interesting part—when they let the AI do multiple rounds, like trial-and-error, success jumps to sixty-six percent to eighty-one percent. So the AI can learn from feedback. The problem is, it’s slow. It takes one point five to two point seven times more compute than just doing a brute-force scan.
Jane: That’s the paradox, right? The AI is better at finding the answer, but it’s not better at finding it cheaply. And that’s a big deal because in the real world, you don’t just want the right answer—you want it before the project deadline.
Tom: So what’s the takeaway for practitioners? If you need a quick preview, an LLM’s guess might be fine for low accuracy. But for high-stakes work, you’re better off letting the AI call a search algorithm instead of reasoning on its own.
Jane: Exactly. And they also compared against Bayesian optimization, which is a classic method for this kind of problem. The LLMs actually matched its success rate, but they were more efficient at low accuracy because they use physics knowledge to start in a good region.
Tom: So LLMs aren’t useless—they’re just not cost-effective when left to their own devices. That’s a nuanced result, and it points to a clear direction: build agents that know when to reason and when to call a tool.
Jane: Right. And that’s exactly what the paper’s ablations explore. We’ll get into those next.
Tom: Can’t wait. Let’s take a quick break.
Improvements: Tom: Welcome back. So, Jane, we talked about the main results, but this paper goes deeper. They ran a bunch of ablation studies to figure out what could actually improve these AI agents. And the findings are pretty counterintuitive.
Jane: Oh, the in-context learning one really got me. You’d think showing the AI examples of past successful simulations would help, right? And it does—in single-round mode, success goes up by seventeen percent to twenty-nine percent. But in multi-round mode, it actually hurts.
Tom: Yeah, that’s wild. The examples anchor the AI to the parameter values it saw, so it stops exploring. It’s like giving someone a map but then they refuse to look for a better route because the map shows one path.
Jane: And the worst part is, even when the examples include cost information, the AI doesn’t use it well. They tested a variant without cost data, and the success was similar, but the efficiency didn’t improve. So just telling the AI "this was cheap" isn’t enough to make it think about cost.
Tom: So what does that mean for people building these systems? A lot of folks are betting on retrieval-augmented generation—just stuff the prompt with relevant examples. This paper says that’s not a complete solution. It might even backfire in interactive settings.
Jane: Exactly. They also tested reasoning effort—like telling GPT-five to think harder before answering. And you know what? It made no difference. More reasoning didn’t lead to better parameter choices.
Tom: That’s a kick in the gut for anyone hoping we just need bigger models or more thinking time. The bottleneck isn’t reasoning depth; it’s that the AI doesn’t have the right grounding. It’s guessing from memory, not deriving from physics.
Jane: Right. And they also looked at whether you could transfer knowledge between simulators. Like, if you fine-tune on a cheap fluid dynamics solver, does that help with an expensive plasma solver? The answer is no—there’s no correlation between tasks, even within the same parameter type.
Tom: So each solver is its own beast. That limits the promise of fine-tuning on cheap simulators to save money on expensive ones.
Jane: But there’s one bright spot. When they let the LLM provide an initial guess for a search algorithm, success went up by eight point six percent. The problem is the cost doubled because the AI’s guesses are too conservative—it picks safer, more expensive values.
Tom: So the AI is good at getting you into the right neighborhood, but not at finding the cheapest house on the block. That’s a really practical insight for designing hybrid systems.
Jane: Exactly. The future isn’t pure LLM reasoning or pure search—it’s knowing how to combine them. And that’s what makes this benchmark so valuable. It gives us a way to measure progress on exactly that.
Tom: Alright, let’s wrap this up with some big-picture thoughts in our final segment.
Conclusion: Tom: And we’re back for the final stretch on *SimulCost*. Jane, I think we’ve covered a lot, but let’s pull it all together. What’s the one thing you want listeners to remember?
Jane: I think it’s that cost-awareness is not a nice-to-have—it’s the core problem. This paper shows that LLMs can be accurate, but they’re not economical. And in scientific computing, economy is everything. Nobody has infinite compute.
Tom: Right. And the benchmark itself is a gift to the community. They open-sourced everything—the solvers, the cost formulas, the task generator. So other researchers can add new simulators and test their own agents.
Jane: That’s the extensible toolkit part. It’s not just a static dataset; it’s a platform. You can create new physics environments and immediately have a cost-aware evaluation setup. That’s going to accelerate a lot of research.
Tom: And I love that they included EPOCH, a real production plasma code. That grounds the benchmark in reality. It’s not just toy problems; it’s something people actually use for fusion research.
Jane: Exactly. And their practical recommendations are so clear. Use LLMs for quick, low-accuracy previews. For high-accuracy work, let them call a search algorithm. Don’t rely on in-context examples for multi-round tasks. And don’t expect fine-tuning to transfer between simulators.
Tom: So what’s the big impact? I think this paper is going to change how we evaluate AI for science. It’s not enough to be right; you have to be right at a price you can afford.
Jane: And that’s a cultural shift. It pushes the field toward building agents that respect resource constraints, which is what real engineers and scientists face every day. It’s a more honest evaluation.
Tom: Well said. So, *SimulCost*—a benchmark that’s going to make AI for physics a lot more practical. Thanks to the authors for putting it out there.
Jane: And thanks to all of you for listening. We’ll be back next time with another paper. Until then, keep your simulations accurate—and your costs low.
Tom: See you, everyone.
Yadi Cao, Sicheng Lai, Jiahe Huang, Yang Zhang, Zach Lawrence, Rohan Bhakta, Izzy F. Thomas, Mingyun Cao, Chung-Hao Tsai, Zihao Zhou, Yidong Zhao, Hao Liu, Alessandro Marinoni, Alexey Arefiev, Rose Yu
University of California San Diego · The Chinese University of Hong Kong, Shenzhen · Peking University · University of California, Los Angeles · ETH Zurich · California Institute of Technology
physics.comp-ph, cs.AI, cs.DC, cs.LG
Submitted: 2026-08-17
Updated: 2026-08-18
Comments: post conference revision version at ICML; update: removed CGYRO due to bug in cases search. Will add back soon
Code: https://github.com/Rose-STL-Lab/SimulCost-Bench
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 71/100
The gist: SimulCost: A Cost-Aware Benchmark for Automating Physics Simulations with LLMs Abstract "a static benchmark and extensible toolkit to facilitate research on improving cost-aware agentic designs for
Key concepts
- Cost-Aware Benchmark
- This benchmark measures both the success of an AI simulation task and the computational cost, such as FLOPs or wall-clock time. It moves beyond simple accuracy checks to assess efficiency, which is crucial in physics simulations where computation can be very expensive.
- Single-Round vs. Multi-Round Mode
- In single-round mode, the AI must guess the correct parameter value on the first try, with success rates being low. In multi-round mode, where the AI can perform trial-and-error learning from feedback, success rates significantly increase but require much more computation.
- Ablation Studies
- These studies test specific components of the system, such as in-context learning or reasoning effort. The findings showed that providing examples or asking the AI to think harder did not improve performance, suggesting the issue is a lack of proper grounding in physics knowledge rather than just model size.
- Tool Calling vs. Reasoning
- The hosts conclude that for high-stakes work, LLMs should not be left to reason on their own for parameter selection. Instead, they are better used when the AI calls a search algorithm to find the answer, as this combination is more cost-effective than pure reasoning.
Terminology
Summary
SimulCost: A Cost-Aware Benchmark for Automating Physics Simulations with LLMs
Abstract
"a static benchmark and extensible toolkit to facilitate research on improving cost-aware agentic designs for physics simulations, and for expanding new simulation environments. Code and data are available at https://github.com/Rose-STL-Lab/SimulCost-Bench."
"Evaluating LLM agents for scientific tasks has focused on token costs while ignoring tool-use costs like simulation time and experimental resources. As a result, metrics like pass@k become impractical under realistic budget constraints. To address this gap, we introduce S IMUL C OST, the first benchmark targeting cost-sensitive parameter tuning in physics simulations. S IMUL C OST compares LLM tuning cost-sensitive parameters against traditional scanning approach in both accuracy and computational cost, spanning 2,643 single-round (initial guess) and 2,304 multi-round (adjustment by trial-and-error) tasks across 11 simulators from fluid dynamics, solid mechanics, and plasma physics, whose costs are analytically defined and platform-independent. A twelfth simulator, a production plasma code measurable only by wall clock, is reported separately. Frontier LLMs achieve 45–62% success rates in single-round mode, dropping to 34–50% under high accuracy requirements, rendering their initial guesses unreliable especially for high accuracy tasks. Multi-round mode improves rates to 66–81%, but LLMs are 1.5–2.7× slower than traditional scanning, making them uneconomical choices. We also investigate parameter group correlations for knowledge transfer potential, and the impact of in-context examples and reasoning effort, providing practical implications for deployment and fine-tuning. We open-source S IMUL C OST as a static benchmark and extensible toolkit to facilitate research on improving cost-aware agentic designs for physics simulations, and for expanding new simulation environments."
Introduction
"Large Language Models (LLMs) show promise for scientific workflows through code generation, complex reasoning, and especially tool calling... By offloading computations to domain-specific tools, LLM agents reduce hallucinations through grounded outputs. This capability has accelerated workflows in scientific domains including chemistry, computational fluid dynamics, and data science."
"However, existing evaluations focus on task correctness and token costs while overlooking tool costs. Metrics like pass@k with large k implicitly treat tool usage as free, which becomes impractical in realistic scientific workflows where simulations or experiments consume significant time or materials. In physics simulations, the numerical parameter choices directly impact both solution quality and cost... Increasing spatial or temporal resolution improves accuracy but with cost scaling quadratically or cubically. Domain experts develop intuition for these trade-offs through experience, balancing cost against quality via informed guesses and iterative refinement. Yet without mechanisms to evaluate such cost-awareness in LLMs under current pass@k and success-rate-only metrics, we risk developing agents that achieve correctness only after numerous unaffordable trials."
"To address this gap, we present S IMUL C OST, the first benchmark designed to evaluate LLMs’ cost-awareness in tuning physics simulators. Unlike existing benchmarks that only measure success rate, S IMUL C OST additionally evaluates computational efficiency. We quantify cost by counting core operation FLOPs for each simulation instance, which correlates with wall-clock time on serial machines. Our benchmark spans 12 simulators across fluid dynamics, heat transfer, solid mechanics, and plasma physics; 11 of them admit such an analytical cost and carry the main results, while the twelfth (EPOCH, a production plasma code) is timed by wall clock and reported separately."
Main Contributions
"(1) the first benchmark measuring both success rate and computational cost for LLM-automated physics simulations; (2) an extensible toolkit with 12 simulators, 11 of them featuring platform-independent cost tracking; (3) evaluation of state-of-the-art LLMs compared against brute-force scanning and Bayesian optimization; (4) ablation studies providing practical guidance on potential improvements such as knowledge transfer, in-context learning, and reasoning effort."
Key Findings
"(1) Frontier LLMs achieve 45–62% success rates in single-round mode, dropping to 34–50% under high accuracy requirements. This indicates that LLMs’ initial guesses are unreliable and only useful for quick previews when users lack parameter intuition. (2) Multi-round mode improves success rates to 66–81%, making it necessary for high-accuracy tasks. However, LLM trial-and-error is 1.5–2.7× slower than brute-force scanning, suggesting practitioners should let LLMs invoke scanning algorithms rather than relying on internal reasoning alone. (3) Common parameters like spatial resolution are easier to tune than solver-specific ones like convergence tolerances. However, there is little correlation between individual tunable parameters even within the same type, indicating that fine-tuning on cheaper simulators is unlikely to transfer to expensive ones. (4) In-context learning improves single-round success by 17–29% but anchors models to demonstrated parameter regimes, degrading multi-round performance. This indicates naive retrieval-augmented generation may not be a complete solution."
Related Work
"LLM benchmarks for science fall into two categories: scientific knowledge evaluation and tool-using agents. The first category includes QA benchmarks such as SciBench, SciEx, and APBench... These benchmarks are limited to problems with closed-form, hand-calculatable solutions. Realistic scientific workflows require complex external tools like numerical simulations, motivating the second category: LLMs as scientific agents."
"Despite growing interest in scientific agents, most benchmarks neglect the cost of tool execution, a critical factor in real scientific workflows where tool costs (such as simulation time) dominate over LLM API costs. For example, ScienceAgentBench, DiscoveryBench, and most agentic systems track only LLM token costs. AstaBench claims 'computational cost' but reports only token costs in USD. Works such as CodePDE and SciCode combine code generation with pass@K metrics. While valuable for assessing code synthesis capabilities, these benchmarks address a different challenge than properly using existing tools, and do not consider execution cost as a metric."
"Several benchmarks address this gap by recording wall-time, e.g., Agent-Lab, MLE-Bench, MLGym, and BioML-bench. However, these measurements conflate LLM inference with tool execution time and vary across hardware, making cross-study comparisons unreliable. Auto-Bench achieves platform-independence by counting interventions as the tool cost, but this assumes uniform cost per intervention and remains in the data science domain. A related line of work adapts whether to call a tool at all, deciding by problem difficulty and trading success rate against efficiency, but the choice is binary and leaves the tool’s own cost unmeasured. CATP-LLM models cost-aware tool planning via function-as-a-service, but is limited to generic tools with relatively fixed costs per function, which breaks down for simulations where parameters like grid resolution alter cost of the same function by orders of magnitude. MetaOpenFOAM is the only prior work targeting physics simulations with cost awareness. However, it is a workflow automation framework rather than a comprehensive benchmark, and records wall-time that conflates LLM inference with solver execution."
"S IMUL C OST bridges these gaps by defining cost based on computational complexity analysis of each solver, providing platform-independent measurements with guaranteed reproducibility. We focus on parameter tuning of physics simulations where the parameters of interest fundamentally influence both simulation accuracy and computational cost."
The SimulCost Toolkit
"Our 12 simulators span fluid dynamics, solid mechanics, and plasma physics. We evaluate both single-round inference, assessing physics and numerical intuition, and multi-round inference, testing adaptive parameter optimization through solver interaction."
"Each solver presents LLMs with parameter tuning tasks: given a physics-based simulation scenario, the LLM must select parameter values that satisfy accuracy requirements while minimizing computational cost. In reality, simulators usually require tuning multiple parameters for a successful execution. However, domain experts usually isolate potential issues upon seeing unsuccessful simulation runs and tune the corresponding parameters one by one. We adopt this pattern and constrain the LLM to tune individual parameters while fixing others to reasonable (but non-optimal) values chosen by rule-of-thumb practices (e.g., CFL= 0.25 for explicit time-stepping). This isolation avoids the complexity of multi-parameter optimization and enables a meaningful grid search (scan) baseline. For each solver, we define the tool cost using computational complexity analysis by counting dominant operations FLOPs, which strongly correlates with wall time on a single-core processor."
Dataset Curation
"Our curation pipeline has four phases: (1) Solver development - developing new or adapting existing solvers, (2) Reference solution search - finding reference parameters for each task using brute-force search, (3) Variation design - scaling up task diversity by combining accuracy thresholds (low/medium/high), initial/boundary conditions, environments, and non-target parameter settings, and (4) Filtering - since experts define the accuracy thresholds consistently for all combinations, some may fail to find a solution that meets the threshold within the maximum scan iterations. We filter these infeasible tasks post-hoc to avoid unnecessarily high computational time for later LLM evaluations."
"All solvers are sourced from textbooks, online repositories, and research papers. Original scripts typically contain bugs or lack cost-related metadata, requiring expert debugging and adaptation for quality assurance, cost calculation, and consistent API design."
Inference Modes
"Expert parameter tuning in physics-based simulations requires two distinct capabilities: (1) physical and numerical intuition for initial parameter estimation and (2) iterative adjustment through interactions with solvers. We design two inference modes to evaluate these skills, measuring both success rate and cost efficiency:
Single-Round Inference requires single-attempt parameter selection using LLM’s prior knowledge to find the optimal parameter balancing accuracy and cost. Since the true optimal in continuous parameter space is unknown, we set the reference as the near-optimal solution with near-minimal cost that meets the accuracy threshold found by scan algorithm.
Multi-Round Inference allows up to 10 trials. Human expertise usually only has patience for 5–10 trials before resorting to systematic sweeps, and our scan baseline uses approximately 20 grid points, so 10 trials gives LLMs half the budget of exhaustive search. Each trial’s simulation feedback includes: a) convergence status, b) RMSE between the current and a finer resolution to judge convergence, and c) accumulated cost is appended to the conversation list. LLMs can terminate early if they deem a solution satisfactory."
Evaluation Metrics
"Success Rate. Success Si ∈ 0, 1 for task i indicates whether the task-dependent distance between the LLM-proposed simulation output and the reference output (found by Algorithms 1 or 2) falls within the accuracy threshold."
"Efficiency. Efficiency measures how well the LLM’s cost compares to the reference cost: Ei = (Cibf / Cisr,mr) × Si, where Cibf is the brute-force reference cost, Cisr,mr is the LLM’s cost, and superscript 'sr' or 'mr' denotes single-round or multi-round mode respectively. For single-round, Cibf,sr is the near minimum-cost solution; for multi-round, Cibf,mr is the accumulated scan cost."
"For aggregated results, we report arithmetic mean for success rate and geometric mean for efficiency over successful samples only (Ei > 0). The geometric mean is standard practice for aggregating performance ratios and speedups: it is scale-invariant, handles values spanning multiple orders of magnitude without being dominated by outliers, and treats multiplicative relationships symmetrically. Failed tasks (Ei = 0) are excluded from efficiency aggregation."
Experiments
"We evaluate state-of-the-art LLMs including GPT-5-2025-08-07, Claude-3.7-Sonnet-2025-02-19, GPT-OSS-120B, Llama-3-70B-Instruct, and Qwen3-32B across the 11 solvers with analytical costs spanning fluid dynamics, solid mechanics, and plasma physics. Our evaluation encompasses both single-round and multi-round inference modes at three accuracy levels (low, medium, high). Temperature is set to 0.0 for all models except GPT-5, which uses its default setting."
Main Results
"In single-round mode, success rates vary substantially across models (45–62%), with GPT-5 leading at 61.5%. Even this best rate falls short of practical reliability, requiring manual intervention every 2–3 simulations. Multi-round inference substantially improves all models, with GPT-5 reaching 80.8% and GPT-OSS-120B 77.1%, and even Llama-3-70B reaching 71%."
Regarding performance across different accuracy levels, single-round success drops sharply as accuracy requirements tighten (64% → 39%), while multi-round remains stable (73% → 71%).
"Efficiency interpretation differs by mode: in single-round, efficiency >1.0 means the LLM beat the optimal solution’s cost. In multi-round, the reference is cumulative brute-force cost, so efficiency ≈1.0 merely matches naive search. In single-round, all models achieve 0.16–0.52 efficiency, using 2–6× optimal compute. In multi-round, only GPT-OSS-120B reaches 0.86 (at low accuracy), while other models cluster around 0.3–0.7, taking 1.5–2.7× brute-force cost on average."
Single-Round vs. Multi-Round Comparison
"We observe that the success rate gain is most pronounced at high accuracy levels, with a mean improvement of +30.5% across models (p < 0.001), precisely where single-round struggles most. Lower accuracy levels also benefit, but to a lesser extent. This asymmetry has a clear explanation: higher accuracy requirements narrow the acceptable parameter range, making 'lucky guesses' increasingly unlikely. Multi-round mode compensates through trial-and-error, making it necessary for high-accuracy tasks."
Task Group Analysis
"Success rate varies substantially across parameter groups in single-round mode (15–68%). Spatial and Tolerance parameters cluster at the high end, likely benefiting from predictable cost-accuracy trade-offs in pre-training data, while Misc parameters show the widest variation due to their solver-specific nature. Multi-round mode compresses this variation (all groups: 56–79%), with Misc parameters gaining most (+28%). This pattern suggests that iterative exploration compensates for missing prior knowledge, benefiting uncommon parameters the most."
"Efficiency patterns reveal a paradox: in single-round, Misc parameters (the hardest to solve) achieve near-optimal efficiency when successful, while Spatial parameters lag far behind. This suggests LLMs can find runnable values for common parameters like grid resolution but lack cost intuition, tending toward 'safer' values that work but waste compute."
"We analyzed within- vs. between-group task correlations to assess knowledge transfer potential. Neither single-round nor multi-round showed significant within-group correlation advantages (p=0.24 and p=0.88, respectively), implying task difficulty is parameter-specific rather than type-driven. This limits the viability of fine-tuning on cheap simulators to improve performance on expensive ones."
In-Context Learning
"ICL improves single-round success but degrades multi-round performance. This pattern holds across all variants, suggesting that examples help narrow initial guess ranges but limit exploration, steering models to shown parameter regimes."
"Comparing variants in single-round mode: Mixed-Accuracy ICL achieves the largest gains in both success (+28.7%) and efficiency (1.74×), likely because diverse examples span a wider parameter range. Cost-Ignorant ICL shows comparable success gains but minimal efficiency improvement, suggesting that cost information in examples is critical for efficiency."
Bayesian Optimization Baseline
"BO achieves comparable aggregate success rates to LLMs but shows higher inter-solver variance. Efficiency-wise, LLMs have an obvious advantage, especially at low accuracy requirements (2.13 vs. 1.05). We found that BO’s exploration strategy tends to choose extreme values early to maximize information gain. However, this triggers early stopping if it selects the 'finer' side of extreme bounds, resulting in high cumulative cost. In contrast, LLMs leverage physics intuition from pre-training to make more informed initial guesses, which is especially helpful at low accuracy requirements where thresholds are more lenient."
Additional Results
"Reasoning Effort: GPT-5’s reasoning effort parameter shows no significant overall impact, as increased reasoning does not lead to better parameter selection. Failure Modes: We identify five recurring failure patterns (false positives, blind exploration, instruction misunderstanding, prior bias, and conservative strategy) to guide future improvements. These patterns reveal a tension between cost-awareness and correctness: LLMs either skip intermediate resolutions to save cost and land on non-converged solutions (blind exploration), or choose unnecessarily fine resolutions 'to be safe' and then make tiny adjustments that hinder exploration (conservative strategy). Additional Baselines: Classical optimization methods (Bisection, Nelder-Mead) achieve similar success rates to brute-force at comparable cost, justifying our baseline choice. Robustness: Prompt sensitivity analysis (25 variants) and unconditional cost metrics leave our conclusions unchanged. Tool-Augmented Tuning: When LLMs provide initial guesses for search algorithms, success improves (+8.6%) but cost doubles due to conservative parameter choices. Wall-Clock-Cost Solver: EPOCH, whose cost cannot be counted in FLOPs, is evaluated on its own and excluded from every aggregate above."
Limitations
"Our benchmark restricts LLMs to text-based solver interactions without access to auxiliary tools. Real-world agents might leverage log parsing or custom timeout logic, but this constraint ensures fair comparison: many LLMs do not yet support full agentic mode reliably, and the choice of framework introduces confounding variables."
"Additionally, we isolate single-parameter tuning rather than joint multi-parameter optimization. Baseline tractability drives this choice: grid search complexity grows as O(nk) where n is the number of choices per parameter and k is the number of parameters. With n = 20 grid points, single-parameter search requires 20 evaluations, while 3-parameter joint search requires 203 = 8,000, making exhaustive reference solutions prohibitively expensive."
Conclusions and Future Directions
"We introduced S IMUL C OST, the first benchmark for evaluating cost-aware parameter tuning capabilities of LLMs in physics simulations. Our main evaluation spans the 11 physics simulators with analytical costs and 5 state-of-the-art LLMs; EPOCH is reported separately. We also open-source S IMUL C OST upon publication, including the static benchmark and extensible toolkit."
Practical recommendations
"(1) LLMs’ initial guesses are generally unreliable (45–62% success) and only useful for quick previews when users lack parameter intuition, accuracy requirements are low, and optimal cost efficiency is not needed. (2) For high-accuracy tasks or finding a guaranteed cost-efficient configuration, multi-round mode becomes necessary (66–81% success), but practitioners should let LLMs invoke scanning algorithms rather than relying on internal reasoning alone, as the latter is 1.5–2.7× slower. (3) Fine-tuning on cheaper simulators is unlikely to transfer to expensive ones due to the lack of cross-parameter correlation, even within the same parameter type. (4) ICL improves single-round performance but anchors models to demonstrated regimes, degrading multi-round exploration. Hence, naive RAG may not be a complete solution. (5) Using examples from a diverse range of accuracy requirements helps exploration, and including cost information helps optimize cost efficiency. (6) LLM-initialized search should be steered toward intermediate parameter values: conservative initial guesses undershoot the optimum by 50–80%, incurring 5–14× cost overhead."
Future directions
"(1) Tool-augmented tuning: equipping LLMs with timeout-based early stopping, callable search algorithms, and multi-modal feedback such as field visualizations to enable richer decision-making. (2) Human-in-the-loop evaluation: user studies measuring how LLM-suggested parameters accelerate expert workflows, other than autonomously tuning, would validate real-world utility. (3) Cost-aware post-training: developing fine-tuning strategies that explicitly optimize for accuracy and computational efficiency. (4) Multi-parameter optimization: enabling LLMs to jointly tune interdependent parameters with adaptive sampling approaches. (5) Parallel computing: extending cost tracking to parallel architectures with analytical complexity models that account for overhead and delays, or careful, reproducible wall-time measurements. (6) Non-monotonic parameter handling: for parameters without clear cost-accuracy monotonicity, experts either draw on mathematical or physics priors or inspect visualizations after the fact for abnormal behavior such as oscillations near shock fronts."
Improvements for AI systems
Based on the paper's findings, here are specific improvements I can implement in AI systems:
Improvement: Replace pass@k and success-rate-only metrics with a dual metric system that tracks both success rate and computational efficiency (cost ratio vs. brute-force baseline).
What the improved system can do: Evaluate LLM agents on physics simulation tasks while penalizing excessive computational resource usage, preventing agents from succeeding
through wasteful trial-and-error.
Improvement: Implement a hybrid decision mechanism that automatically selects between single-round and multi-round modes based on task accuracy requirements and parameter type.
Improvement: Add explicit cost-minimization instructions and include cost information in any in-context examples.
Improvement: Equip the agent with callable brute-force scanning algorithms rather than relying on internal reasoning for multi-round refinement.
Improvement: Implement a diversity check on retrieved examples to prevent anchoring to demonstrated parameter regimes.
Improvement: Add a post-generation validation step that checks if proposed parameters are within a safe intermediate range
rather than extreme values.
Improvement: Implement different strategies for different parameter groups (Spatial, Temporal, Tolerance, Misc).
Improvement: Add cumulative cost monitoring that triggers early termination when costs exceed a threshold relative to expected brute-force cost.
Improvement: Implement a feedback analysis module that extracts actionable patterns from simulation results before proposing next parameters.
Improvement: Add a task-similarity check before applying fine-tuned knowledge from one simulator to another.
Abstract
Evaluating LLM agents for scientific tasks has focused on token costs while ignoring tool-use costs like simulation time and experimental resources. As a result, metrics like pass@k become impractical under realistic budget constraints. To address this gap, we introduce SimulCost, the first benchmark targeting cost-sensitive parameter tuning in physics simulations. SimulCost compares LLM tuning cost-sensitive parameters against traditional scanning approach in both accuracy and computational cost, spanning 2,643 single-round (initial guess) and 2,304 multi-round (adjustment by trial-and-error) tasks across 11 simulators from fluid dynamics, solid mechanics, and plasma physics, whose costs are analytically defined and platform-independent. A twelfth simulator, a production plasma code measurable only by wall clock, is reported separately. Frontier LLMs achieve 45-62% success rates in single-round mode, dropping to 34-50% under high accuracy requirements, rendering their initial guesses unreliable especially for high accuracy tasks. Multi-round mode improves rates to 66-81%, but LLMs are 1.5-2.7x slower than traditional scanning, making them uneconomical choices. We also investigate parameter group correlations for knowledge transfer potential, and the impact of in-context examples and reasoning effort, providing practical implications for deployment and fine-tuning. We open-source SimulCost as a static benchmark and extensible toolkit to facilitate research on improving cost-aware agentic designs for physics simulations, and for expanding new simulation environments. Code and data are available at https://github.com/Rose-STL-Lab/SimulCost-Bench
Sources
- MetaOpenFOAM: an LLM-based multi-agent framework for CFD
- AI Agents in Engineering Design: A Multi-Agent Framework for Aesthetic and Aerodynamic Car Design
- CodePDE: An Inference Framework for LLM-driven PDE Solver Generation
- DiscoveryBench: Towards Data-Driven Discovery with Large Language Models
- Dominant balance-based adaptive mesh refinement for incompressible fluid flows
- MLGym: A New Framework and Benchmark for Advancing AI Research Agents
- Gorilla: Large Language Model Connected with Massive APIs
- Agent Laboratory: Using LLM Agents as Research Assistants
- ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases
- DataSciBench: An LLM Agent Benchmark for Data Science
- A stabilized march approach to adjoint-based sensitivity analysis of chaotic flows
Related papers
- Learning Spectral-Like Mesh-Free Discretisations
- Breaking the Exascale Barrier for the Electronic Structure Problem in Ab-Initio Molecular Dynamics
- Aitomia: Your Intelligent Assistant for AI-Driven Atomistic and Quantum Chemical Simulations
- Metal-Insulator Transition of Solid Hydrogen by the Antisymmetric Shadow Wave Function
- Resonating Valence Bond Quantum Monte Carlo: Application to the ozone molecule
- Quantifying Gibbs measures of disordered crystals up to the solid-liquid phase transition