SimulCost: A Cost-Aware Benchmark and Toolkit for Automating Physics Simulations with LLMs

summary

Video file (mp4)

The gist

SimulCost: A Cost-Aware Benchmark for Automating Physics Simulations with LLMs Abstract "a static benchmark and extensible toolkit to facilitate research on improving cost-aware agentic designs for

In short

The episode discusses 'SimulCost,' a benchmark and toolkit for automating physics simulations with LLMs. The hosts detail how LLMs struggle with cost-aware parameter tuning, finding that they are inaccurate for high-precision tasks but can be useful for low-accuracy previews. The paper emphasizes the need to build agents that respect resource constraints.

Key concepts

Cost-Aware Benchmark
This benchmark measures both the success of an AI simulation task and the computational cost, such as FLOPs or wall-clock time. It moves beyond simple accuracy checks to assess efficiency, which is crucial in physics simulations where computation can be very expensive.
Single-Round vs. Multi-Round Mode
In single-round mode, the AI must guess the correct parameter value on the first try, with success rates being low. In multi-round mode, where the AI can perform trial-and-error learning from feedback, success rates significantly increase but require much more computation.
Ablation Studies
These studies test specific components of the system, such as in-context learning or reasoning effort. The findings showed that providing examples or asking the AI to think harder did not improve performance, suggesting the issue is a lack of proper grounding in physics knowledge rather than just model size.
Tool Calling vs. Reasoning
The hosts conclude that for high-stakes work, LLMs should not be left to reason on their own for parameter selection. Instead, they are better used when the AI calls a search algorithm to find the answer, as this combination is more cost-effective than pure reasoning.

Terminology used across episodes

This episode discusses

The paper

SimulCost: A Cost-Aware Benchmark and Toolkit for Automating Physics Simulations with LLMs · Read on arXiv

Yadi Cao, Sicheng Lai, Jiahe Huang, Yang Zhang, Zach Lawrence, Rohan Bhakta, Izzy F. Thomas, Mingyun Cao, Chung-Hao Tsai, Zihao Zhou, Yidong Zhao, Hao Liu, Alessandro Marinoni, Alexey Arefiev, Rose Yu

University of California San Diego · The Chinese University of Hong Kong, Shenzhen · Peking University · University of California, Los Angeles · ETH Zurich · California Institute of Technology

Evaluating LLM agents for scientific tasks has focused on token costs while ignoring tool-use costs like simulation time and experimental resources. As a result, metrics like pass@k become impractical under realistic budget constraints. To address this gap, we introduce SimulCost, the first benchmark targeting cost-sensitive parameter tuning in physics simulations. SimulCost compares LLM tuning cost-sensitive parameters against traditional scanning approach in both accuracy and computational cost, spanning 2,643 single-round (initial guess) and 2,304 multi-round (adjustment by trial-and-error) tasks across 11 simulators from fluid dynamics, solid mechanics, and plasma physics, whose costs are analytically defined and platform-independent. A twelfth simulator, a production plasma code measurable only by wall clock, is reported separately. Frontier LLMs achieve 45-62% success rates in single-round mode, dropping to 34-50% under high accuracy requirements, rendering their initial guesses unreliable especially for high accuracy tasks. Multi-round mode improves rates to 66-81%, but LLMs are 1.5-2.7x slower than traditional scanning, making them uneconomical choices. We also investigate parameter group correlations for knowledge transfer potential, and the impact of in-context examples and reasoning effort, providing practical implications for deployment and fine-tuning. We open-source SimulCost as a static benchmark and extensible toolkit to facilitate research on improving cost-aware agentic designs for physics simulations, and for expanding new simulation environments. Code and data are available at https://github.com/Rose-STL-Lab/SimulCost-Bench

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "SimulCost: A Cost-Aware Benchmark and Toolkit for Automating Physics Simulations with LLMs".

Jane: The paper was written by Yadi Cao, Sicheng Lai, Jiahe Huang, Yang Zhang, Zach Lawrence et al. from University of California San Diego and The Chinese University of Hong Kong, Shenzhen and Peking University and University of California, Los Angeles and ETH Zurich and California Institute of Technology.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone. Today we’re digging into a fresh arXiv paper called *SimulCost: A Cost-Aware Benchmark for Automating Physics Simulations with LLMs*. Jane, I gotta say, the title alone got me excited—it’s tackling something we don’t hear enough about in the AI-for-science space.

Jane: Absolutely, Tom. And I think the key word there is "cost-aware." Most benchmarks for AI agents just ask, "Did you get the right answer?" They ignore how much compute or time it took to get there. This paper says, hey, in physics simulations, the cost of running the simulation is often way bigger than the cost of the AI thinking.

Tom: Right, and that’s a huge blind spot. You could have an AI that nails the answer but only after burning a supercomputer’s worth of time. That’s not useful in the real world. So SimulCost is trying to measure both success and efficiency.

Jane: Exactly. They built a benchmark with twelve different physics simulators—covering fluid dynamics, solid mechanics, and plasma physics. And for eleven of them, they can actually count the cost analytically, like a formula for how many operations the simulation runs. That’s clever because it’s platform-independent.

Tom: So you’re saying they don’t just time it on one machine? Because that would be hard to reproduce.

Jane: Precisely. They count the FLOPs, the core operations, which is a much more stable measure. The twelfth simulator, a production plasma code called EPOCH, is too complex for that, so they report it separately using wall-clock time.

Tom: Got it. And the authors—this is a big team from UCSD, CUHK-Shenzhen, Peking University, UCLA, ETH Zurich, and Caltech. Led by Yadi Cao and Rose Yu. They’ve got a lot of domain experts on board, which shows in the detail.

Jane: It really does. They have twelve experts from physics and mechanical engineering helping design the tasks. That’s not something you can fake.

Tom: So what’s the big picture here? Why should listeners care about a benchmark for tuning simulation parameters?

Jane: Because simulations are everywhere—designing airplanes, predicting weather, modeling fusion reactors. If we want AI to help scientists and engineers, we need to know if it can do the job without wasting resources. This paper is the first real attempt to measure that.

Tom: And I bet the results are eye-opening. Let’s get into what they actually found in the next segment.

Jane: You bet. Stay tuned.

Summary: Tom: We’re back with *SimulCost* and, Jane, the results are honestly a bit sobering. They tested five frontier LLMs—GPT-five Claude-three point seven-Sonnet, Llama-three-70B, Qwen3-32B, and GPT-OSS-120B—on over two thousand six hundred single-round tasks and two thousand three hundred multi-round tasks.

Jane: And in single-round mode, where the AI has to just guess the right parameter value on the first try, the best model only hit about sixty-two percent success. That means even the best AI is wrong nearly four times out of ten.

Tom: Yeah, and it gets worse when you demand high accuracy. Success rates drop to between thirty-four percent and fifty percent. So if you need a precise simulation, the AI’s initial guess is basically a coin flip.

Jane: That’s a strong finding. It tells us that LLMs don’t have reliable physics intuition for picking parameters like grid resolution or time step size. They might know the general ballpark, but not the sweet spot that balances accuracy and cost.

Tom: But here’s the interesting part—when they let the AI do multiple rounds, like trial-and-error, success jumps to sixty-six percent to eighty-one percent. So the AI can learn from feedback. The problem is, it’s slow. It takes one point five to two point seven times more compute than just doing a brute-force scan.

Jane: That’s the paradox, right? The AI is better at finding the answer, but it’s not better at finding it cheaply. And that’s a big deal because in the real world, you don’t just want the right answer—you want it before the project deadline.

Tom: So what’s the takeaway for practitioners? If you need a quick preview, an LLM’s guess might be fine for low accuracy. But for high-stakes work, you’re better off letting the AI call a search algorithm instead of reasoning on its own.

Jane: Exactly. And they also compared against Bayesian optimization, which is a classic method for this kind of problem. The LLMs actually matched its success rate, but they were more efficient at low accuracy because they use physics knowledge to start in a good region.

Tom: So LLMs aren’t useless—they’re just not cost-effective when left to their own devices. That’s a nuanced result, and it points to a clear direction: build agents that know when to reason and when to call a tool.

Jane: Right. And that’s exactly what the paper’s ablations explore. We’ll get into those next.

Tom: Can’t wait. Let’s take a quick break.

Improvements: Tom: Welcome back. So, Jane, we talked about the main results, but this paper goes deeper. They ran a bunch of ablation studies to figure out what could actually improve these AI agents. And the findings are pretty counterintuitive.

Jane: Oh, the in-context learning one really got me. You’d think showing the AI examples of past successful simulations would help, right? And it does—in single-round mode, success goes up by seventeen percent to twenty-nine percent. But in multi-round mode, it actually hurts.

Tom: Yeah, that’s wild. The examples anchor the AI to the parameter values it saw, so it stops exploring. It’s like giving someone a map but then they refuse to look for a better route because the map shows one path.

Jane: And the worst part is, even when the examples include cost information, the AI doesn’t use it well. They tested a variant without cost data, and the success was similar, but the efficiency didn’t improve. So just telling the AI "this was cheap" isn’t enough to make it think about cost.

Tom: So what does that mean for people building these systems? A lot of folks are betting on retrieval-augmented generation—just stuff the prompt with relevant examples. This paper says that’s not a complete solution. It might even backfire in interactive settings.

Jane: Exactly. They also tested reasoning effort—like telling GPT-five to think harder before answering. And you know what? It made no difference. More reasoning didn’t lead to better parameter choices.

Tom: That’s a kick in the gut for anyone hoping we just need bigger models or more thinking time. The bottleneck isn’t reasoning depth; it’s that the AI doesn’t have the right grounding. It’s guessing from memory, not deriving from physics.

Jane: Right. And they also looked at whether you could transfer knowledge between simulators. Like, if you fine-tune on a cheap fluid dynamics solver, does that help with an expensive plasma solver? The answer is no—there’s no correlation between tasks, even within the same parameter type.

Tom: So each solver is its own beast. That limits the promise of fine-tuning on cheap simulators to save money on expensive ones.

Jane: But there’s one bright spot. When they let the LLM provide an initial guess for a search algorithm, success went up by eight point six percent. The problem is the cost doubled because the AI’s guesses are too conservative—it picks safer, more expensive values.

Tom: So the AI is good at getting you into the right neighborhood, but not at finding the cheapest house on the block. That’s a really practical insight for designing hybrid systems.

Jane: Exactly. The future isn’t pure LLM reasoning or pure search—it’s knowing how to combine them. And that’s what makes this benchmark so valuable. It gives us a way to measure progress on exactly that.

Tom: Alright, let’s wrap this up with some big-picture thoughts in our final segment.

Conclusion: Tom: And we’re back for the final stretch on *SimulCost*. Jane, I think we’ve covered a lot, but let’s pull it all together. What’s the one thing you want listeners to remember?

Jane: I think it’s that cost-awareness is not a nice-to-have—it’s the core problem. This paper shows that LLMs can be accurate, but they’re not economical. And in scientific computing, economy is everything. Nobody has infinite compute.

Tom: Right. And the benchmark itself is a gift to the community. They open-sourced everything—the solvers, the cost formulas, the task generator. So other researchers can add new simulators and test their own agents.

Jane: That’s the extensible toolkit part. It’s not just a static dataset; it’s a platform. You can create new physics environments and immediately have a cost-aware evaluation setup. That’s going to accelerate a lot of research.

Tom: And I love that they included EPOCH, a real production plasma code. That grounds the benchmark in reality. It’s not just toy problems; it’s something people actually use for fusion research.

Jane: Exactly. And their practical recommendations are so clear. Use LLMs for quick, low-accuracy previews. For high-accuracy work, let them call a search algorithm. Don’t rely on in-context examples for multi-round tasks. And don’t expect fine-tuning to transfer between simulators.

Tom: So what’s the big impact? I think this paper is going to change how we evaluate AI for science. It’s not enough to be right; you have to be right at a price you can afford.

Jane: And that’s a cultural shift. It pushes the field toward building agents that respect resource constraints, which is what real engineers and scientists face every day. It’s a more honest evaluation.

Tom: Well said. So, *SimulCost*—a benchmark that’s going to make AI for physics a lot more practical. Thanks to the authors for putting it out there.

Jane: And thanks to all of you for listening. We’ll be back next time with another paper. Until then, keep your simulations accurate—and your costs low.

Tom: See you, everyone.

More episodes

← Home