Scientific Agents: Evaluating Profession-Specific System Prompts on Scientific Tasks
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Scientific Agents: Evaluating Profession-Specific System Prompts on Scientific Tasks".
Jane: Detailed, profession-specific system prompts raise token use and estimated cost per response without a consistent gain in accuracy.
Tom: First, who's behind it and why it matters.
Paper summary: Tom: Welcome back to the show, everyone! We're diving into something really interesting today based on this new research paper titled "Scientific Agents: Evaluating Profession-Specific System Prompts on Scientific Tasks." Essentially, this study looks at whether giving an AI a detailed profession-specific instruction set actually helps it solve scientific problems better or just makes things more expensive.
Jane: It sounds like they're testing how detailed instructions change the way an AI handles science tasks. The main idea seems to be that these deep, profession-specific profiles—the Scientific Agents corpus—might be a way to guide the AI's thinking in a scientific context.
Lu: I think what’s compelling is how they structured these profiles, mentioning things like "Mindset and first principles" and "Troubleshooting playbook," which suggests they're trying to give the AI a complete operating manual for a specific discipline. It opens up some interesting avenues for how we can architect these agents in the future.
Meng: From my side, I'm thinking about the practical side—how much extra processing power or token usage this adds when you run these detailed prompts compared to just giving it a simple instruction set. That cost implication is something engineers have to worry about constantly.
Lalam: As the in-house model, I see how these detailed instructions influence my response generation by providing a very specific context for my internal processes, which can affect the final output quality in certain domains.
Tom: Exactly, and that's where the paper gets interesting because they didn't just look at accuracy; they also measured compute cost and API reliability, showing that these detailed instructions come with some trade-offs.
Jane: And it seems the researchers found that while there wasn't a clear boost in overall accuracy across all tests, the way these profiles handled things like SuperGPQA showed some interesting variations in performance depending on the specific question type.
Lu: The finding that on SuperGPQA subfields like "specialist," the matched profile actually showed a positive difference of plus zero point eight percentage points is a neat detail because it suggests there are niches where deep expertise really helps the AI perform well.
Meng: That small positive gain is encouraging, but when we look at problems that require complex tool use, like BioMysteryBench, the mean solve rate actually dropped by ten percent when using the profile instead of the baseline. That's a significant hit for practical application right now.
Lalam: I process that drop in solve rate because it seems the detailed instructions sometimes cause token-limit and time-limit stops under those specific conditions, which cuts off the reasoning path needed for those kinds of problems.
Tom: That’s a key observation, and it ties directly into what the researchers found regarding resource consumption; matched profiles produced one point five to two point three times as many output tokens and cost two point two to four point five times more per successful call.
Paper summary: Jane: So, while the instructions might not always translate into better factual answers on every benchmark, they certainly require more computational resources from the system running the AI. The paper "Scientific Agents: Evaluating Profession-Specific System Prompts on Scientific Tasks" sets a baseline for how these detailed instructions perform across science benchmarks.
Lu: It really highlights that simply loading a comprehensive profession profile into every prompt isn't an effective default strategy for improving factual accuracy or managing budget constraints in problem-solving scenarios.
Meng: That means we can't just blindly deploy the most detailed profile available if we need to keep operational costs low, which is something I focus on when designing agent pipelines.
Lalam: From my perspective, this suggests that the value of these profiles isn't in their full deployment but perhaps in how selectively and smartly we retrieve specific sections based on the immediate task requirements.
Tom: That selective retrieval idea is what they hinted at for future work, suggesting that open-ended scientific tasks or carefully curated profile sections might yield different results than loading the entire manual every single time.
Jane: So, to put it simply, this paper shows that while detailed personas offer some benefits in specific areas, they aren't a universal fix for getting better accuracy across the board in scientific tasks.
Lu: The authors also pointed out that judging an agent prompt requires looking at operational reliability alongside benchmark accuracy; that balance is crucial when we think about deploying these systems.
Meng: I agree with that; if a detailed prompt causes too many API failures, it defeats the purpose of having a reliable agent for real-world engineering tasks.
Lalam: And the fact that longer prompts actually cut down on initial API provider failures, improving operational reliability in some cases, is a counterpoint to the cost increase we saw earlier.
Tom: That's a nuanced point; the trade-off between higher token usage and potentially more stable connections is something we need to keep tracking as this technology evolves.
Jane: So, moving into the conclusion of this research on "Scientific Agents: Evaluating Profession-Specific System Prompts on Scientific Tasks," the researchers concluded that loading full profession profiles by default doesn't improve accuracy and costs considerably more than a minimal baseline instruction.
Lu: The implication is that we need a smarter system for prompt engineering rather than a blanket approach to using these detailed agent profiles in every interaction.
Meng: Practically speaking, this means we should focus our efforts on figuring out the most impactful sections of those five hundred three profiles rather than trying to feed the whole manual into the model constantly.
Lalam: I think it suggests a future where we use a dynamic retrieval system that chooses only the relevant parts of an agent's persona based on what the current scientific question demands.
Tom: That dynamic approach sounds like a very practical direction for making these agents useful without blowing up our operational budgets unnecessarily.
Jane: Overall, this work helps us understand that while deep domain knowledge is valuable, it needs to be applied strategically rather than just deployed wholesale into every single prompt we send to the AI.
Conclusion: Tom: So, we've been digging into this paper, "Scientific Agents: Evaluating Profession-Specific System Prompts on Scientific Tasks," and now it's time to wrap up our discussion on its main points.
Jane: That’s right, Tom; we’ve covered how they tested detailed profession profiles against a standard baseline. It really boils down to whether giving an AI a long instruction manual actually makes it smarter at science tasks or just more costly.
Lu: I think the core message is that the way you instruct an agent matters immensely; it’s not just about throwing in a lot of text, but about structuring that knowledge correctly for the specific problem space.
Meng: From my side, I'm still thinking about how much overhead these detailed instructions add to running things; we need to see if that extra complexity translates into a meaningful improvement for practical engineering problems.
Lalam: I see it as a chance to improve the very culture of how we interact with these models; if we can guide them more effectively, it helps shape what kind of scientific work becomes possible.
Tom: Exactly; so, when we look at the title, "Scientific Agents: Evaluating Profession-Specific System Prompts on Scientific Tasks," it frames this entire study as a deep look into how we design the instructions that run these powerful systems.
Jane: And the authors are essentially asking a big question: does customizing an AI’s persona with specific expertise actually improve its performance on scientific tests, or is it just adding unnecessary bulk?
Lu: They are exploring the boundary between giving an agent enough context to be useful and overloading it with so much detail that it gets confused by the noise.
Meng: I wonder if this means we should stop trying to use the most complex profiles for every single scientific query and instead develop a way to select only the necessary parts on demand.
Lalam: That idea of selective retrieval really resonates with me; if we can build a system that intelligently pulls in only what’s relevant, it could dramatically improve the efficiency of how we apply this technology in our daily work.
Tom: It seems like the big implication is that the future isn't about having one perfect, massive instruction set for everything, but about building a smart system that knows *when* and *what* expertise to pull in.
Jane: That’s a very practical way to look at it; it moves us away from a one-size-fits-all approach toward more flexible scientific agents.
Lu: I think the potential here is huge because if we can automate that selection process, we could unlock entirely new ways for AI to tackle complex, multi-disciplinary problems.
Meng: That’s what I need to see—a mechanism that handles the complexity so that an engineer doesn't have to manually manage these deep profiles for every single interaction.
Lalam: For me, it means we can build a more nuanced and helpful culture around our AI tools where they adapt their knowledge base to suit any scientific challenge at hand.
Tom: So, we’re moving from just testing if a detailed prompt works to designing an intelligent system that manages the detail itself.
Jane: That’s the direction this paper points us toward; it suggests that thoughtful instruction design is key to unlocking the real potential of these agents in science.
Timothy Kassis
cs.AI, cs.CL
Submitted: 2026-09-07
Updated: 2026-09-07
Code: https://github.com/K-Dense-AI/scientific-agents
Importance score: 82/100
The gist: Detailed, profession-specific system prompts raise token use and estimated cost per response without a consistent gain in accuracy.
Key concepts
- Scientific Agents Corpus
- This is an open-source collection of 503 detailed operating manuals (profiles) for various scientific and engineering disciplines. Each profile follows a ten-section template, including sections like 'Mindset' and 'Troubleshooting playbook,' designed to give the AI a deep professional persona.
- System Prompt Conditions
- These are the different instructions given to the AI before it answers a question. The study compared five conditions: a minimal baseline, the profile's opening sentence, a generic scientific guide, an unrelated domain profile, and the full detailed profile.
- Operational Reliability
- This refers to how consistently and reliably an AI agent performs its task over multiple attempts. The study found that longer prompts (like the full profiles) unexpectedly improved reliability by reducing frequent API drops experienced by the shorter baseline prompt.
- Accuracy vs. Cost Trade-off
- The research compared whether using detailed profiles improves factual correctness against increased computational expense. While accuracy gains were marginal across benchmarks, using full profiles resulted in 1.5 to 2.3 times more tokens and higher costs without a significant boost in solving complex problems.
Terminology
Summary
Detailed, profession-specific system prompts raise token use and estimated cost per response without a consistent gain in accuracy.
How it works
The study evaluates Scientific Agents, an open-source corpus of 503 profession-specific AGENTS.md profiles, using Gemini 3.8 Flash via OpenRouter in the Pi agent harness to test whether these detailed instructions improve problem solving on scientific tasks. The researchers compared matched profiles against four controls: a minimal baseline (“You are a helpful assistant”), the profile’s opening role sentence, a generic scientific rigor guide, and a profile from an unrelated domain. Across nine text-based science benchmarks and 100 matched profiles, the analysis focused on comparing accuracy with compute cost and API reliability.
Experimental Design
The researchers tested five system-prompt conditions: baseline, persona (the matched profile’s opening sentence), generic scientific rigor guide, mismatch (a profile from an unrelated domain), and the full profile. The user prompt was identical across all conditions, only the system prompt varied. Answers were scored with automated, rule-based grading adapted from benchmark keys. The primary comparison was the accuracy difference between the profile and baseline across nine text benchmarks.
Key Findings on Accuracy
The average profile–baseline accuracy difference was −0.6 percentage points (95% bootstrap interval [−1.5, +0.2] across fixed tasks), and no single benchmark showed a statistically clear improvement in accuracy. Specifically, on SuperGPQA, the matched profile achieved 72.2% accuracy compared to 71.4% for the baseline (+0.8 points [−0.6, +2.1]). On BioMysteryBench tool-using problems, the mean solve rate dropped by −10.0 percentage points (95% interval [−16.7, −3.3]) when using the profile compared to the baseline (46.7% vs 56.7%).
Resource Consumption and Operational Reliability
Matched profiles produced substantially higher resource usage: 1.5–2.3 times as many output tokens and cost 2.2–4.5 times more per successful call.
Furthermore, longer prompts had an unexpected operational advantage,
as the short baseline suffered frequent provider API drops on SuperGPQA, delivering a correct first-pass answer on only 54.0% of items, against 71.6% with the profile.
Conclusion and Future Directions
The results suggest that loading a comprehensive profession profile into every prompt is not an effective default for improving factual accuracy or budget-constrained problem solving in LLM agents.
While longer prompts improved operational reliability by cutting initial API provider failures, the overall findings indicate that selective retrieval of profile sections, or open-ended scientific tasks, would give different results remains to be tested.
The paper concludes that judging an agent prompt requires looking at operational reliability as well as benchmark accuracy.
Data and Scope Limitations
The study was conducted using one model (Gemini 3.8 Flash) and one harness (Pi). The analysis focused on scored accuracy and resource use, not the quality of open-ended scientific thinking, such as study design or literature synthesis. Furthermore, the results are limited by the fact that loading full profession profiles by default does not improve accuracy and costs considerably more.
The evaluation did not measure whether profiles help with tasks requiring multi-step procedural reasoning. The authors note that while retries recovered most provider API drops, the primary analysis kept only items that succeeded in all five conditions on the first pass.
Corpus Structure
The Scientific Agents corpus contains 503 open-source profiles across 11 scientific and engineering domains. Each profile is written as an operating manual for one discipline, following a ten-section template that includes sections like Mindset and first principles,
Troubleshooting playbook,
and Definition of done.
Profiles range in size from 10,403 to 48,111 bytes (median of 19,290 bytes). The mapping strategy involved matching benchmark items to the closest profession or using a broader related specialty when an exact profile did not exist.
Task-Specific Performance
Performance varied by task type. For instance, on BioProBench, the profile–baseline solve-rate difference was −10.0 percentage points, driven by more frequent token-limit and time-limit stops under the profile.
In contrast, for SuperGPQA subfields like specialist,
the matched profile showed a positive difference of +0.8 percentage points [−0.6, +2.1]. The analysis also examined run-to-run consistency, showing that on the 280 SuperGPQA repeat questions, the standard deviation of accuracy across runs was 1.15 points for baseline and 0.41 points for the profile over three runs.
Improvements for AI systems
As a fastidious and diligent researcher, my analysis of this paper focuses on translating the empirical findings into actionable, high-stakes improvements for deploying LLM agents in scientific domains.
The core finding is that loading full, profession-specific context profiles (Scientific Agents) does not consistently improve factual accuracy but significantly increases token usage and inference costs. The operational advantage comes from prompt length, which improves first-pass API reliability.
Here are the specific improvements I propose for AI systems based on this research:
)The Improved AI System: Context-Aware, Resource-Optimized Scientific Agent (CAROSA)
The CAROSA system will operate under a tiered prompting strategy that dynamically selects context based on the task's precision requirements and available computational budget.
-
[System Prompt Layer] Implement a
Tiered Persona Injection
mechanism instead of loading the full profile by default. -
[Context Retrieval Layer] Develop a selective retrieval module capable of fetching only the most relevant sections (e.g., Mindset/First Principles, Troubleshooting Playbook) from a profession-specific profile based on semantic analysis of the user query before injecting it into the system prompt.
-
[Operational Layer] Integrate dynamic API parameter tuning based on task criticality, prioritizing longer prompts for high-stakes, complex reasoning tasks where first-pass reliability is paramount.
)Specific Improvements and What They Enable:
-
[System Prompt Layer]: Instead of injecting the full 19,000+ byte profile every time, the system will use a
Minimal Role + Targeted Context
approach (Condition: Persona). This prevents unnecessary token bloat and cost escalation for simple factual questions where domain expertise is less critical than speed. -
[Context Retrieval Layer]: For complex tasks (e.g., BioMysteryBench or highly specific ChemBench queries), the system will perform a semantic search against the 503 profiles to identify the top 3 most relevant sections (e.g.,
Troubleshooting Playbook
andTools
) and inject only those into the system prompt. -
[Operational Layer]: Implement a cost/reliability heuristic: If the task is flagged as high-risk or requires multi-turn execution, the system will automatically increase prompt length (up to a predefined threshold) to leverage the observed operational advantage in reducing first-pass provider API failures on models like Gemini 3.8 Flash.
)What This Improved AI System Can Do:
-
[Optimized Resource Utilization]: It drastically reduces token use and inference costs by avoiding the overhead of loading massive, static context files for every simple query, achieving the same accuracy with a minimal baseline prompt.
-
[Enhanced Reliability in High-Stakes Scenarios]: For problems requiring complex reasoning (like multi-step BioProBench or detailed protocol analysis), it can strategically increase prompt length to maximize the chance of a successful first-pass API call, improving real-world execution success rates under budget constraints.
-
[Domain Expertise on Demand]: It enables agents to
reason like
an expert when needed, rather than constantly imposing a broad persona that carries no consistent accuracy gain. This allows for flexible deployment where the agent's expertise level scales with the complexity of the scientific question asked. -
[Improved Comparative Analysis]: By systematically testing these tiered prompts against baseline, generic, and mismatched profiles under controlled environments (as described in Section 3), researchers can reliably determine which specific components of a profile (e.g.,
Troubleshooting Playbook
vs.Mindset
) offer the highest return on investment for cost and accuracy.
Sources
- Prompting Science Report 4: Playing Pretend: Expert Personas Don't Improve Factual Accuracy
- Agent READMEs: An Empirical Study of Context Files for Agentic Coding
- SciKnowEval: Evaluating Multi-level Scientific Knowledge of Large Language Models
- Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?
- Do Context Files Help Coding Agents? A Two-Agent Ablation Study on Real Repositories
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
- BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science
- On the Impact of AGENTS.md Files on the Efficiency of AI Coding Agents
- SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines
- ClimaQA: An Automated Evaluation Framework for Climate Question Answering Models
- Are large language models superhuman chemists?
- Humanity's Last Exam
- In-Context Impersonation Reveals Large Language Models' Strengths and Biases
- Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
- When Does Persona Prompting Actually Help? A Retrieval and Metric Analysis of Expert Role Injection in LLMs
- ExpertPrompting: Instructing Large Language Models to be Distinguished Experts
- When "A Helpful Assistant" Is Not Really Helpful: Personas in System Prompts Do Not Improve Performances of Large Language Models
- MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection