Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing

arXiv:2608.07437 · cs.AI · Submitted 2026-08-07 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Paper Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing".

Jane: The paper was written by Jiacheng Miao, Jin Mu, Guanhua Chen and James Zou from Stanford University and University of Wisconsin–Madison.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Paper summary: Tom: Welcome back, everyone. Today we're looking at a paper that tackles something that sounds dry but is actually urgent — whether eye agents can be trusted to run a statistical hypothesis test and get the conclusion right. The authors built a new benchmark of real scientific tasks, then trained their own open-weight model to pass it.

Jane: And the reason it's urgent is that these coding agents are already being used to inspect datasets, write analysis code, and produce reports end to end. If they make subtle inferential mistakes, the fluently written conclusion can still be completely wrong. The paper shows exactly that happening with a frontier model.

Lu: What struck me is the specific failure mode. The model doesn't write broken code. The code runs, produces a precise p-value, and the conclusion still ends up wrong — because the chosen test doesn't fit the data. And existing benchmarks don't catch it, because they rarely check whether the analysis is statistically valid, only whether it executes.

Meng: And the example they open with is perfect. A cancer genomics dataset where a few outliers make linear regression look highly significant, but a rank-based test correctly fails to reject the null. GPT-5 point 4 literally noticed the outliers and the warning signs, then ran the linear regression anyway.

Lalam: So the paper does two main things. It builds P-Bench, 425 real hypothesis-testing tasks with expert-verified answer keys spanning economics, biology, and medicine. And it trains Fisher-R1, an open-weight agent that beats GPT-5 point 4 and DeepSeek-V4-Pro on the strictest scoring. A 14-billion-parameter model doing that tells you something about where the capability gap actually is.

Tom: And that gap lives in statistical judgment — knowing which test is valid given the assumptions of the data. The paper says the method must match the question and the data, otherwise a precise p-value can support the wrong scientific conclusion. That's the thread we'll pull through the whole episode.

Lu: Another thing about the framing — they're not claiming agents can replace statisticians. They're showing the current generation isn't reliable enough even for well-defined single tests. That's a lower bar than people assume.

Jane: It's a good thread. The opening pages lay out the failure mode in detail. And the numbers they cite frame everything else.

Page 1 of the paper: Jane: So the opening section carries one core claim — LLM agents frequently make subtle inferential errors that lead to incorrect conclusions, even when the executed analysis is technically correct.

Tom: Even when the code runs?

Jane: Exactly. The code runs, the p-value is precise, and the conclusion is still wrong, because the test itself wasn't valid for the data. And existing benchmarks fail to capture this, because they rarely assess whether a reported p-value is statistically valid given the assumptions underlying the data.

Lu: And they quantify it with the headline numbers. On P-Bench, Fisher-R1-14B shows a 21 percent average relative improvement in single-trial success over DeepSeek-V4-Pro. And on the most challenging tasks, that gain goes up to 26 percent.

Meng: But the more telling number is the baseline failure rate. Frontier models need that much improvement because they start from a low bar on strict scoring. And the hypothesis-testing workflow they describe is familiar to any scientist — take a question and a dataset, translate it into a testable hypothesis, select a test, compute a p-value, draw a conclusion.

Jane: I love that they show the actual trace in Figure 1. GPT-5 point 4 sees that tumor purity reaches a value of 2 point 47, which should be impossible, sees the linear model looking strong while the rank-based one looks weak, and still reports the linear association as significant.

Tom: Wait — the dataset had a tumor purity above 2 point 5?

Jane: Yes, an outlier that extreme should have been a red flag. Fisher-R1 instead runs the Spearman test, gets p = 0 point 085, and correctly fails to reject the null. The frontier model produces a false discovery; the trained 7-billion-parameter model gets it right.

Lalam: That example carries the whole motivation. It shows why the benchmark and the training pipeline exist, and why conclusion-only evaluation was overstating how reliable these agents are. The p-value has to be statistically valid, not just computable.

Tom: So the next thing they define is exactly that contract — what the agent must deliver and what the hidden answer key checks.

Page 2 of the paper: Jane: So we've seen the failure mode and the scale of it. Now the paper defines the problem formally. The agent receives a scientific question, a dataset, and a data description — but no prescribed method. It has to choose the statistical test, execute the analysis in an R environment, report a p-value, and return a reject or fail-to-reject decision at a pre-specified significance level.

Lu: The multi-turn loop matters too. The agent writes code, sees the output, gets warnings and diagnostics, and can revise. It's not a one-shot answer; it's an iterative process with real feedback, which mirrors how a human analyst works.

Meng: One detail worth keeping — the setting is open-ended by design. The task doesn't tell the agent which test to run. The agent must autonomously select an analysis strategy, and then the output is judged against a hidden answer key that the agent never sees.

Jane: And the related work section makes the landscape clear. Factoid benchmarks ask for a number. Workflow benchmarks check whether code executes and the final answer matches. StatQA handles method selection, but in multiple-choice format without running the analysis. Scientific claim verification checks claims against abstracts, but never reconstructs the underlying statistics.

Lalam: P-Bench sits in the gap, because the answer key isn't transcribed from what a paper claims. It's computed from a logged execution of the canonical reference analysis on the real dataset, then audited by domain experts. That grounding is what makes these tasks usable for both evaluation and training.

Tom: And you need that trust, because the benchmark is the measuring stick for everything else. Which brings up the pipeline question — how do you build 425 verified tasks without an army of human annotators?

Lu: Yeah, that's the construction pipeline. And that's exactly what comes next.

Page 3 of the paper: Tom: So now the construction pipeline. P-Bench draws from three source families — economics papers with datasets on Harvard Dataverse, biology papers with data on cBioPortal, and authoritative biostatistics teaching materials from Vanderbilt. Every task is anchored to an analysis that domain experts already computed and acted on.

Jane: Then the three-stage pipeline. Reproduce the reference analysis on a clean machine and log the execution. Filter out anything that can't be reproduced. Package a self-contained task with a structured answer key in the P-Bench format. So the p-value in the key comes from a logged run of canonical code, not from reading the paper's prose.

Lu: The expert audit closes the loop. Statisticians independently verify that the analysis request, the released data subset, and the answer key all align with the original source. Tasks that fail the review get repaired or removed.

Lalam: And one detail I appreciate — the benchmark deliberately includes realistic statistical traps. Outliers, heteroskedasticity, clustered observations are all in there. So the benchmark checks whether the model computes a number correctly, and it also checks whether the model notices when the data are trying to mislead it.

Meng: The composition numbers tell the same story. 425 tasks total, 203 easy and 222 hard. Hard tasks either belong to tricky method families — Cox regression, instrumental variables, Tobit — or they carry adversarial data-quality perturbations. And the method coverage spans 17 categories, with no single category exceeding 19 percent.

Jane: Those perturbations really matter. Textbook assumption violations that any trained statistician would check for — and the benchmark shows agents walking straight into them. On P-Hard, GPT-5 point 4's strict accuracy drops to 30 point 5 from 64 point 7 on easy.

Tom: That easy-to-hard drop is the measurement gap the field needed. So the benchmark exists and the failure is quantified. But 425 tasks are nowhere near enough to train an agent from scratch — that's where the synthetic data generator comes in.

Page 4 of the paper: Tom: So the synthetic task generator is the key to scaling. It's a Cartesian grid over six factors — statistical method, domain scenario, sample size, effect size, prompt style, and seed. The full corpus comes to 8,642 tasks with balanced coverage across all of them.

Jane: The clever part is the answer key. An LLM writes simulation code that generates a dataset, and then the canonical statistical method runs on its own simulated data. The p-value and the reject or fail-to-reject decision that come out become the ground truth. So the reward signal is verified by construction, not by human labeling.

Lu: And the effect-size axis has three regimes — null, borderline, medium. The borderline regime is auto-calibrated by simulation, so tasks land in a genuinely ambiguous significance range. That's where statistical judgment actually gets tested.

Meng: The data-quality perturbations carry over into training too. Missing values, extreme observations, invalid entries — the model has to notice and handle them rather than mechanically applying a test. That mirrors the traps in P-Bench, so both sets stress the same skills.

Jane: Then the SFT stage. They collect expert trajectories from Claude Sonnet 4 point 6 following a fixed five-step workflow — basic exploration, detailed exploration, assumption checking, method selection and analysis, and the conclusion. Automatic quality control keeps only trajectories with a valid multi-turn trace, a parseable conclusion, and agreement with the ground-truth decision.

Tom: So the quality bar is concrete. The trajectory has to match the ground truth on significance and stay within one order of magnitude on the p-value itself.

Jane: Exactly. That filtering is what turns noisy teacher behavior into a clean warm start. They keep about 83 point 5 percent of the 4,611 teacher trajectories.

Lalam: And all of this happens without touching P-Bench itself. The evaluation set stays out of the training corpus by construction, which is what makes the final generalization claims meaningful.

Tom: So the SFT gives the model discipline — the shape of good statistical behavior, checking assumptions before committing to a method. But SFT alone doesn't push p-value accuracy far enough. The real gains come from the reinforcement learning stage, where the reward is tied to the actual statistical outcome.

Page 5 of the paper: Jane: So the reinforcement learning stage is where the accuracy gains come from, and the reward function is the star. Two components — a p-value closeness score and a conclusion correctness score — gated by a hard format constraint. If the trajectory doesn't contain reasoning, executable code, and a parseable final answer, the reward is zero.

Lu: The p-value component uses a z-score transformation, and that's a thoughtful design choice. Raw p-values are compressed near zero, so the difference between 0 point 5 and 0 point 6 looks identical to the difference between 0 point 1 and ten to the minus ten. But the first pair is barely a change in evidence, while the second spans orders of magnitude. The z-scale spreads out the region where differences actually matter.

Meng: They also avoid rewarding method choice directly, because multiple procedures can be defensible for the same question. Instead, the outcome-grounded reward carries that signal indirectly — pick the wrong test, get a p-value that deviates from the reference, and the score drops. The weights, 0 point 9 on the p-value and 0 point 1 on the conclusion, reflect that the conclusion check mainly verifies consistency with the reported p-value.

Lalam: That's the right call for scientific practice — you don't want the model penalized for a defensible alternative test. But it means the p-value comparison has to carry the whole inferential load, and the z-space scoring is what makes that work.

Jane: The algorithm is DAPO, with decoupled clipping and dynamic sampling. Groups where every rollout gets the same reward are discarded and re-sampled, so updates only happen where there's real signal. And the results table shows it working — Fisher-R1-14B beats GPT-5 point 4 on three of the four strict metrics, including 33 point 0 versus 30 point 5 on P-Hard pass@1.

Tom: The stability gain is striking too. The 7B backbone's standard deviation on P-Easy raw accuracy collapses from plus or minus 8 point 5 to plus or minus 1 point 6. So the model is both more accurate and more consistent across rollouts.

Lu: And the most important number is the gap between raw and strict for the frontier baselines. GPT-5 point 4 scores 58 point 3 raw on P-Hard pass@1, but only 30 point 5 strict. It gets the reject or fail-to-reject direction right almost twice as often as it produces a p-value close to the canonical analysis. Conclusion-only evaluation was overstating reliability.

Tom: So the reward design targets exactly that gap. But anytime a model improves this much, the obvious worry is memorization — did it just learn synthetic prompt templates? The paper addresses that head-on.

Page 6 of the paper: Jane: The answer to the memorization worry comes in two parts. First, the ablation. SFT plus DAPO is the best configuration on every metric. DAPO on the raw backbone does lift P-Hard strict pass@1 from 13 point 2 to 25 point 2, but it plateaus well below the full system's 30 point 6. And SFT alone improves pass@3 without reliably improving pass@1.

Lu: Right, so the SFT warm start gives the policy a broad, plausible distribution of solutions, and the RL sharpens it toward accurate p-values. The combination is what gets both. That's a nice illustration of why these two stages complement each other.

Meng: Then the generalization check. They embed prompts in semantic space and compare similarity between P-Bench prompts and the training pool. Train-to-train similarities form a tight high-similarity band, while eval-to-train similarities sit clearly below it. So the 425 benchmark tasks have no near-duplicates in the training corpus.

Jane: And that fits the construction of the two sets. Training is synthetic simulations; evaluation is real-world data from published analyses. The performance gain reflects transfer, not retrieval of memorized prompts.

Tom: The discussion is honest about limits too. P-Bench evaluates a single hypothesis test per task. Extending it to multi-test pipelines with multiple-comparison correction is the natural next step. And they flag the harder question of making agents justify assumptions explicitly and know when no single test is adequate.

Lalam: I want to highlight their framing around misuse. A more reliable statistical agent can catch method-driven errors before they reach the literature — that's the reproducibility win. But it can also lend false legitimacy to weak claims. So they release the benchmark and the model as evaluation and oversight tools, not as substitutes for human statistical review.

Meng: Which is reassuring given how much training data was synthetic. The whole claim is that synthetic verified tasks teach real-world judgment, and the similarity analysis is exactly the check you'd want to see.

Tom: All the pieces line up — the benchmark, the training signal, the evidence against memorization, and a clear sense of what's still missing. That's the picture worth walking away with.

Conclusion: Tom: So let's close the loop. The paper's central point is that current LLM agents look fluent at hypothesis testing but are statistically unreliable, and the field didn't notice because the benchmarks weren't measuring inferential validity. P-Bench fixes the measurement, and Fisher-R1 shows that training with verified rewards closes a large part of the capability gap.

Jane: The results are hard to argue with. A 14-billion-parameter open-weight model beating GPT-5 point 4 on the strict hard metrics — that's about the training signal, not the size of the model. The z-space reward and the SFT warm start are the ingredients that make the difference. And the stability gain, the standard deviation dropping from 8 point 5 to 1 point 6, is what makes it usable in practice.

Lu: For me, the lasting contribution is the verifiability pipeline. Every answer key is grounded in a logged execution of canonical code and audited by domain experts. That's what makes it possible to train on synthetic tasks and still trust the real-world evaluation — and it's a template other benchmark builders should follow.

Meng: And the benchmark design leaves an agent nowhere to hide. Seventeen method categories, realistic traps, task families requiring Cox regression, instrumental variables, and mixed-effects models — you can't default to a single recipe and score well.

Lalam: Bigger picture: as autonomous research agents get more ambitious, statistical reasoning becomes the constraint on whether we can trust them in high-stakes settings like clinical trials and policy evaluation. This paper points at that bottleneck and shows a path through it, while insisting that human statistical review stays in the loop.

Jane: And they release P-Bench and Fisher-R1 as open tools, so the next round can build on them — multi-test pipelines, uncertainty over method choice, knowing when no single test is adequate.

Lu: The opening example stays with me, though. A frontier model seeing outliers, noting the warning signs, and still reporting the linear result as significant. That's the failure mode we should all be watching for in the next wave of tools.

Tom: That's a good place to leave it. Thanks for listening, everyone — we'll see you for the next paper.

Jiacheng Miao, Jin Mu, Guanhua Chen, James Zou

Stanford University · University of Wisconsin–Madison

cs.AI

Submitted: 2026-08-07

Updated: 2026-08-10

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 84/100

The gist: The paper addresses a critical gap in the use of large language model (LLM) agents for scientific hypothesis testing.

Key concepts

P-Bench
A benchmark of 425 real hypothesis-testing tasks from economics, biology, and medicine, each with an expert-verified answer key. It tests whether an agent can select a valid statistical test, run it, and report a p-value that matches the canonical analysis, not just execute code.
Strict scoring
An evaluation metric that requires the agent's p-value to be close to the reference p-value (within one order of magnitude) and the reject/fail-to-reject decision to match. This contrasts with raw scoring, which only checks the conclusion direction, revealing that frontier models often get the conclusion right but the p-value wrong.
Outcome-grounded reward
In reinforcement learning, the reward is based on the statistical outcome—the p-value closeness and conclusion correctness—rather than directly rewarding the choice of method. This allows multiple defensible tests but penalizes deviations from the reference p-value, encouraging statistically valid analysis.
Synthetic task generator
A pipeline that creates 8,642 training tasks by simulating datasets from a grid of factors (method, domain, sample size, effect size, etc.). The answer key is generated by running the canonical statistical method on its own simulated data, ensuring ground truth is verified by construction without human labeling.

Terminology

Summary

The paper addresses a critical gap in the use of large language model (LLM) agents for scientific hypothesis testing. While LLM agents can now automate the full empirical workflow—inspecting datasets, generating code, running analyses, and drafting reports—the authors demonstrate that these agents frequently make subtle inferential errors that lead to incorrect conclusions despite correctly executed analyses. The paper makes two main contributions: (1) P-Bench, a benchmark of 425 open-ended, realistic hypothesis-testing tasks spanning economics, biology, and medicine, and (2) Fisher-R1, an open-weight LLM agent trained via supervised fine-tuning (SFT) and reinforcement learning (RL) on synthetic tasks with verified statistical rewards. The authors report that "Fisher-R1-14B substantially improves over its backbone and outperforms strong proprietary and open-source baselines, including GPT-5.4 and DeepSeek-V4-Pro, achieving a 21% average relative improvement in single-trial success over DeepSeek-V4-Pro, with gains up to 26% on the most challenging tasks."

The motivating example concerns a cancer genomics task testing whether mutation burden is associated with tumor purity. The dataset contains high-leverage outliers: a linear analysis reports a highly significant association driven by these points, whereas a rank-based test does not. The authors show that a frontier LLM agent recognizes these warning signs yet proceeds with linear regression, producing runnable code and a fluent but incorrect conclusion. In contrast, a Spearman test correctly fails to reject the null. In the accompanying figure, GPT-5.4 observes that tumor purity reaches 2.47 and that linear regression has warning signs but says keeps linear regression anyway, concluding 'Despite this, the linear association test supports a positive association,' whereas Fisher-R1 notes Given non-normality and outliers, Spearman rank correlation is appropriate, correctly failing to reject the null.

The core claim is that "Existing evaluations do not directly measure this ability. General data-analysis benchmarks often reward a plausible answer or an executable workflow, but they rarely isolate the inferential method itself: which test was executed, whether the reported p-value is grounded in that execution, and whether the conclusion follows from the computed evidence."

Each task provides the agent with a scientific question q, dataset D, and data description s—but not the statistical method. The agent must choose a method, execute the analysis, return a p-value p̂, and a decision δ̂. Evaluation compares the output against a hidden answer key k* = (p*, δ*). The agent interacts with an R environment through a multi-turn loop, where actions may contain R code, intermediate analysis decisions, or the final answer; the environment executes code actions and returns observations such as data summaries, warnings, model outputs, diagnostics, test statistics, and p-values. The authors justify R over Python because "the canonical implementations of the statistical methods P-Bench covers — Cox proportional-hazards regression, mixed-effects models, instrumental-variable estimators, robust and rank-based tests — are most mature and standardized in R packages such as survival, lme4, rms, and AER."

P-Bench tasks are anchored to a statistical claim from a high-quality source, such as peer-reviewed top scientific journals and canonical course notes. It is built from three source families: "economics papers in top journals with datasets hosted on Harvard Dataverse, peer-reviewed biology papers with datasets hosted on cBioPortal, and authoritative teaching materials from Vanderbilt Biostatistics with datasets distributed through R packages. The construction pipeline involves three stages: (1) reproduce the reference analysis on a clean machine and log the execution; (2) filter out analyses that could not be reproduced, and ground each remaining claim in the execution log; and (3) generate a self-contained executable task paired with a structured answer key. Answer keys are trustworthy because they are computed from canonical reference code (not transcribed from paper text), cross-validated against the published claim, and every task is audited by domain experts."

P-Bench contains "425 high-quality open-ended hypothesis-testing tasks, stratified into Easy (203) and Hard (222) splits using deterministic, metadata-only rules: a task is Hard if it belongs to a reference-method family that requires careful model specification and assumption checking (e.g., Cox regression, IV/2SLS, Tobit model) or it carries an adversarial data-quality perturbation. The benchmark covers 17 hypothesis-testing method categories spanning commonly used tests (t-test, χ2, OLS coefficient test, Fisher exact test) and specialized methods routinely used in real published analyses (e.g., Cox proportional-hazards regression, IV / 2SLS estimators, mixed-effects models, log-rank tests, Mann–Whitney, Tobit model, probit model). No single category exceeds 19% of tasks. P-Bench also contains realistic statistical traps, including outliers, heteroskedasticity, and clustered observations, where naive methods yield confident-but-wrong results."

Training uses synthetic tasks with executable data-generating processes, known target analyses, and programmatically checkable answer keys. Each task is generated by combining six independent factors: a statistical method (M), a domain scenario (S m), a sample size (N), an effect size (E), a prompt style (P), and a random seed (K), forming the Cartesian grid M × S m × N × E × P × K. Each method is paired with multiple domain scenarios; effect size takes three regimes (null, borderline, medium), where the borderline regime is auto-calibrated by simulation, so tasks fall in a genuinely ambiguous significance range. Data-quality perturbations such as missing values and outliers are added. The full corpus contains 8,642 tasks. Critically, the answer key is the output of the canonical method on the simulated data, not the true generative parameter. This aligns the supervision target with what a correct analysis would actually produce.

Teacher trajectories are generated by Claude-Sonnet-4.6 on a randomly sampled subset of synthetic tasks, yielding 4,611 trajectories in total, following a fixed five-step workflow: basic EDA, detailed EDA, assumption checking, method selection and analysis, and conclusion. Quality control retains a trajectory only if it produces a valid multi-turn trace and a parseable conclusion, matches the ground-truth decision, and reports a p-value agreeing with the answer key in significance at α = 0.05 and within one order of magnitude. After filtering, 3,851 of 4,611 trajectories (≈83.5%) are kept for SFT, trained with masked autoregressive negative log-likelihood.

Starting from the SFT-initialized policy, the authors use the Decoupled Clip and Dynamic Sampling Policy Optimization (DAPO) algorithm with asymmetric clipping range (ϵl, ϵh) = (0.20, 0.28), dynamic sampling that discards groups with zero variance, and a group-relative advantage estimator. The reward is two outcome-grounded components, gated by a hard format constraint: R(o) = I valid(o)(w p r p(o) + w c r c(o)) with w p = 0.9, w c = 0.1. The p-value component compares the reported p̂ with the key value p* on a two-sided normal test-statistic scale: r p(o) = exp(−min z(p̂), 5 − min z(p*), 5 / σ) with σ = 1 and z(p) = Φ−1(1 − p/2). The authors justify the z-scale because "raw p-values are heavily compressed near 0, where the most informative differences in evidence live... the pair p = 0.5, p = 0.6 and the pair p = 0.1, p = 10−10 both have ∆p ≈ 0.1, yet the first represents essentially no change in evidence while the second spans many orders of magnitude in statistical significance. The conclusion component rc ∈ 0,1 checks the reject/fail-to-reject decision at α = 0.05. There is no explicit method-correctness term because multiple statistical procedures may be defensible for the same hypothesis-testing task; instead, an inappropriate method will typically produce a p-value that deviates from the reference analysis, and therefore receive a lower z-space p-value closeness reward."

Fisher-R1 uses Qwen2.5-Coder-7B and 14B backbones. Baselines include GPT-5.4, DeepSeek-V4-Pro, GPT-OSS-120B, Qwen3-Coder-30B, Qwen3-32B, and DataMind (trained for data analysis). Two metrics are used: Raw measures whether the final reject/fail-to-reject decision matches ground truth; Strict additionally requires the reported p-value to be numerically close in two-sided z-space: "Strict = Raw ∧ z(p̂) − z(p*) < 0.5." Three independent rollouts per task are sampled; pass@1 averages per-trial success and pass@3 counts a task solved if any trial succeeds.

Key results from Table 1:

  • Fisher-R1-7B on P-Easy: Raw pass@1 rises from 61.4 (backbone) to 87.0; Strict pass@1 rises from 36.3 to 65.7. On P-Hard: Raw pass@1 from 37.5 to 63.4; Strict pass@1 from 13.2 to 30.6.

  • Fisher-R1-14B achieves the best P-Hard Strict scores in the table (33.0 pass@1, 45.5 pass@3) and P-Easy Strict pass@1 of 64.2 (pass@3 75.9).

  • GPT-5.4 scores P-Hard Raw pass@1 58.3 but only 30.5 Strict; Fisher-R1-14B exceeds GPT-5.4 on three of four Strict metrics, trailing only narrowly on P-Easy Strict pass@1 (64.2 vs 64.7).

  • Fisher-R1 "beats DeepSeek V4 Pro, GPT-OSS-120B, Qwen-3-Coder-30B, and Qwen-3-32B on every Strict metric, indicating that for hypothesis testing, targeted training with a verified statistical reward is more effective than scaling the underlying coding model."

  • Run-to-run standard deviations drop sharply (e.g., ±8.5 → ±1.6 on P-Easy Raw pass@1 for the 7B), so Fisher-R1 is not only more accurate but also more stable across rollouts.

The authors emphasize the Raw–Strict gap: "Across all baselines, Strict scores are far below Raw scores, especially on P-Hard. GPT-5.4, for example, scores 58.3 Raw but only 30.5 Strict on P-Hard pass@1: it reports the correct reject/fail-to-reject direction nearly twice as often as it produces a p-value close to the canonical analysis." This holds under a relaxed threshold of ∆z < 1 (Table 5), leading to the conclusion that conclusion-only evaluation overstates agent reliability in open-ended hypothesis testing, and that explicitly rewarding p-value accuracy is what gets the model to ground its decision in the executed analysis.

Table 2 shows combining the SFT warm-start with DAPO is essential: SFT+DAPO achieves the best score on every metric, substantially above either stage alone. DAPO applied directly to the backbone lifts P-Hard Strict pass@1 from 13.2 to 25.2 but plateaus well below SFT+DAPO (30.6). SFT alone improves pass@3 by broadening the solution distribution, but does not reliably improve pass@1.

To verify the gains reflect generalization, the authors compute for each sample, the mean cosine similarity to its top-5 nearest neighbors in the training pool using text-embedding-3-small. The results show that "train-to-train similarities form a tight high-similarity 'memorization band' that reflects the templated structure of generated RL prompts, while eval-to-train similarities sit clearly below this band. The 425 P-Bench tasks, therefore, have no near-duplicate counterparts in the RL corpus, indicating that gains on P-Bench reflect generalization rather than retrieval of training prompts."

SFT uses LlamaFactory: 3 epochs, learning rate 1e−5, cosine scheduler, warmup ratio 0.1, batch size 16, cutoff length 8192 tokens. RL uses verl: 1 epoch, learning rate 1e−6, 20 warmup steps, batch size 16, max prompt length 2048 tokens, max response length 4096 tokens, asymmetric clipping (0.2, 0.28), rollout temperature 0.7, top-p 1.0, group size G = 8; inference uses temperature 0.3, top-p 0.9. All training runs on a single node with NVIDIA H200 GPUs; the full RL training run completes in about 36 hours for the 14B model.

The authors conclude that even current frontier models still lack reliable statistical reasoning for open-ended hypothesis testing. They note P-Bench currently evaluates a single hypothesis test per task and propose extending it to multi-test pipelines with multiple-comparison correction as natural next steps. They also caution: "Reliable statistical agents can improve scientific reproducibility by catching method-driven errors before they enter the literature. They could also lend false legitimacy to weak claims. We therefore release P-Bench and Fisher-R1 as evaluation and oversight tools, not as substitutes for human statistical review. Future work should study how to make agents explicitly justify assumptions, quantify uncertainty over method choice, and recognize when no single hypothesis test is adequate for a scientific question."

Improvements for AI systems

Improvements to AI systems:

  1. Reward statistical evidence, not just conclusions. Train agents with a two-component reward: (a) numerical closeness of the reported p-value to the canonical p-value, computed in two-sided normal z-space (z(p)=Φ−1(1-p/2), clipped at ±5, with exponential decay), and (b) binary correctness of reject/fail-to-reject at α=0.05. This forces the model to ground its decision in the actual executed analysis rather than producing a plausible-sounding conclusion.

  2. Gate rewards on format validity. Multiply the outcome reward by I valid(o) so that only parseable, structurally valid responses receive credit. This reduces hallucinated conclusions and enforces a disciplined final-answer format (method, p-value, decision).

  3. Train with a two-stage SFT + RL pipeline on synthetic tasks with executable ground truth. Generate a large grid of tasks by crossing statistical method, domain scenario, sample size, effect size, prompt style, and random seed. Use canonical R implementations to compute the answer key from the simulated data, not from generative parameters, so the supervision target matches what a correct analysis would produce.

  4. Use the borderline effect-size regime. Auto-calibrate effect sizes by simulation so that a substantial fraction of training tasks fall in a genuinely ambiguous significance range. This teaches the agent to perform careful assumption checking and method selection under uncertainty, instead of defaulting to a familiar test.

  5. Encourage assumption checking explicitly in the multi-turn loop. Structure agent trajectories as: basic EDA → detailed EDA → assumption checking → method selection and analysis → conclusion. Filter retained teacher trajectories by requiring the final decision to match ground truth and the reported p-value to be within one order of magnitude and significance-consistent with the key.

  6. Select methods that are robust to data-quality perturbations. Inject outliers, missing values, heteroskedasticity, and clustered observations into synthetic training tasks. Reward p-value agreement in z-space, which implicitly penalizes inappropriate methods because an ill-chosen test will produce a p-value that diverges from the canonical reference analysis.

  7. Use the P-Bench benchmark for evaluation. Measure both Raw accuracy (decision match) and Strict accuracy (decision match plus numerical p-value closeness). This exposes the common failure mode where models report the correct direction but a p-value far from the canonical one—showing that conclusion-only evaluation overstates reliability.

  8. Apply DAPO with asymmetric clipping and dynamic sampling. Use clipping bounds (εl=0.20, εh=0.28), group size 8, discard zero-variance groups, and employ a group-relative advantage estimator. This stabilizes RL and yields large gains over either SFT alone or RL without the warm start.

What the improved AI system can do:

  • Perform open-ended hypothesis testing reliably. Given a scientific question, a dataset, and a data description—but no prescribed method—it chooses an appropriate statistical test (e.g., Spearman instead of linear regression when outliers and non-normality are present), executes the analysis in R, and returns a p-value and reject/fail-to-reject decision grounded in the computed evidence.

  • Avoid subtle inferential errors. Unlike frontier models that recognize warning signs yet proceed with linear regression, the improved system detects high-leverage outliers, non-normality, heteroskedasticity, and clustered observations, then selects robust methods such as rank-based tests, Cox regression, IV/2SLS, mixed-effects models, or Tobit/probit models as appropriate.

  • Produce numerically accurate p-values. The z-space reward makes reported p-values fall close to canonical reference values, not merely match significance direction. On P-Bench, the 14B model reaches 64.2 Strict pass@1 on easy tasks and 33.0 on hard tasks, surpassing GPT-5.4 on three of four Strict metrics and beating DeepSeek-V4-Pro by 21% average relative improvement in single-trial success, with gains up to 26% on the hardest tasks.

  • Generalize beyond training templates. The agent does not retrieve memorized prompts: P-Bench tasks have no near-duplicate counterparts in the RL corpus, and gains reflect learned statistical reasoning rather than pattern matching. Cosine-similarity analysis confirms eval-to-train similarity sits far below the memorization band.

  • Behave consistently across runs. Run-to-run standard deviation drops dramatically (e.g., from ±8.5 to ±1.6 on P-Easy Raw pass@1 for the 7B model), making the agent reliable for repeated analyses and reproducible scientific workflows.

  • Serve as an oversight tool. It can audit human or AI-generated analyses by independently re-running hypothesis tests, flagging method-driven errors before they enter the literature—while still being designed for evaluation and oversight, not as a replacement for human statistical review.

Abstract

Reliable hypothesis testing is the foundation of many empirical scientific claims. Large language model (LLM) agents are increasingly used to automate this process, as they can inspect datasets, generate code, and produce analyses end-to-end. However, we show that they frequently make subtle inferential errors that lead to incorrect conclusions despite correctly executed analyses. Existing benchmarks fail to capture this failure mode, as they rarely assess whether a reported p-value is statistically valid given the assumptions underlying the data. We address this gap by building P-Bench, a benchmark comprising 425 open-ended, realistic hypothesis-testing tasks spanning economics, biology, and medicine. Each task requires an agent to select a statistical method, compute a p-value, and draw a conclusion given only a scientific hypothesis and a dataset. We further introduce Fisher-R1, an open-weight LLM agent trained for rigorous hypothesis testing using synthetic tasks and reinforcement learning. On P-Bench, Fisher-R1-14B substantially improves over its backbone and outperforms strong proprietary and open-source baselines, including GPT-5.4 and DeepSeekV4-Pro, achieving a 21% average relative improvement in single-trial success over DeepSeek-V4-Pro, with gains up to 26% on the most challenging tasks. Our results demonstrate that current LLM agents lack reliable statistical reasoning for hypothesis testing and that reinforcement learning on tasks with verified statistical reward substantially improves reliability.

Sources

Related papers