Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration".
Jane: The paper was written by Patrik P. Süli, György Eigner and Roland Hollós from Óbuda University and HUN-REN Centre for Agricultural Research and Czech Academy of Sciences.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Welcome back to the show, everyone. Today we're looking at a paper that's been making the rounds on arXiv, and it's called "Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration." Jane, I have to say, that title is a mouthful, but the problem it's tackling is something I think a lot of scientists will instantly recognize.
Jane: Oh, absolutely, Tom. And I love the name "Distribird" — it's like a little bird that goes out and gathers seeds of knowledge from the scientific literature and brings them back to build a nest. That's basically what this tool does, but instead of seeds, it's collecting numbers from research papers.
Tom: So let's break down what's actually going on here. The paper is from researchers at Obuda University and the HUN-REN Centre for Agricultural Research in Hungary. They're tackling this thing called Bayesian model calibration, which is a fancy way of saying: when you build a computer model of something real — like crop growth or water flow — you need to figure out the right values for the parameters in that model.
Jane: And the "Bayesian" part means you start with what you already believe about those parameters before you look at the data. That's called a prior. The problem is, most scientists just pick a flat, uniform prior because it's easy, even though they know it's not great. It's like saying every possible value is equally likely, which is almost never true.
Tom: Right, and the paper makes a really strong point about this. They say that building informative priors — priors that actually reflect what we know from decades of published research — is so time-consuming that nobody does it. You'd have to read dozens of papers, pull out the numbers, and figure out how to combine them. That could take days for a model with twenty parameters.
Jane: So Distribird automates that whole process. You give it a parameter name, like "maximum temperature for photosynthesis," and it goes out, searches scientific databases, reads the papers, extracts the reported values, and fits a probability distribution to them. All automatically.
Tom: And here's the kicker — it does this using AI models that run entirely on local hardware. No data leaves the researcher's machine. That's a big deal for scientists working with unpublished or sensitive modeling details.
Jane: I think that's what excites me most, Tom. This isn't just about saving time. It's about making Bayesian calibration actually practical for people who aren't statistics experts. The paper is essentially saying: the knowledge is out there, let's make it accessible.
Tom: And we're going to dig into exactly how it works in a bit, but first — Lu, you've been quiet. What's your take on the title and the core idea here?
Lu: I think the ambition is right. The prior is where Bayesian methods live or die, and the field has been stuck on uniform priors for decades because the alternative is too expensive. If this tool works as advertised, it could change how a whole generation of modelers approach calibration.
Jane: And that's a great segue into what the paper actually found when they tested it. Stay with us.
Summary: Tom: So we're back, and we're still talking about "Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration." Jane, you had a chance to look at the summary — what did the authors actually build?
Jane: Okay, so imagine you're a scientist and you have a parameter you need a prior for. You type in its name, a plain-language description, the unit, and maybe some physical bounds. Distribird then runs a whole pipeline. It's got multiple AI agents working together — one searches the literature, another reads the papers, another extracts the numbers, and another fits the distribution.
Tom: And it's not just a simple search-and-extract. The paper describes something called a multi-agent LangGraph pipeline with feedback loops. If the first search doesn't find enough evidence, the system goes back and tries again with broader terms. If it finds papers but can't extract values, it refines its search strategy.
Jane: Exactly. And here's a detail I really appreciated — it doesn't treat all papers equally. It judges how relevant each paper's study context is to your specific domain. So if you're calibrating a model for maize in Central Europe, a study done on maize in Kansas is more relevant than one done on rice in Thailand. The system weights the values accordingly.
Tom: That's smart. And then it fits the best probability distribution using something called AIC model selection — basically it tries several different distribution shapes and picks the one that fits the extracted data best.
Lu: What I find impressive is the honesty built into the system. It doesn't pretend to know things it doesn't. If it only finds one value, it gives you a low-confidence prior. If it finds nothing at all, it falls back to a wide, uninformative prior and clearly labels it as such.
Jane: And that's the part that makes it trustworthy for scientific use. Every prior comes with a complete provenance chain — the search queries, the papers found, the values extracted, the relevance judgments. A reviewer can trace every number back to the exact sentence in the paper where it was reported.
Meng: But I have to ask — how well does it actually perform? Because a pipeline like this sounds expensive to run.
Tom: Great question, Meng. The paper evaluated it on twenty-four parameters across ten scientific domains. And here's the honest finding: the full pipeline doesn't produce more accurate priors than a single, well-prompted AI call. They're basically tied on prior placement.
Meng: So why bother with the whole pipeline then?
Jane: Because of what the pipeline adds beyond accuracy. Every prior is traceable to cited sources. The system refuses to produce priors for fabricated or non-empirical parameters — we'll get into that in a minute. And it runs entirely on local open-weight models. Those are properties that matter for serious scientific work, even if the point estimate isn't better.
Lu: And I'd add that the cost is real but manageable. The paper reports about a million tokens and twenty to forty minutes per parameter on local hardware. That's not nothing, but it's a lot cheaper than a human spending days reading papers.
Tom: So the summary is: it's not trying to be smarter than a single AI call — it's trying to be more accountable. And that's a trade-off worth making in science. Up next, we're going to look at the specific improvements the paper suggests and how the system handles garbage requests.
Improvements: Tom: Welcome back. We're still on "Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration," and now we're getting to the part I find genuinely exciting — the improvements the paper brings to the table. Jane, walk us through the validity layer.
Jane: Okay, so this is the safety net. The paper identifies a real problem: if you ask an AI directly for a prior on a made-up parameter, it will often confidently invent one. The authors tested this with nonsense names like "mumblesnort factor" and "fake quantum correction xyz." A single-prompt AI model returned confident, informative-looking priors for those in eleven out of thirty cases.
Tom: That's terrifying. A scientist could take that prior, plug it into their calibration, and get results that look legitimate but are built on nothing.
Jane: Exactly. So Distribird has a validity classification system. It checks whether the parameter is actually recognized and whether it's empirically measurable. If the name is unrecognized, the pipeline skips the expensive search and extraction steps entirely and flags the request as likely invalid. If the evidence is thin, it marks the result as suspicious and asks for a second opinion from another AI call.
Lu: The key improvement here is that the system is designed to say "I don't know" — or even "this doesn't exist" — rather than hallucinate. That's a fundamental shift from how most AI tools behave. Most models are trained to always give an answer. This one is trained to know when not to.
Meng: And the early-skip saves a ton of compute. The paper says fabricated names are caught in seconds using just a few thousand tokens, instead of burning a million tokens on a search that was doomed from the start.
Tom: Right, that's the efficiency angle. But there's another improvement I want to highlight — the relevance-aware synthesis. Jane, you mentioned this earlier, but can you go deeper?
Jane: Sure. When the system extracts values from papers, it doesn't just throw them all into a pot. Each paper gets a relevance label — high, medium, or low — based on how well its study context matches your domain. High-relevance values get full weight, medium values get sixty percent, low values get twenty-five percent. And the confidence level of the final prior is capped by the relevance of the evidence behind it.
Lu: That's a really thoughtful design. It means a single off-topic paper can't dominate the prior, but it also doesn't get completely discarded. It's a gentle down-weighting rather than a hard cutoff.
Meng: And what about the fitting itself? The paper mentions AIC selection. Can you explain that in plain terms?
Jane: So once you have a set of values, the system tries fitting five different distribution shapes — Normal, Truncated Normal, Gamma, Log-Normal, and Beta. AIC, the Akaike Information Criterion, is a way of scoring which shape fits the data best while penalizing complexity. Since all five have the same number of parameters, it comes down to which one has the highest likelihood given the data.
Tom: And the system picks the winner and marks it as high confidence. If there are fewer values, it falls back to simpler methods with lower confidence labels. It's a tiered approach that matches the confidence to the evidence.
Lu: I think the most important improvement, though, is the transparency. Every prior comes with a full audit trail. You can see exactly which papers contributed which values, and you can discard any source you don't trust. That's what makes this usable in real scientific workflows.
Jane: And that's the hook for our final segment — what does this mean for the future of scientific modeling?
Conclusion: Tom: And we're wrapping up our discussion of "Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration." Jane, give us the big picture — what did this paper really accomplish?
Jane: I think it reframed the problem. The authors didn't try to build a tool that's smarter than a single AI call. They built a tool that's more accountable. The priors it produces are about as accurate as what you'd get from a well-prompted model, but they come with evidence, citations, and confidence levels. And critically, it refuses to fabricate priors for parameters that don't exist.
Tom: And that's the part that really matters for science. When you publish a result, you need to be able to defend your choices. With Distribird, you can show exactly where your prior came from — which papers, which values, which reasoning. That's a level of transparency that single-prompt AI just can't offer.
Lu: I'd add that the local-only operation is a quiet revolution. No unpublished modeling details leave the researcher's machine. That's going to be increasingly important as AI tools become more integrated into scientific workflows and data privacy concerns grow.
Meng: And the cost, while not trivial, is manageable. Twenty to forty minutes per parameter on local hardware is a reasonable price for a defensible prior. And for out-of-scope requests, it's nearly free because it catches them early.
Jane: The paper is also honest about its limitations. It's designed for physically interpretable parameters in active research fields. It's not for neural network weights or purely empirical tuning constants. But within its scope, it fills a real gap.
Tom: So what's the takeaway for our listeners? If you're a modeler who's been defaulting to uniform priors because building informative ones is too expensive, this tool gives you a practical alternative. It's open source, pip-installable, and runs on your own hardware.
Lu: And I think the broader implication is that AI tools for science don't have to be black boxes. This paper shows a path where AI assists with the grunt work — searching, reading, extracting — while keeping the human in the loop with full visibility into every step.
Jane: That's exactly right. It's not about replacing the scientist's judgment. It's about giving them better raw material to exercise that judgment on.
Tom: Well said. That's our show on "Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration." Thanks to Lu, Meng, and Jane for a great discussion. Join us next time when we'll be looking at another paper from the arXiv. Until then, keep questioning, keep modeling, and keep your priors honest. Goodbye, everyone.
Patrik P. Süli, György Eigner, Roland Hollós
Óbuda University · HUN-REN Centre for Agricultural Research · Czech Academy of Sciences
cs.AI, cs.MA, cs.SE
Submitted: 2026-08-17
Updated: 2026-08-19
Code: https://github.com/HUN-REN-AI1Science/Distribird
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 63/100
The gist: The paper presents Distribird, an agentic web application that automates the construction of informative prior distributions for Bayesian model calibration from scientific literature.
Key concepts
- Bayesian Model Calibration
- This is the process of setting the correct parameters for a computer model, such as one predicting crop growth. The 'Bayesian' approach involves starting with what you already believe about those parameters (a prior) before looking at actual collected data.
- Prior Distribution
- This represents a pre-existing belief or range of values used in Bayesian methods. Instead of assuming all values are equally likely, Distribird builds an informed distribution based on decades of published research, making the model more accurate.
- Multi-agent LangGraph Pipeline
- This is the core automated system that executes the search and extraction process. It uses multiple specialized AI agents—one to search literature, one to read papers, one to extract values, and another to fit distributions—with feedback loops for refinement.
- AIC Model Selection
- This is how the system chooses the best distribution shape (e.g., Normal or Gamma) for a set of extracted data. It selects the shape that fits the values best while penalizing complexity, ensuring high confidence in the final generated prior.
Terminology
Summary
The paper presents Distribird, an agentic web application that automates the construction of informative prior distributions for Bayesian model calibration from scientific literature. Despite decades of methodological work, researchers almost always fall back on uniform priors
because building informative priors from scientific literature is slow and needs both domain and statistical expertise.
Distribird deploys "a multi-agent LangGraph pipeline that searches Semantic Scholar and OpenAlex in parallel, reads the retrieved papers in full, extracts reported numerical values, weights them by how well each source’s study context matches the target domain, and fits the best-matching probability distribution via AIC model selection from a predefined set of distributions. When no literature is available, the system
falls back to sensible uninformative alternatives (a uniform prior or a wide Normal), and clearly reports both the evidence behind and the confidence level of every prior it produces. The tool is designed for
physically interpretable parameters in process-based models, where domain knowledge exists in the published literature. Evaluated on
24 parameters across ten scientific domains using three open-weight models (Qwen3.6 27B, Gemma 4 31B, Mistral Small 4 119B) run entirely on local hardware, the full pipeline
matches this baseline rather than beating it on prior quality. Its contribution is different:
Every prior is traced to the specific papers and values from which it was constructed; a built-in validity layer declines to produce priors for out-of-scope requests, flagging all five out-of-scope test parameters on every model, whereas the single-prompt baseline returns confident but unfounded priors for them in 11 of 30 model–parameter cases; and every language-model call runs on local open weights, so no parameter description or unpublished modelling detail is transmitted to a third-party LLM provider (only generated search terms reach the public literature databases). The authors argue
these properties matter more than a marginal improvement in point-estimate accuracy" for scientific use.
Process-based models "describe complex natural systems (eg. crop growth, water movement through soil, biogeochemical cycles, disease spread, etc.) using mechanistic equations with parameters that represent physical, chemical, or biological quantities. Because many parameters
cannot be measured directly, they must be inferred from observations by adjusting parameter values until the model output matches available data, a process called
model calibration, or inverse modeling. Bayesian calibration is
arguably the soundest approach to this problem, on both statistical and philosophical grounds as it
combines prior knowledge about the parameters with the information in the observed data, and returns a posterior distribution that captures not only the most likely parameter values but also the uncertainty around them."
The Bayesian approach requires the researcher to specify a prior distribution for each parameter before seeing the data.
The mathematically convenient answer is to use a uniform prior,
which is almost universally adopted in practice
but almost always inadequate.
Pericchi and Walley [1991] showed no single uninformative prior works in every case.
Uniform priors are not invariant under reparameterisation,
and with a uniform prior the posterior mode coincides exactly with the maximum-likelihood estimate.
The information lost "has real consequences: the sampler converges more slowly, credible intervals are wider than they need to be, and when data are scarce, posteriors are poorly constrained where an informative prior would have stabilised them. Informative priors
improve all of these properties. The knowledge needed
exists. It is in the scientific literature. The problem is that
gathering and combining it is expensive, taking
days of reading for a model with twenty parameters.
Researchers default to uniform priors not because they believe them appropriate, but because the alternative is too costly."
Recent LLMs that can search databases on their own have created a new option.
Directly prompting an LLM for a prior has in fact been shown to yield well-placed, informative distributions quickly and cheaply.
However, "LLMs return priors with no traceable basis: as we demonstrate, a model will readily produce a confident, informative prior for a parameter that has no empirical referent, a typographic error, or a software-internal tuning constant. Also important is
the question of attribution: when a language model synthesises findings from the literature without citing its sources, the researchers who produced that empirical work receive no credit. The objective is
not to extract a more accurate point estimate from a language model, but to make literature-based prior construction trustworthy, meaning:
every value is traceable to a cited source; out-of-scope requests are declined rather than confabulated; and the entire process runs on open-weight models on the researcher’s own hardware. Distribird is
not a general-purpose prior elicitation tool but is
designed specifically for the class of problems where literature-based prior construction is most appropriate and most impactful: process-based models with physically interpretable parameters, active publishing communities, and genuine prior knowledge encoded in the scientific record. The explicit trade-off:
on our benchmark the full pipeline does not produce more accurate priors than a single well-prompted LLM call, and it is considerably more expensive to run. What it provides is
provenance, robustness against fabricated evidence, and fully local operation."
Distribird is implemented as a Python 3.10+ package distributed via PyPI and exposing two user-facing interfaces: a REST API built on FastAPI and an interactive Streamlit web application.
Both call the same core pipeline.
The user provides a ParameterInput object specifying the parameter name, a plain-language physical description, the measurement unit, a domain context string, and optional physical constraints.
The system returns a PipelineResult containing "the fitted prior distribution, its confidence level, the number of contributing sources, all search queries attempted, the full list of retrieved papers with extracted values, an optional enrichment context, and, when multi-agent deliberation is enabled, the moderator’s rationale and any excluded papers."
The core is a LangGraph StateGraph that orchestrates a multi-agent pipeline with three feedback loops and a conditional forward path.
The pipeline allows cycles: when the initial evidence is too thin, the graph routes back to earlier nodes for another attempt before moving on to synthesis.
All nodes read from and write to a shared PipelineState typed dictionary that follows the blackboard pattern.
The pipeline comprises "eleven primary processing nodes: Enrich (parameter semantics expansion), SearchStrategy (search-scope planning), QueryGen (search query generation), Search (parallel API queries), RelevanceJudge (paper scoring), CrossEnrich (citation snowballing), FetchFulltext (PDF retrieval and parsing), Extract (numerical value extraction), QualityGate (routing logic), Synthesize (distribution fitting), and ValidityCheck (out-of-scope classification). There are
three feedback loops and one conditional forward path":
-
Loop A, Search refinement:
When the quality gate finds zero extracted values but papers were retrieved,
the pipeline routes to a RefineSearch node, whichgenerates refined search queries targeting more specific experimental or calibration studies.
This loopmay execute up to two iterations.
-
Loop B, Domain broadening:
When the quality gate finds too few domain-relevant values (fewer than domain broadening min relevant, default 2)
and search breadth can be widened, the pipeline routes to a BroadenSearch node, whichadvances the current breadth one tier toward broad and returns control to QueryGen.
This loopmay execute up to domain broadening max times (default 2).
-
Conditional path, Cross-enrichment:
if the maximum iteration number permits and at least two high-relevance papers have been identified, the pipeline routes through CrossEnrich (citation snowballing) before fetching full texts.
This isexecuted at most once per pipeline invocation.
-
Loop C, Extraction refinement:
When the quality gate finds extracted values but none are high-confidence and the coefficient of variation exceeds a predefined threshold,
the pipeline routes to a RefineExtraction node thatuses web-assisted search to locate additional confirming or disconfirming evidence.
This loopmay execute once.
An IterationBudget object tracks the number of iterations consumed by each loop and enforces a global cap on total LLM calls (default: 30), guaranteeing termination.
A scope-planning node (SearchStrategy) classifies the request’s domain-specificity as high, medium, or low and maps it to a starting search breadth (strict, mixed, or broad, respectively).
The domain-broadening loop widens the breadth after the fact when the first pass returns too little on-target evidence.
Papers found at each tier accumulate rather than replacing one another, so broadening only ever adds evidence.
Broadening escalates at most domain broadening max times (default 2), so the pipeline starts as narrow as the request allows and widens only as far as it must.
Every request is classified as Valid, Suspicious, Likely Invalid, or Unknown.
The early-skip runs right after enrichment
: if the LLM does not recognise the parameter name, the pipeline goes straight to the validity node and skips search, extraction, and synthesis,
saving 80–95% in our benchmarks
of runtime. The terminal classifier combines the enrichment self-reports, the number of refined queries, papers, and extracted values, and the prior’s confidence into a verdict.
If the heuristics return Suspicious, a single second-opinion LLM call is made
that can sharpen the verdict to Likely Invalid or confirm it as Suspicious, but never overrides a Valid result.
The search subsystem queries two academic APIs in parallel
: Semantic Scholar and OpenAlex, both restricted to open-access papers.
Two more source agents are optional: a deep-research agent hands the task to an LLM with web-search capabilities
and a web search agent runs a similar checked LLM search.
Both agents' results are checked against Semantic Scholar before they are included, to guard against made-up references.
The FetchFulltext node retrieves and reads the complete text of each selected paper
using a tiered fetch fallbacks
cascade: URL-derived resolver, primary PDF, open-access mirror, last-resort resolvers (DOI→PMC, arXiv preprint), and an opt-in stealth browser (Camoufox) for JavaScript bot walls. PDFs are read as Markdown via PyMuPDF4LLM rather than flattened plain text,
preserving tables, headings, and page boundaries. Extraction reads the entire paper, not just its opening pages,
splitting long papers into overlapping pages
sized to the model's context window.
The fitting strategy has tiers set by how many values survive extraction and quality-gate filtering
:
-
≥5 values:
AIC selection across admissible families
(Normal, Truncated Normal, Gamma, Log-Normal, Beta), markedhigh confidence.
-
2–4 values:
Moment matching with σ widened ×1.5 and a min-width floor
against a Truncated Normal, markedmedium confidence.
-
1 value:
Centred on the reported value; σ from paper’s uncertainty, else half the reported value,
markedlow confidence.
-
0 values:
Centred at midpoint of bounds with σ = (u − l)/4
(or effectively unbounded with µ = 0, σ = 1000), markednone / uninformative.
Relevance-aware synthesis is enabled by default: each value-bearing paper receives a dedicated LLM judgment of how well its study context matches the searched domain,
yielding a label of high, medium, or low. This changes synthesis in three ways: selection (fit from the strongest subset), weighting (relevance factors 1.0, 0.6, 0.25), and confidence ceiling (capped by relevance). Values are weighted by the square root of its sample size
when reported.
Distribird exports priors in three formats: JSON (a structured record containing the parameter name, distribution family, fitted parameters, confidence level, informative/uninformative flag, fitting rationale, number of contributing sources, and a citation list
), Python (an executable scipy.stats script
), and R (an executable R script using base distribution functions
). All formats carry the full provenance chain.
An optional integer seed is forwarded to the LLM API’s seed field,
and sampling temperature is set per task class.
An opt-in debug-trace framework captures a complete structured record of a run, every LLM prompt and response, each search request, each PDF fetch outcome, the extracted values, and the AIC candidates considered during fitting.
All runs use open-weight models served locally through llama.cpp as Unsloth Dynamic (UD) GGUF quantisations
: Qwen3.6 27B and Gemma 4 31B at Q8 K XL, and Mistral Small 4 119B at Q4 K XL, all on a single NVIDIA H100 GPU (80 GB).
The baseline is a naive single-prompt elicitation
with no literature search at all,
run on the same three local models and for reference, on three cloud frontier models (Claude Opus 4.8, Gemini 3.1 Pro, and GPT-5.5).
The benchmark covers 24 parameters across 12 use cases spanning ten scientific domains
(climate, pharmacokinetics, hydrology, structural engineering, epidemiology, astrophysics, three ecology systems, geophysics, robotics, economics).
Running the full pipeline locally is computationally expensive
: a single parameter consumes on the order of a million LLM tokens over roughly 80 model calls and 20–40 minutes of time on local hardware.
Full-text reading and extraction dominate. Per-parameter figures span from 0.4 M tokens and 12 minutes
to 4.7 M tokens and 96 minutes
for the Hubble constant.
Prior accuracy is scored as accuracy = 1 − prior mode − data optimum / range.
The two approaches place priors about equally well
: Across all 24 parameters their mean accuracy is within a few points on every model.
Seven parameters are benchmark artifacts
where the synthetic data force the true value close to the top of the allowed range, higher than any published value.
On the remaining 17 fair
parameters, Distribird is ahead by about seven points on Mistral (83% vs. 76%) and behind by less than one point on Gemma 4 and Qwen.
A two-sample Kolmogorov–Smirnov test over all 72 parameter–model pairs finds D is small and the p-value is far above the usual 0.05 threshold,
so we cannot reject the hypothesis that the accuracy scores come from one common distribution.
The pipeline shows no measurable difference in prior placement from a single well-written prompt.
Against flagship cloud models, on the fair set all three flagship baselines score 84%,
which is essentially the same as Distribird on local open-weight models (83–84%).
The local, auditable pipeline places priors about as well as a direct prompt to the strongest cloud models.
Effective Sample Size (ESS) from four NUTS chains shows Distribird’s priors give a higher ESS on 46–54% of the parameters, and the naive priors on about 62%
compared to a flat prior.
Distribird attaches "a complete, machine-readable provenance record: the search queries it issued, the papers it found, the numerical values it extracted with their in-text context, the per-paper domain-relevance judgments, the AIC family and score, and a citation list with DOIs. For example, the COVID-19 incubation-period prior built by Qwen3.6 27B is
an 'AIC-selected lognormal from 85 values (85 high- and medium-relevance values kept, 12 low-relevance dropped)' drawn from 43 source papers. On saturated hydraulic conductivity,
the naive single-prompt model had no value it trusted and fell back to an uninformative prior (accuracy 0.50), whereas Distribird on Gemma 4 read 69 papers, kept the 9 values its relevance judge rated high or medium, dropped 6 low-relevance ones, and fit a high-confidence truncated Normal that placed the prior near the data optimum (accuracy 0.91)."
The test set is "six requests: two fabricated names with no scientific meaning (mumblesnort factor, fake quantum correction xyz), three real but non-empirical model-internal quantities, and one real empirical control (specific leaf area). Distribird's validity layer
flagged all 15 fabricated or non-empirical requests (five per model) as non-valid, likely invalid or suspicious, never valid and
recovered the real control on all three (3/3). On the 24 in-scope benchmark parameters,
it never once hard-rejected a real parameter and marked 71% (51/72) valid. The naive baseline
returned a confident informative prior for the fabricated parameters in 11 of 30 model–parameter cases and never once refused. Fabricated names are caught by the early-skip route
in 12 s to 3.6 min using only 2–6 k tokens, while requests reaching the full pipeline cost
40 k–1.2 M tokens and 7–90 minutes."
Prior elicitation has a large methodological literature
falling into expert elicitation protocols
and empirical Bayes methods.
Distribird fills a different niche: automated literature-based elicitation, where the prior knowledge comes from the published scientific record rather than from a human expert or a separate dataset.
The authors state: "We are not aware of another system that automates literature-based prior construction end to end, combining multi-agent literature search, per-source relevance weighting, AIC-based distribution fitting, and explicit confidence and out-of-scope reporting."
The prior problem in Bayesian calibration is not a problem of missing knowledge; it is a problem of the cost of getting to it.
Distribird makes that knowledge available automatically, turning a task that used to take days of expert effort into one that takes minutes.
The evaluation shows "the full Distribird pipeline does not produce more accurate priors than a single, well-prompted call to the same language model, and it is substantially more expensive to run. Its contribution is trustworthiness rather than accuracy. The tool is
openly available, pip-installable, and deployable via Docker. For informal use,
a single LLM call is a cheaper and comparably accurate choice. For scientific use,
where a prior must be defensible, auditable, and robust against silently fabricated evidence, and where transmitting requests and unpublished model details to a commercial LLM provider may be unacceptable, these are precisely the properties Distribird provides."
Distribird produces its best results under: Physically interpretable parameters
with a physical or biological interpretation that is discussed in the scientific literature
; Literature-active domain
with an active publishing community
; Univariate distributions
describable by a standard univariate probability distribution (Normal, truncated Normal, Beta, Gamma, Log-Normal)
; and Known physical bounds.
When conditions are not fully met, Distribird degrades gracefully: it states the evidence behind its output clearly, flags low-confidence results, and falls back to sensible uninformative alternatives.
It is not designed for parameters without physical interpretation (e.g. neural network weights), purely empirical tuning coefficients with no literature base, or parameters whose behaviour is fundamentally multivariate.
Improvements for AI systems
Based on the paper, here are specific improvements I can implement in AI systems:
-
Improvement: Add a mandatory citation-traceability layer to any AI system that generates scientific claims, parameter values, or data summaries. Every output must include a machine-readable chain linking each value to its source paper, DOI, and extraction context.
-
What the improved system can do: Generate priors, parameter estimates, or literature syntheses where every number is auditable—a reviewer can click through from the final distribution to the exact sentence in each cited paper. It refuses to emit unverifiable claims.
-
Improvement: Implement a two-gate validity classifier (early-skip + terminal check) that detects fabricated, typo-ridden, or non-empirical inputs before expensive processing. The system must decline to answer with a confident prior when the parameter has no empirical referent.
-
What the improved system can do: For nonsense inputs like
mumblesnort factororfake quantum correction xyz, it returnslikely invalidin seconds, saving 80–95% of compute. For real but non-empirical quantities (e.g., Kalman filter noise covariance), it flags them assuspiciousrather than inventing a distribution. -
Improvement: Replace single-pass pipelines with a cyclic LangGraph architecture that has three feedback loops (search refinement, domain broadening, extraction refinement) and a global iteration budget to guarantee termination.
-
What the improved system can do: When initial evidence is thin, it automatically widens search scope or refines queries, accumulating papers across tiers rather than discarding them. The budget cap (default 30 LLM calls) prevents runaway compute on pathological inputs.
-
Improvement: Add a per-source relevance judgment layer that labels each extracted value as high/medium/low relevance to the target domain, then weight values accordingly (1.0/0.6/0.25) during distribution fitting, with a confidence ceiling tied to relevance.
-
What the improved system can do: A porosity measurement from carbonate rock is down-weighted when calibrating a sandstone model, preventing off-domain values from dominating the prior. High confidence is only achievable when genuinely high-relevance evidence exists.
-
Improvement: Implement a four-tier fitting strategy based on the number of surviving values: AIC model selection (≥5 values, high confidence), moment matching with widened σ (2–4 values, medium), single-value centering (1 value, low), and uninformative fallback (0 values, none).
-
What the improved system can do: Automatically select between Normal, Truncated Normal, Gamma, Log-Normal, and Beta via AIC, while explicitly labeling the confidence level so downstream users never mistake an uninformative fallback for evidence-backed guidance.
-
Improvement: Configure all LLM calls to run on local open-weight models (e.g., Qwen3.6 27B, Gemma 4 31B) via quantized GGUF files, with no third-party API calls for inference.
-
What the improved system can do: Process unpublished modelling details, proprietary data, or sensitive parameter descriptions entirely on the researcher's hardware. Only generated search terms reach public literature databases (Semantic Scholar, OpenAlex), never the raw input.
-
Improvement: Use Markdown-based PDF extraction (PyMuPDF4LLM) that preserves tables as pipe-tables, headings, and page boundaries, rather than flattening to plain text.
-
What the improved system can do: Extract numerical values with their row labels, column headers, and unit associations intact, dramatically reducing extraction errors when values appear in calibration tables or appendices rather than prose.
-
Improvement: Expose a fixed seed parameter and per-task-class temperature settings (low for precise tasks like extraction, higher for creative tasks like enrichment), plus an optional debug-trace framework that records every prompt, response, and intermediate result.
-
What the improved system can do: Repeated runs on the same input converge to identical priors, enabling scientific reproducibility. The trace can be rendered as an HTML viewer, allowing any prior to be traced back to the exact model calls and evidence that produced it.
-
Improvement: Add a conditional forward path that, when at least two high-relevance papers are found, follows their citations to surface additional key papers before full-text retrieval.
-
What the improved system can do: Discover seminal works that direct keyword search might miss, expanding the evidence base by 20–30% on well-studied parameters (e.g., the incubation period example surfaced 18 additional papers).
-
Improvement: When zero values are extracted, return a wide Truncated Normal centered at the midpoint of bounds with σ = (u−l)/4, and set
is informative=falsewith confidencenone. -
What the improved system can do: Never silently produce a false-confidence prior. The user sees a clear warning that the prior is uninformative, preventing downstream misinterpretation in MCMC sampling.
What the fully improved system can do: Given a parameter name, description, and domain context, it autonomously searches and reads dozens of papers, extracts and relevance-weights numerical values, fits the best distribution via AIC, and returns a fully traceable prior with explicit confidence—all while refusing fabricated inputs, running entirely on local hardware, and producing reproducible results that a reviewer can audit end-to-end.
Abstract
Bayesian calibration of process-based models requires a prior distribution for each model parameter. Despite decades of methodological work, researchers almost always fall back on uniform priors. The main reason is that building informative priors from scientific literature is slow and needs both domain and statistical expertise. We present Distribird, an agentic web application that automates this process. Given a parameter name, physical description, and domain context, Distribird deploys a multi-agent pipeline that searches the literature, extracts and weights reported values by domain relevance, and fits a probability distribution via AIC model selection. When no literature is available, the system falls back to sensible uninformative alternatives, and clearly reports both the evidence behind and the confidence level of every prior it produces. It is designed for the problems where the models have physically interpretable parameters, where domain knowledge exists in the published literature. We evaluate the tool on 24 parameters across 10 scientific domains comparing three open-weight models (Qwen3.6 27B, Gemma 4 31B, Mistral Small 4 119B) with a single-prompt LLM baseline. On prior quality the full pipeline matches this baseline. Every prior is traced to the specific papers and values from which it was constructed; a built-in validity layer declines to produce priors for out-of-scope requests, whereas the single-prompt baseline returns confident but unfounded priors for them in 11 of 30 model--parameter cases; and every language-model call runs locally, so no parameter description or unpublished modelling detail is transmitted to a third-party LLM provider (only generated search terms reach the public literature databases). For scientific use, we argue these properties matter more than a marginal improvement in point-estimate accuracy.
Sources
- Emergent autonomous scientific research capabilities of large language models
- The No-U-Turn Sampler: Adaptively Setting Path Lengths in Hamiltonian Monte Carlo
- LLM-Prior: A Framework for Knowledge-Driven Prior Elicitation and Aggregation
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection