Evidence-Based Scientific Question Discovery: A Framework with Historical Backtesting
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Evidence-Based Scientific Question Discovery".
Jane: This paper presents a framework for scientific question discovery, which the authors argue is a distinct and bottlenecked capability within the scientific enterprise.
Tom: First, who's behind it and why it matters.
Paper discussion segment 1 — Tom and Jane discuss title and authors: Tom: We've talked about the concept, but let's look at who wrote this. The paper, "Evidence-Based Scientific Question Discovery: A Framework with Historical Backtesting," was put out by Hui Mao, an independent researcher from UPenn.
Jane: That’s interesting because it comes from a researcher who seems deeply embedded in the scientific community. When you hear "historical backtesting" in the title, it sounds like they’re testing the system against real, established science before they even use it for new things.
Lu: The authors are clearly drawing on existing ideas about how AI can perform ideation, looking at things like novelty search and intrinsic motivation which we see in other related work. It shows a focus on structuring the evidence itself rather than just feeding raw data into a black box.
Meng: So, this isn't just another wrapper around an existing large language model; it seems to be about creating a specific pipeline for scientific inquiry that respects the structure of how scientists actually talk to each other.
Lalam: Exactly, and that structure is what matters. When we look at how this framework handles evidence as provenance-carrying claims, it’s a direct attempt to model the social and epistemic process of scientific dialogue.
Paper discussion segment 2 — Tom and Jane discuss paper's summary and implications: Tom: So, we know it turns a corpus into ranked questions, but what’s the actual mechanism? The core idea is representing evidence as claims, detecting tensions between papers that share objects or topics, and then having humans adjudicate those tensions before refining them into actual research questions.
Jane: That tension detection part sounds incredibly sophisticated. It suggests that AI can be trained to spot subtle conflicts or agreements between different pieces of literature without just flagging everything as noise.
Lu: The way they type those tensions—using words like "contradicts" or "qualifies"—is brilliant because it moves beyond simple binary conflict; it captures the nuance of scientific disagreement. That vocabulary itself is an empirical result of their calibration review, which suggests a data-driven approach to defining scientific relationships.
Meng: From an engineering standpoint, I wonder how robust that human adjudication step is going to be when you scale it up across millions of papers. It’s a critical dependency for the whole system's quality.
Lalam: That human adjudication is where the cultural impact really hits. If we can automate the initial sorting and refinement of these tensions, it means the heavy lifting of reading and synthesizing literature is done, allowing scientists to focus on interpreting what those refined questions actually mean.
Paper discussion segment 3 — Tom and Jane discuss improvements and implications: Tom: The authors propose several improvements, and the historical backtest results are really compelling. They showed that questions generated before two thousand twenty-one were substantively engaged by literature published between two thousand twenty-one and two thousand twenty-six proving the system’s predictive power against unseen data.
Jane: That historical validation is a huge deal because it provides a way to measure the quality of question discovery objectively against future research, which is something we haven't really done before. It turns question value into a measurable outcome.
Lu: They also break down the ranking into two stages: scientific priority and execution priority, where they define a specific formula involving significance, tension strength, feasibility, novelty, and information gain. That separation is crucial for guiding research strategy.
Meng: The separation of priorities is what I find most useful for practical application. It lets you prioritize questions that are scientifically high-impact even if they are operationally hard, rather than getting stuck on easy but meaningless tasks.
Lalam: And the improvements they suggest, like adding a clarity gate for questions and checking for data sufficiency, directly address the failure modes they found in their initial run. It shows that refining the question generation process itself is just as important as the initial evidence gathering.
Conclusion: Tom: So, to wrap this up, this paper on "Evidence-Based Scientific Question Discovery: A Framework with Historical Backtesting" shows how we can move AI beyond just answering known questions and into the realm of proposing new ones. It’s about building a traceable system that respects the structure of scientific knowledge.
Jane: Right, it’s a framework that turns evidence into ranked, falsifiable questions by rigorously checking tensions and testing those questions against future literature to see if they hold up. The implication is that we can start automating the very first step of scientific inquiry in a highly structured way.
Lu: It gives us a concrete way to formalize "interestingness" not just through an agent's internal experience, but through the structure of the existing scientific record. This moves us closer to having AI act as a true co-scientist by analyzing the evidence stream itself.
Meng: From my side, it means we can focus our engineering efforts on building better mechanisms for tension detection and that two-stage ranking protocol, because those are where the measurable performance gains will actually happen in a real research setting.
Lalam: Ultimately, this work shows how we can improve the culture of science by giving researchers tools that help them see the landscape of unsolved problems more clearly, proving that question-asking competence can be decomposed into auditable and measurable stages.
Hui Mao
cs.DL, cs.AI
Submitted: 2026-07-29
Updated: 2026-08-12
Comments: 13 pages, 4 figures
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 62/100
Key concepts
- Evidence as Claims
- The framework represents evidence not just as raw data, but as claims. This means treating pieces of literature as assertions that can be compared and analyzed for relationships between different scientific ideas.
- Tension Detection
- This is a sophisticated mechanism where the AI is trained to spot subtle conflicts or agreements between papers sharing topics. It looks for nuanced disagreements, using specific vocabulary like 'contradicts' or 'qualifies,' rather than just simple binary conflicts.
- Two-Stage Ranking Protocol
- The system ranks questions in two stages: scientific priority and execution priority. This separation helps guide research strategy by allowing prioritization of high-impact questions even if they are operationally difficult to execute.
Terminology
Summary
Summary
This paper presents a framework for scientific question discovery, which the authors argue is a distinct and bottlenecked capability within the scientific enterprise. The authors state: “Current AI systems are optimized for answering questions; the scientific enterprise is bottlenecked earlier, at discovering the questions worth investigating.” The framework transforms a traceable, reproducible, scope-controlled research corpus into ranked, falsifiable research questions. The core methodology involves representing evidence as provenance-carrying claims, detecting and typing cross-paper tensions, human adjudication of those tensions, refining surviving signals into questions, and ranking them via a two-stage protocol that separates scientific priority from execution priority.
The framework is instantiated on exoplanet atmospheres, chosen because it “uniquely combines literature, structured catalogs, and space-telescope archives.” The central evaluation is a historical backtest: questions are generated from evidence available before a cutoff date (2020-12-31) and judged against literature published after that cutoff (2021–2026), which the system never saw. The authors report that “all questions generated from evidence available before 2021 were substantively engaged by the 2021–2026 literature the system never saw: two were answered—including one whose premise the community later explicitly refuted—and the top-ranked question is independently posed and still open.”
The framework consists of six stages. First, evidence representation: knowledge enters as three layers (literature, catalog, observation), with the atomic unit being a claim tuple: c = (text, evidence, assumptions, uncertainty, objects, datasets, location, tier), where tier records verification provenance (human gold, model matched gold, or model only). Second, an evidence graph with typed tensions: claims, objects, datasets, and assumptions become nodes; cross-paper claim pairs sharing an object and topic term become tension candidates. Crucially, candidate tensions are typed by human adjudication using the vocabulary contradicts, qualifies, challenges-method, explains-discrepancy, complements, supports. This vocabulary is described as “an empirical result of the calibration review,” since only 1 of 16 top tension candidates was a genuine contradiction. Third, question generation: questions are generated only from prioritized signal classes—(P1) human-confirmed observational tensions, (P2) methodological challenges, (P3) qualifications across independent datasets, and (P4) single-dataset conclusions from trusted claims. Fourth, two-stage ranking: Stage A is a clarity gate (presupposed conclusions, statistics/physics conflation, explicit comparison, falsifiable outcome, bounded scope, stateable failure condition); Stage B ranks survivors by scientific priority = 0.35 S + 0.25 T + 0.20 F + 0.10 N + 0.10 G, where S (significance) is hard-capped by signal tier, T is tension strength, F is feasibility, N is novelty (deliberately down-weighted), and G is expected information gain. Execution priority is reported separately.
Corpus discipline is enforced via frozen manifests that pin scope, queries, versions, and a cutoff date. The historical validation protocol involves: (1) Corpus A containing only evidence up to the cutoff, enforced end-to-end with a contamination guard; (2) running the pipeline to produce ranked questions; (3) collecting a validation corpus of post-cutoff literature under its own manifest, retrieving relevant abstracts by embedding similarity, and classifying each question’s fate as answered, partially addressed, posed but open, or not addressed, plus stating whether the premise was validated, refuted, or left untested.
The experiments use manifest exoplanet atmospheres v1 (cutoff 2020-12-31). Corpus A contains 2,512 papers (literature metadata), 500 core full-text papers, 51,616 citation edges, 6,324 catalog objects, 306,729 observation metadata records, 2,364 paper–object links, and 553 catalog planets with archival spectra. Claim extraction was calibrated: a domain reviewer annotated 20 papers yielding 37 gold claims; model extraction was calibrated over three prompt revisions, with the final prompt enforcing 2–5 load-bearing claims with a priority order (conclusion > interpretation > methodological challenge > decisive quantity > secondary measurement). Semantic deduplication matched all 37 gold claims to model counterparts, producing a canonical table of 353 claims (37 model matched gold, 316 model only).
Tension detection produced 161 cross-paper candidates over an 807-node, 1,237-edge graph. Human adjudication of the top 16 candidates yielded: 7 supports, 4 complements, 2 challenges-method, 1 qualifies, 1 explains-discrepancy, and 1 contradicts. The single confirmed contradiction is the HD 189733 b wind-velocity discrepancy: ∼8 km s−1 from optical sodium (Wyttenbach et al., 2015) versus ∼1.7 km s−1 from near-infrared CO/H2O (Brogi et al., 2016).
Question generation produced 11 candidates; editorial curation merged one near-duplicate pair and rewrote three questions. All 10 questions passed the Stage-A gate. Stage-B ranking placed the terminator-heterogeneity family and the wind-discrepancy question statistically tied at the top (scientific priority 8.13 vs. 8.13), with opposite execution profiles (execution priority 9.1 for the heterogeneity family vs. 6.9 for the wind discrepancy). The seven P4 robustness questions ranked in a 6.6–7.1 band.
Historical validation used 1,891 unique 2021–2026 papers. Results: 2 answered, 1 posed but open, 7 partially addressed, 0 not addressed. Three outcomes are emphasized. First, a premise refuted: the question of whether HD 209458 b’s strongly subsolar terminator water abundance (MacDonald and Madhusudhan, 2017) is atmospheric reality or retrieval artifact was answered by 2025 reanalyses finding solar-consistent water. Second, the top question (WASP-12 b terminator heterogeneity vs. water evidence) is independently posed and open. Third, the TRAPPIST-1 CO2 detectability prediction question is being tested by JWST/NIRSpec PRISM programs, with stellar contamination currently blocking a decisive verdict.
Failure analysis identified three failure modes for the seven partial verdicts: (1) blocked by a shared upstream obstacle (e.g., stellar contamination), with the failure source being that “the falsification stage checks data availability but not data sufficiency”; (2) questions demanding joint analyses no single study performs (e.g., the wind-discrepancy question needs contemporaneous optical–near-infrared spectroscopy); (3) questions broader than any single study answers (class-level conditions and multi-factor sensitivity questions). The authors also note two selection effects contributing to zero not-addressed outcomes: tensions among highly-cited claims concern heavily-observed targets, and the validation corpus used target-specific queries.
Threats to validity include: single domain (astronomy only), abstract-only extraction, single human adjudicator, judge bias (LLM judge sharing a model family with extraction and ranking stages), and small sample size with no baseline. The authors state the results “support a specific, testable position: question-asking competence can be decomposed into auditable stages—evidence representation, tension detection, refinement, prioritization—each of which can be measured and improved independently, with historical backtesting as the end-to-end score.”
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in an AI system, and what the improved system can do:
1. Time-Frozen Corpus with Contamination Guard
-
Improvement: Add a
cutoff dateparameter to the system's data ingestion pipeline. All training, retrieval, and generation processes must filter every data source (literature, catalogs, archives) by this timestamp. Store corpus manifests immutably; any scope change mints a new corpus ID rather than mutating the old one. -
What it can do: The system can generate questions or hypotheses from a strictly time-bounded knowledge state, then be evaluated against post-cutoff literature it never saw. This eliminates data leakage and enables historical backtesting as a built-in evaluation metric.
2. Provenance-Carrying Claim Extraction
-
Improvement: Replace plain-text extractions with a structured claim tuple:
(text, evidence, assumptions, uncertainty, objects, datasets, location, tier). Enforce a priority order for extraction (conclusion > interpretation > methodological challenge > decisive quantity > secondary measurement). Merge coupled measurements into single claims. Normalize to third person. Run an omission self-check after each extraction. -
What it can do: Every generated question or hypothesis carries a full evidence trail—source document, exact location, assumptions, and verification tier (human gold, model matched gold, model only). This makes every output auditable and reproducible, and prevents the system from manufacturing spurious conflicts by duplicating claims.
3. Typed Tension Detection with Human Adjudication
-
Improvement: Replace binary
contradicts/does not contradict
classification with a six-type vocabulary:contradicts, qualifies, challenges-method, explains-discrepancy, complements, supports. Run a high-recall candidate rule first (shared object + shared topic term), then require human adjudication on the top-ranked candidates before any tension can drive question generation. Commit all human verdicts as data that feeds back into the evidence graph. -
What it can do: The system avoids the common failure of collapsing
adds limiting conditions
intonegates.
In the paper's calibration, only 1 of 16 top candidates was a true contradiction; the rest were supports, complements, or methodological challenges. The system can now distinguish between a genuine conflict (e.g., factor-of-five wind discrepancy) and a benign complementarity (e.g., terminator vs. dayside measurements), preventing false-positive question generation.
4. Two-Stage Ranking Separating Scientific and Execution Priority
-
Improvement: Implement Stage A as a clarity gate (presupposed conclusions, statistics/physics conflation, explicit comparison, falsifiable outcome, bounded scope, stateable failure condition)—binary pass/fail, not a score. Implement Stage B as a weighted scientific priority score:
0.35*Significance + 0.25*TensionStrength + 0.20*Feasibility + 0.10*Novelty + 0.10*InformationGain. Hard-cap Significance by signal tier (P1: 10, P2: 9, P3: 8.5, P4: 7.5) so single-dataset robustness checks cannot outrank confirmed tensions. Report Execution Priority separately. -
What it can do: The system produces a ranked list where scientific importance and operational difficulty are explicitly decoupled. A question can be scientifically first (e.g., the wind-discrepancy question) while being operationally harder (execution 6.9 vs. 9.1 for a reanalysis-ready question). This prevents the system from recommending easy-but-trivial work over hard-but-important work.
5. Historical Validation with Premise Refutation Detection
-
Improvement: After generating questions from a pre-cutoff corpus, collect a post-cutoff validation corpus under a separate manifest. For each question, retrieve the most relevant post-cutoff abstracts by embedding similarity. Classify each question's fate as
answered,partially addressed,posed but open, ornot addressed. Additionally, classify the premise asvalidated,refuted, oruntested. Force the judge to cite only retrieved records. -
What it can do: The system can measure its own question quality against ground truth—what the field actually did. The strongest success mode is a refuted premise: the system flagged a published conclusion as fragile (e.g., HD 209458 b's subsolar water) before the community overturned it. This turns question discovery from a subjective exercise into a measurable, outcome-based benchmark.
6. Falsification Screen with Data-Sufficiency Check
-
Improvement: Extend the falsification stage beyond
does archival data exist
todoes the data have sufficient signal-to-noise and known systematics handling to actually test the question.
Consult the evidence graph for known systematics on the target (e.g., stellar contamination for M-dwarfs) before generating the question. If an obstacle is known, generate the question with the obstacle named and a contamination-robust test demanded. -
What it can do: The system avoids generating questions that the community will engage but cannot answer due to a shared upstream obstacle (e.g., TRAPPIST-1 CO2 detectability blocked by stellar contamination). It either flags the obstacle upfront or reframes the question to demand a test that is robust to it.
7. Granularity Matching for Question Scope
-
Improvement: Add a granularity criterion to the clarity gate: the question's scope must match the granularity at which the field publishes. A question spanning a full parameter space (e.g.,
conditions for patchy-cloud fits
) should be split into instance-level sub-questions that a single study can address. Flag questions that are broader than any single study can answer. -
What it can do: The system produces questions that are actionable at the level of a single paper, rather than questions that require a coordinated multi-study program. This reduces the
partially addressed
failure mode where the community advances pieces but never answers the question as posed.
8. Baseline-Controlled Engagement Metrics
-
Improvement: When reporting engagement rates (e.g.,
0 of 10 questions ignored
), also run two baselines: (a) random future-work questions sampled from review papers, and (b) LLM-generated questions without the evidence graph. Compare engagement rates across all three. Report the confound of target popularity (highly-cited targets are observed intensively regardless of question quality). -
What it can do: The system's headline metric—
all questions engaged by subsequent literature
—becomes interpretable. Without baselines, the 100% engagement rate could be explained by selection effects (popular targets, targeted validation queries). With baselines, the system can claim a genuine effect size, not just an existence claim.
Summary of what the improved AI system can do:
-
Generate falsifiable research questions from a time-frozen corpus, with full provenance and auditable evidence trails.
-
Distinguish genuine contradictions from benign complementarities, methodological challenges, and qualifications—reducing false-positive question generation.
-
Rank questions by scientific priority separately from execution feasibility, preventing easy-but-trivial work from outranking hard-but-important work.
-
Evaluate its own question quality against post-cutoff literature, including detecting when a question's premise was later refuted (the strongest success signal).
-
Avoid generating questions blocked by known systematics (e.g., stellar contamination) by checking data sufficiency, not just data availability.
-
Produce questions at the granularity of a single publishable study, reducing the
partially addressed
failure mode. -
Measure its engagement rate against baselines, making its performance claims statistically meaningful rather than anecdotal.
Sources
- Accelerating scientific discovery with Co-Scientist
- PaperQA: Retrieval-Augmented Generative Agent for Scientific Research
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- OpenAlex: A fully-open index of scholarly works, authors, venues, institutions, and concepts
- Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
- Language agents achieve superhuman synthesis of scientific knowledge
- Galactica: A Large Language Model for Science