Inspire: Benchmarking Scientific Literature Search for Open Research Problems
cs.IR, cs.AI
Submitted: 2026-09-27
Updated: 2026-09-27
Terminology
Sources
- PreScience: A Dataset and Benchmark for Scientific Forecasting
- SciPaths: Forecasting Pathways to Scientific Discovery
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
- MIR: Methodology Inspiration Retrieval for Scientific Research Problems
- Multi-Turn Agentic Scientific Literature Search via Workflow Induction
- ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition
- WebGPT: Browser-assisted question-answering with human feedback
- BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models
- Learning to Predict Future-Aligned Research Proposals with Language Models
- BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
- Scientific reasoning does not reliably translate into scientific forecasting in frontier AI
- AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery
- HyBIRD: Hyperbolic Bridge Retrieval and Diagnosis for Methodology Inspiration Retrieval
- ReAct: Synergizing Reasoning and Acting in Language Models
- Detail Matters: Mamba-Inspired Joint Unfolding Network for Snapshot Spectral Compressive Imaging
Related papers
- The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing
- SCAR: Semantic Continuity-Aware Retrieval for Efficient Context Expansion in RAG
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
- RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
- Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval
- UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG