SABER-Math: Automated Benchmark for Information Retrieval Evaluation in Mathematics
cs.IR, cs.AI, cs.CL, cs.LG
Submitted: 2026-06-29
Updated: 2026-09-02
Comments: Accepted at The 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP), Hungary 2026, 34 pages
Code: https://github.com/project-numina/aimo-progress-prize
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: As agentic AI systems tackle more complex mathematical tasks, they increasingly rely on information retrieval (IR) to search problem databases, theorem libraries, and educational resources.
Terminology
Abstract
As agentic AI systems tackle more complex mathematical tasks, they increasingly rely on information retrieval (IR) to search problem databases, theorem libraries, and educational resources. However, choosing the right retriever remains difficult, as it is infeasible to directly isolate its effect on downstream performance. On the other hand, existing retrieval-specific benchmarks often fail to capture fine-grained mathematical relevance, penalizing relevant documents. We address this gap by introducing SABER-Math, the first fully automated benchmark for evaluating mathematical IR without expert annotation. Starting from 283K high-school-level math problems with solutions, SABER-Math builds challenging reranking tasks in three steps: (i) first, LLMs extract concise solution summaries and mathematical topics for each problem; (ii) then, per-query relevant documents are discovered using ontology topic-based and lexical solutions-summary-based similarities, and (iii) finally, a Swiss-style LLM preference tournament produces fine-grained relevance ratings for the documents. We evaluate lexical retrievers, specialized mathematical retrieval systems, and recent embedding models. We find that while modern embedding models substantially outperform classical and math-specific baselines, even the strongest systems struggle in symbol-heavy domains like Algebra and Calculus. Importantly, we show that general-purpose IR benchmarks such as MTEB do not reliably predict mathematical performance, especially for recent embedding models, highlighting the need for math-specific retrieval benchmarks.
Sources
- ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate
- gpt-oss-120b & gpt-oss-20b Model Card
- jina-embeddings-v5-text: Task-Targeted Embedding Distillation
- Training Verifiers to Solve Math Word Problems
- MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval
- Polyrating: A Cost-Effective and Bias-Aware Rating System for LLM Evaluation
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- Llama-Embed-Nemotron-8B: A Universal Text Embedding Model for Multilingual and Cross-Lingual Tasks
- SBI-RAG: Enhancing Math Word Problem Solving for Students through Schema-Based Instruction and Retrieval-Augmented Generation
- Large Language Models as Annotators: Enhancing Generalization of NLP Models at Minimal Cost
- Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
- From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review
- ChatGPT Outperforms Crowd-Workers for Text-Annotation Tasks
- RocqStar: Leveraging Similarity-driven Retrieval and Agentic Systems for Rocq generation
- Gemini Embedding: Generalizable Embeddings from Gemini
- Retrieval-augmented Generation to Improve Math Question-Answering: Trade-offs Between Groundedness and Human Preference
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- MIRB: Mathematical Information Retrieval Benchmark
- Automated Conjecture Resolution with Formal Verification
- From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
Related papers
- The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing
- SCAR: Semantic Continuity-Aware Retrieval for Efficient Context Expansion in RAG
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
- RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
- Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval
- UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG