FormalTCS: Benchmarking End-to-End Frontier Formal Theoretical Computer Science Research of Large Language Models
cs.CL
Submitted: 2026-08-20
Updated: 2026-09-23
Code: https://github.com/zirui-HIT/FormalTCS
License: http://creativecommons.org/publicdomain/zero/1.0/
The gist: Large language models (LLMs) have shown growing potential for automated theoretical computer science (TCS) research, yet existing benchmarks remain far from realistic research settings.
Terminology
Abstract
Large language models (LLMs) have shown growing potential for automated theoretical computer science (TCS) research, yet existing benchmarks remain far from realistic research settings. We introduce, an expert-validated benchmark for evaluating LLMs on frontier, end-to-end TCS research. contains 143 instances drawn from papers accepted to STOC, FOCS, SODA, and COLT in 2025-2026, preserving paper-specific definitions, assumptions, and proof dependencies, with expert-verified Lean formalizations and proofs. Evaluations of leading LLMs reveal that current models remain far from reliably completing the full research pipeline. In particular, autoformalization is the sharpest bottleneck: the best model achieves only 11.5 on translating natural-language claims into formal theorem statements, compared with 28.6 Pass@8 when proving human-provided formal statements. Building on, we further develop an automated TCS research framework that generates, formalizes, filters, and proves new claims. Of 64 generated claims, only 6 ultimately pass expert evaluation and proof verification, indicating that beyond formalization, limited research taste remains another major barrier to autonomous TCS research.
Sources
- Bolzano: Case Studies in LLM-Assisted Mathematical Research
- AI4Research: A Survey of Artificial Intelligence for Scientific Research
- TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research Ability
- Towards Autonomous Mathematics Research
- Theory-Scale Auto-Formalization of Logics for Computer Science
- Time Travel in LLMs: Tracing Data Contamination in Large Language Models
- The Surprising Effectiveness of Membership Inference with Simple N-Gram Coverage
- BlueprintRepair: Typed Local Edits for Failed Lean Proof Blueprints
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
- DeepSeek-Prover-V2: Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal Decomposition
- Accelerating Scientific Research with Gemini: Case Studies and Common Techniques
- LeanMarathon: Toward Reliable AI Co-Mathematicians through Long-Horizon Lean Autoformalization
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering