PBEBench: A Multi-Step Programming by Examples Reasoning Benchmark inspired by Historical Linguistics
cs.CL
Submitted: 2025-05-29
Updated: 2026-09-16
License: http://creativecommons.org/licenses/by/4.0/
The gist: Although many benchmarks evaluate the reasoning abilities of Large Language Models (LLMs) within domains such as mathematics, coding, or data wrangling, few abstract away from domain specifics to
Terminology
Abstract
Although many benchmarks evaluate the reasoning abilities of Large Language Models (LLMs) within domains such as mathematics, coding, or data wrangling, few abstract away from domain specifics to examine reasoning as a capability in and of itself. We contribute a novel type of benchmark evaluating the inductive reasoning capabilities of LLMs that is inspired by the forward reconstruction task from historical linguistics but is formulated in an extremely simple, general way (in the form of Programming by Examples). The task involves generating a cascade of simple string rewrite programs to transform a given list of input strings into a list of desired output strings. We present a fully automated pipeline that programmatically generates problems of this type with controllable difficulty, enabling scalable evaluation of reasoning models while avoiding contamination. Using this approach, we construct two benchmarks: PBEBench-Lite, which efficiently stratifies models of varying capabilities, and PBEBench, which requires models to induce programs similar in complexity to those constructed by historical linguists. Our experiments reveal a substantial performance gap between models that leverage test-time compute or LCoT (long chain-of-thought) reasoning and those that do not. Moreover, although recent models show promise, the solve rate for both of them drops below 5% for hard instances of the PBEBench dataset (ground truth cascade lengths of 20 and 30, respectively), falling well short of realistic historical linguistics requirements even with computationally expensive, popular scaling techniques from the PBE and reasoning literature. Additionally, we also study the effectiveness of different scaling strategies and the impact of various hyperparameters on the difficulty of the generated data using gpt-oss-120b, the best-performing open-source model.
Sources
- Program Synthesis with Large Language Models
- Evaluating Large Language Models Trained on Code
- On the Measure of Intelligence
- De Finetti for mathematics undergraduates
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- RobustFill: Neural Program Learning under Noisy I/O
- Comment on The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
- Large Language Models as Test Case Generators: Performance Evaluation and Enhancement
- KOR-Bench: Benchmarking Language Models on Knowledge-Orthogonal Reasoning Tasks
- DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
- s1: Simple test-time scaling
- Programming by Examples Meets Historical Linguistics: A Large Language Model Based Approach to Sound Law Induction
- Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks
- Can Large Language Models Code Like a Linguist?: A Case Study in Low Resource Sound Law Induction
- Code2Bench: Scaling Source and Rigor for Dynamic Benchmark Construction
- gpt-oss-120b & gpt-oss-20b Model Card
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
- CLUTRR: A Diagnostic Benchmark for Inductive Reasoning from Text
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering