Playing log(N)-Questions over Wikipedia Abstracts: How Per-Round Errors Compound Under Information Asymmetry
cs.CL
Submitted: 2026-09-16
Updated: 2026-09-22
Comments: 29 pages
Code: https://github.com/ppotash/logn-questions
License: http://creativecommons.org/licenses/by/4.0/
The gist: We evaluate six frontier language models on the two-agent (N) -Questions game.
Terminology
Abstract
We evaluate six frontier language models on the two-agent (N) -Questions game. A questioner sees N Wikipedia lead paragraphs and must identify a secretly chosen target using exactly 2 N yes/no questions. An answerer sees only the target and the question, and replies with one word. Both roles run on the same provider, so the game measures how well a model communicates with itself across an information asymmetry. We run 408 games over document sets of 4 to 1024 paragraphs at a total API cost of 363. One model finishes well behind the others: Claude Opus 5 wins 28 of 68 games, against 45 to 56 for GLM-5.3, GPT-5.6 Sol, Grok 4.6, Gemini 3.8 Flash and Kimi K3. The leading five are only marginally separable. Pooling those five, win rate declines with set size at r=-0.973 and is fit by a single per-round reliability parameter. The form is win=p 2 N with p=0.928. Losses divide into answer errors and discrimination failures in roughly equal measure, and models almost never name a document their own evidence excludes. Every unanimous answer error from the weakest model was inspected: 32 of 34 are ``No'' answers, on properties stated in the document's first sentence, under an instruction that explicitly warns against defaulting to ``No''. Information per question, estimated from answer balance, correlates with win rate at r=+0.88. The only two models to extract a full bit per question are the only two that partition on document titles, a strategy absent below N = 32 and used in a quarter of questions above it. Reasoning-token expenditure varies 4.5 times across models with little relation to success, and the trace grows as the candidate set shrinks without a matching gain in reliability.
Sources
- Longformer: The Long-Document Transformer
- Generating Long Sequences with Sparse Transformers
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- Emergent Multi-Agent Communication in the Deep Learning Era
- MoBA: Mixture of Block Attention for Long-Context LLMs
- Playing log(N)-Questions over Sentences
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
- Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
- Inverse Scaling in Test-Time Compute
- Measuring AI Ability to Complete Long Software Tasks
- LLMs Get Lost In Multi-Turn Conversation
- LLM Evaluators Recognize and Favor Their Own Generations
- Codenames as a Benchmark for Large Language Models
- ToMBench: Benchmarking Theory of Mind in Large Language Models
- Revealing Hidden Mechanisms of Cross-Country Content Moderation with Natural Language Processing
- Constitutional AI: Harmlessness from AI Feedback
- Language Models (Mostly) Know What They Know
- NoLiMa: Long-Context Evaluation Beyond Literal Matching
- Kimi K3: Open Frontier Intelligence
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering