LogicSkills: A Structured Benchmark for Formal Reasoning in Large Language Models
cs.AI, cs.CL
Submitted: 2026-02-06
Updated: 2026-09-06
Comments: EMNLP 2026, 13 pages, 5 figures
Code: https://github.com/wo/tpg
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: Large language models perform well on many logical reasoning benchmarks, but it remains unclear which core logical skills they truly master.
Terminology
Abstract
Large language models perform well on many logical reasoning benchmarks, but it remains unclear which core logical skills they truly master. To address this, we introduce LogicSkills, a benchmark that isolates three fundamental logical skills: (i) formal symbolization translating premises into first-order logic; (ii) countermodel construction showing that an argument is logically invalid by constructing a finite countermodel; and (iii) validity assessment determining whether a conclusion follows from a set of premises. Items are drawn from the two-variable fragment of first-order logic without identity and are presented in both English and a Carrollian nonce-word language. All instances are solver-verified with Z3 for correctness and non-triviality. Across conventional instruction-tuned LLMs, performance is high on validity assessment but substantially lower on formal symbolization and countermodel construction, highlighting that high task-level accuracy can mask weaknesses in core logical skills. In contrast, recent reasoning-tuned models perform strongly across all three tasks, suggesting a more systematic logical skill profile.
Sources
- Propositional Interpretability in Artificial Intelligence
- Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models
- Logical Reasoning over Natural Language as Knowledge Representation: A Survey
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection