Quantifying Logical Consistency in Transformers via Query-Key Alignment
cs.CL, cs.AI, cs.IT, cs.LG, math.IT, math.LO
Submitted: 2025-02-24
Updated: 2025-02-24
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models (LLMs) have demonstrated impressive performance in various natural language processing tasks, yet their ability to perform multi-step logical reasoning remains an open challenge.
Terminology
Abstract
Large language models (LLMs) have demonstrated impressive performance in various natural language processing tasks, yet their ability to perform multi-step logical reasoning remains an open challenge. Although Chain-of-Thought prompting has improved logical reasoning by enabling models to generate intermediate steps, it lacks mechanisms to assess the coherence of these logical transitions. In this paper, we propose a novel, lightweight evaluation strategy for logical reasoning that uses query-key alignments inside transformer attention heads. By computing a single forward pass and extracting a "QK-score" from carefully chosen heads, our method reveals latent representations that reliably separate valid from invalid inferences, offering a scalable alternative to traditional ablation-based techniques. We also provide an empirical validation on multiple logical reasoning benchmarks, demonstrating improved robustness of our evaluation method against distractors and increased reasoning depth. The experiments were conducted on a diverse set of models, ranging from 1.5B to 70B parameters.
Sources
- Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla
- Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought
- Testing the General Deductive Reasoning Capacity of Large Language Models Using OOD Examples
- FOLIO: Natural Language Reasoning with First-Order Logic
- Reasoning Circuits in Language Models: A Mechanistic Interpretation of Syllogistic Inference
- Correcting Negative Bias in Large Language Models through Negative Attention Score Alignment
- Attention Heads of Large Language Models: A Survey
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering