Reasoning or a Semblance of it? A Diagnostic Study of Transitive Reasoning in LLMs
cs.CL
Submitted: 2024-10-26
Updated: 2024-10-26
Comments: To appear in EMNLP Main 2024
DOI: 10.18653/v1/2024.emnlp-main.650
License: http://creativecommons.org/licenses/by/4.0/
The gist: Evaluating Large Language Models (LLMs) on reasoning benchmarks demonstrates their ability to solve compositional questions.
Terminology
Abstract
Evaluating Large Language Models (LLMs) on reasoning benchmarks demonstrates their ability to solve compositional questions. However, little is known of whether these models engage in genuine logical reasoning or simply rely on implicit cues to generate answers. In this paper, we investigate the transitive reasoning capabilities of two distinct LLM architectures, LLaMA 2 and Flan-T5, by manipulating facts within two compositional datasets: QASC and Bamboogle. We controlled for potential cues that might influence the models' performance, including (a) word/phrase overlaps across sections of test input; (b) models' inherent knowledge during pre-training or fine-tuning; and (c) Named Entities. Our findings reveal that while both models leverage (a), Flan-T5 shows more resilience to experiments (b and c), having less variance than LLaMA 2. This suggests that models may develop an understanding of transitivity through fine-tuning on knowingly relevant datasets, a hypothesis we leave to future work.
Sources
- Scaling Instruction-Finetuned Language Models
- Training Verifiers to Solve Math Word Problems
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Larger language models do in-context learning differently
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering