Understanding LLM Failures: A Multi-Tape Turing Machine Analysis of Systematic Errors in Language Model Reasoning
cs.CL
Submitted: 2026-01-27
Updated: 2026-09-15
Comments: 8 pages, 1 page appendix; v2 added Acknowledgements; v3 is 7 pages. Substantially revised exposition and claims; title changed; related work and references updated; no new empirical results
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models (LLMs) exhibit failure modes on seemingly trivial tasks.
Terminology
Abstract
Large language models (LLMs) exhibit failure modes on seemingly trivial tasks. We propose a formalisation of LLM interaction using a deterministic multi-tape Turing machine, where each tape represents a distinct component: input characters, tokens, vocabulary, model parameters, activations, probability distributions, and output text. The model enables precise localisation of failure modes to specific pipeline stages, revealing, e.g., how tokenisation obscures character-level structure needed for counting tasks. The model clarifies why techniques like chain-of-thought prompting help, by externalising computation on the output tape, while also revealing their fundamental limitations. This approach provides a rigorous, falsifiable alternative to geometric metaphors and complements empirical scaling laws with principled error analysis.
Sources
- Scaling Laws for Neural Language Models
- The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
- Transformers Learn Shortcuts to Automata
- The Expressive Power of Transformers with Chain of Thought
- Autoregressive Large Language Models are Computationally Universal
- Why Do Large Language Models (LLMs) Struggle to Count Letters?
- Neural Turing Machines
- Are Transformers universal approximators of sequence-to-sequence functions?
- Training Compute-Optimal Large Language Models
- Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering