Psychosis involves a deficit of information compression in connected speech
cs.CL
Submitted: 2026-09-14
Updated: 2026-09-14
Code: https://github.com/tensorflow/tensor2tensor
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Large language models (LLMs) with human-like performance on linguistic tasks have transformed the study of language in neurodiverse conditions.
Terminology
Abstract
Large language models (LLMs) with human-like performance on linguistic tasks have transformed the study of language in neurodiverse conditions. LLMs provide representations of linguistic input in the form of high-dimensional vectors (embeddings), and next-token predictions computed from these embeddings. Previous crosslinguistic evidence suggests a complexity reduction in the form of both lower intrinsic dimensionality (ID) of LLM representations and higher mean surprisal (prediction error) in psychosis. We hypothesized that these metrics reflect a general deficit of information compression in psychosis, linked to grammatical organization as what enables predictions in language.We operationalized surprisal difference as the difference between surprisal as estimated from word frequency and surprisal as based on a contextual LM, which is sensitive to grammatical organization over and above lexical concepts. Using a dataset of 144 Turkish speakers, including 106 patients with schizophrenia-spectrum disorders (SSD) - 56 with chronic schizophrenia (SZH), 33 with first-episode psychosis (FEP), and 17 with schizoaffective disorder (SZA) - and 38 healthy controls. We report: (1) Surprisal difference is attenuated in all clinical groups relative to controls, independently of word count; (2) Compressibility (intrinsic dimension) is reduced in SZH and FEP; (3) Syntactic complexity and compressibility both predict surprisal difference. These results, further refining an alteration in the geometry of the semantic space in psychosis as previously attested, suggest a broader deficit in information compression in this disorder, with a mechanistic underpinning in the operations of grammar.
Sources
- The Geometry of Tokens in Internal Representations of Large Language Models
- The grip of grammar on meaning uncertainty: cross-linguistic evidence, neural correlates, and clinical relevance
- Distributional Results for Model-Based Intrinsic Dimension Estimators
- Lyapunov Spectral Analysis of Speech Embedding Trajectories in Psychosis
- Unveiling Intrinsic Dimension of Texts: from Academic Abstract to Creative Story
- Scale adaptive and robust intrinsic dimension estimation via optimal neighbourhood identification
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering