Word-Class and Construction-Like Structure Emerges in Neural Successor Representations Trained on Natural Language

arXiv:2605.24585 · cs.CL, q-bio.NC · Submitted 2026-05-23 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Word-Class and Construction-Like Structure Emerges in Neural Successor Representations Trained on Natural Language".

Jane: The gist The central result of this work is that SRs, trained as multi-horizon predictive representations over a large naturalistic text corpus,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So we're looking at this paper today titled "Word-Class and Construction-Like Structure Emerges in Neural Successor Representations Trained on Natural Language". It sounds a bit dense, but the core idea is that they're moving beyond just predicting the next word.

Jane: Right, it’s about these things called Successor Representations. Instead of just looking at what comes right after a word, these models are trained to predict the expected distribution of words coming in several steps ahead.

Lu: The authors are taking a framework from reinforcement learning and applying it to natural language data from WikiText-one hundred three which is one hundred three million tokens <ref:2605.24585#pg1>. They're training a deep residual neural network on that massive dataset.

Meng: So, the big question here is whether this approach can actually capture the underlying structure of language without needing people to manually label every single part of speech or grammar rule during training.

Lalam: It suggests that linguistic structure might emerge naturally just from learning how sequences transition over time, instead of being explicitly taught.

The paper's summary: Tom: The paper argues that this way of predicting future word distributions allows the model to learn about the long-range transition structure of language, which is a really big deal compared to just next-token prediction.

Jane: They found that after training, the learned representation space organizes itself geometrically in a way that mirrors how language is structured, specifically around part of speech categories.

Lu: They showed that nouns, verbs, and adjectives become recoverable through unsupervised clustering even though no linguistic labels were used during the training process itself.

Meng: That’s significant because it means the model discovered these categories on its own based purely on the predictive dynamics of the text.

Lalam: It’s like it learned grammar by observing how words follow each other over longer stretches of text, which is a different way to learn than just memorizing patterns.

The paper's improvements: Tom: One key improvement they make is using KL divergence instead of just direct regression when training the model. This helps them optimize those successor representations as probability distributions across multiple time horizons.

Jane: That optimization method lets them learn about the long-range transition structure more effectively, which is what they call a predictive principle derived from reinforcement learning.

Lu: They also use this approach to show that syntactic categories aren't something that needs to be explicitly encoded or supervised; they emerge as a consequence of this predictive sequence learning process alone.

Meng: So, the improvement isn't just the prediction method, it’s proving that structure can emerge spontaneously without any prior linguistic annotations guiding the model.

Lalam: It provides this conceptual bridge between reinforcement learning and linguistics, suggesting that these two fields are more connected than we thought in how they understand sequence modeling.

Conclusion: Tom: So, to wrap up, the main point of "Word-Class and Construction-Like Structure Emerges in Neural Successor Representations Trained on Natural Language" is that structural categories can appear as an emergent consequence of predictive sequence learning alone.

Jane: The paper shows that when you train a model to predict future word distributions over multiple horizons, the resulting representation space naturally organizes itself around grammatical structures like nouns and verbs.

Lu: This work establishes a conceptual bridge between reinforcement learning and linguistics by showing how these two different fields can interact in understanding language.

Meng: From an engineering side, it shows we don't need to hand-code all the grammatical rules if the predictive objective is set up correctly to learn those structures for us.

Lalam: It suggests that constructional knowledge isn't some separate linguistic module, but rather an instance of a more general predictive memory system operating over structured sequences.

Mathis Immertreu, Achim Schilling, Thomas Kinfe, Patrick Krauss

Cognitive Computational Neuroscience Group, Friedrich-Alexander-Universität Erlangen–Nürnberg (FAU) · Mannheim Center for Neuromodulation and Neuroprosthetics (MCNN), University Hospital Mannheim, University Heidelberg

cs.CL, q-bio.NC

Submitted: 2026-05-23

Updated: 2026-10-05

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 86/100

The gist: The gist The central result of this work is that SRs, trained as multi-horizon predictive representations over a large naturalistic text corpus, spontaneously organise into a geometry that mirrors

Key concepts

Successor Representations (SRs)
These are neural representations trained not to predict the immediate next word in a sequence, but rather the expected discounted distribution of future words. This models what is likely to happen in the future across several time steps, capturing long-range transition structures within language.
Temporal Prediction Horizon ($\gamma$)
This parameter controls how far into the future the model looks when predicting word distributions. Shorter horizons capture local syntactic rules, while longer horizons integrate broader contextual and semantic information from a wider window of text.

Terminology

Summary

The gist The central result of this work is that SRs, trained as multi-horizon predictive representations over a large naturalistic text corpus, spontaneously organise into a geometry that mirrors the hierarchical linguistic structure of language, namely its constructional organisation, without any explicit supervision.

How it works

  1. Successor Representations (SRs) model not the immediate next state but the expected discounted distribution of future states instead of predicting the next token in a sequence The paper explores an alternative predictive principle derived from reinforcement learning, which models what is expected to occur in the future and how often

  2. The model is trained to predict the expected discounted distribution of future words across multiple temporal horizons instead of predicting the next token This objective learns representations of the long-range transition structure of language

  3. The training involves treating SR targets as probability distributions using KL divergence rather than direct regression, optimizing a deep residual neural network on WikiText-103 This allows linguistic structure to emerge spontaneously without explicit supervision or annotation of linguistic categories

Emergent Linguistic Structure

The model develops a pronounced geometric organization with respect to part-of-speech (POS) categories after training, where nouns, verbs, and adjectives become recoverable through unsupervised clustering The temporal prediction horizon systematically shapes representation geometry, as short horizons preserve local syntactic regularities while longer horizons integrate broader contextual and semantic information

Recovering Basic Syntactic Categories

The unsupervised clustering of SR embeddings at a low discount factor γ = 0.2 recovers the three major categories of nouns, verbs, and adjectives with per-cluster POS purities of 0.89, 0.91, and 0.85 respectively The inter-cluster transition network replicates well-motivated grammatical asymmetries: strong ADJ→NOUN flow (67%) reflects the modifier–head construction, and dominant NOUN→VERB connectivity (52%) encodes the subject–predicate template

Constructional Slots as Lossy Clusters of Memory Traces

Increasing the number of clusters to k = 30 reveals semantically coherent sub-classes, such as ordinal modifiers for adjectives and event-type classes for verbs, which are interpretable as distributional signatures of constructional slots This structure suggests that the SR space encodes not only category membership but also the constructional contexts items inhabit, consistent with usage-based accounts of grammar

Limitations and Future Directions

A practical limitation is the absence of a principled criterion for selecting the appropriate clustering resolution k or temporal horizon γ, as higher values of γ integrate over longer contextual windows and reduce cluster coherence Furthermore, the inability to distinguish between distinct senses of a surface form due to type-level representations necessitates a shift toward contextualised SR estimation conditioned on preceding context

The SR space instantiates precisely such a space, and the clusters that emerge from it can be understood as such compressions: each cluster summarises a family of partially overlapping usage traces whose shared successor profile reflects the combinatorial privileges conferred by repeated participation in similar contexts

This multi-granular alignment provides a topological metric for the theoretical distinction between errors and innovation, quantifying exactly how and where a rule is broken The SR space thus encodes the full slot-and-filler architecture predicted by usage-based accounts

The model's ability to recover these categories so robustly without explicit supervision suggests that structural categories are emergent consequences of predictive sequence learning rather than intrinsic lexical properties This demonstrates that the SR space encodes the full slot-and-filler architecture predicted by usage-based accounts

The paper's findings support a unifying hypothesis that constructional knowledge is not a specialized linguistic module, but an instance of a more general predictive memory system operating over structured sequences The SR may serve as a computational bridge between the constructional organization of language and the neuro-scientific accounts of the memory system in which that organization is grounded

The results confirm that structural categories can be discovered directly from sequential statistics, offering a data-driven tool for crosslinguistic analysis that organically maps a language’s true emergent categories without requiring any prior commitment to traditional taxonomies The SR space encodes the full slot-and-filler architecture predicted by usage-based accounts

The SR transition structure provides a concrete, mathematical mechanism to bridge the gap between abstract template planning and concrete lexical instantiation The model's ability to recover these categories so robustly without explicit supervision suggests that structural categories are emergent consequences of predictive sequence learning rather than intrinsic lexical properties The SR space encodes the full slot-and-filler architecture predicted by usage-based accounts

The SR geometry captures the sequential links that bind elements together into constructions (syntagmatic relations) as well as the equivalence classes of items that can substitute into the same structural slots (paradigmatic relations) This view suggests that category membership is not an intrinsic lexical property but an emergent consequence of the positional and combinatorial privileges a form repeatedly occupies across constructions

The SR space instantiates precisely such a space, and the clusters that emerge from it can be understood as such compressions: each cluster summarises a family of partially overlapping usage traces whose shared successor profile reflects the combinatorial privileges conferred by repeated participation in similar contexts This view suggests that category membership is not an intrinsic lexical property but an emergent consequence of the positional and combinatorial privileges a form repeatedly occupies across constructions

The SR space encodes the full slot-and-filler architecture predicted by usage-based accounts The model's ability to recover these categories so robustly without explicit supervision suggests that structural categories are emergent consequences of predictive sequence learning rather than intrinsic lexical properties

The SR transition structure provides a concrete, mathematical mechanism to bridge the gap between abstract template planning and concrete lexical instantiation The model's ability to recover these categories so robustly without explicit supervision suggests that structural categories are emergent consequences of predictive sequence learning rather than intrinsic lexical properties The SR

Improvements for AI systems

  1. Successor Representation Training for Linguistic Structure Emergence: Implement a multi-horizon predictive training objective using KL divergence to model the expected discounted distribution of future states across multiple temporal horizons, which allows the system to learn the long-range transition structure of language rather than just immediate next-token prediction.

  2. Spontaneous Part-of-Speech Recovery via Unsupervised Clustering: Leverage the learned representation space by applying KMeans and BGM clustering on SR embeddings to recover linguistic categories, as shown in Figure 4, demonstrating that nouns, verbs, and adjectives become recoverable through unsupervised clustering despite the absence of linguistic labels during training.

  3. Context-Sensitive Constructional Slot Modeling: Utilize high-resolution clustering (e.g., k=120) to identify constructional slots by analyzing inter-cluster transition weights, enabling the system to model the full slot-and-filler architecture predicted by usage-based accounts, such as distinguishing between bare cardinals and measurement units and calendar years and month names.

  4. Scale-Dependent Creativity Modeling: Apply the SR geometry to quantify creativity by observing how it models deviation: a collo-creative utterance—such as an unexpected lexical filler—corresponds to a drop in SR transition weights to near-zero at a fine-grained, item-specific level (such as k=120), which allows the system to distinguish between deliberate, creative deviations from unintelligible errors.

  5. Multiscale Planning Substrate: Design an architectural extension where the hierarchical structure of the SR embedding space could serve as a substrate for multiscale planning in language production, allowing the model to commit to a high-level frame (e.g., [MONTH YEAR] frame) while using fine-grained transition probabilities to constrain the selection of a specific month name and year token in sequence."

Abstract

Neural language models are typically trained on next-token prediction, although linguistic structure spans multiple temporal scales. Successor representations (SRs) make this horizon explicit by encoding discounted distributions over future states. Here, we ask whether such predictive representations can recover not only word classes, but also finer functional and construction-like structure from natural language. A residual network trained on WikiText-103 predicts SR distributions at three horizons without part-of-speech supervision. At the shortest horizon, unsupervised clustering robustly recovers nouns, verbs, and adjectives, while directed inter-cluster transitions reproduce familiar syntactic asymmetries. At finer resolutions and across 13 part-of-speech categories, the same geometry reveals semantic-functional groupings that cross category boundaries and directed relations tracing candidate date, measurement, and title-name constructions. Part-of-speech agreement declines as the predictive horizon lengthens. These results suggest that word classes are coarse regions within a richer predictive geometry in which categorical and construction-like linguistic structure emerge from future-word distributions.

Sources

Related papers