The relational processing limits of classic and contemporary neural network models of language processing

arXiv:1905.05708 · cs.CL, cs.AI · Submitted 2019-05-12 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "The relational processing limits of classic and contemporary neural network models of language processing".

Jane: The paper was written by Guillermo Puebla, Andrea E. Martin and Leonidas A. A. Doumas from University of Edinburgh and Universidad de Tarapacá and Max Planck Institute for Psycholinguistics.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Alright, welcome back to the show, everyone. We've got a fascinating paper today that's been making the rounds, and it's called "The relational processing limits of classic and contemporary neural network models of language processing." Jane, I have to say, just the title alone is a mouthful, but it's really about something we all care about — can machines actually reason about relationships between things?

Jane: Absolutely, Tom. And I love that this paper is taking on both the old-school models and the new fancy deep learning ones. The authors — Guillermo Puebla, Andrea Martin, and Leonidas Doumas — they're basically saying, "Hey, we've got these two very different approaches to language, and we want to see if either of them can actually handle relational reasoning."

Tom: Right, and for our listeners who might not be deep in the weeds of cognitive science, let's break that down. Relational reasoning is when you understand that "the fox is chasing the hen" — you know the fox is the chaser and the hen is the chased. It's not just about the words; it's about the roles they play.

Jane: Exactly. And the paper's core claim is that neither the classic Story Gestalt model from the 90s, nor a modern Sequence-to-Sequence model with attention — the kind that powers things like Google Translate — can truly handle this. They can memorize patterns, but they can't dynamically bind a role to a filler when the situation changes.

Tom: So it's like, you can teach a model that "Anne tipped the waiter," but if you suddenly say "Lois tipped the waiter," and Lois was never an agent in a restaurant story during training, the model just freezes up?

Jane: Precisely. And that's what they call the "concept violation" test. The model was trained with certain characters restricted from certain scripts, and when those characters showed up anyway, the model couldn't adapt. It just defaulted to the most common characters it had seen.

Tom: That's wild. And the modern model, the Seq2seq with attention, it actually did better on that specific test, which is interesting. But hold on, we're getting ahead of ourselves. Let's talk about why this matters for the field.

Jane: Well, it matters because there's been this big debate for decades — can neural networks handle structured, relational knowledge, or do we need symbolic systems? And this paper is a pretty direct challenge to the idea that these networks, even the deep ones, are doing anything like human reasoning.

Tom: Yeah, and the authors are pretty clear that the failures come down to one thing: the binding problem. That's the ability to say, "this specific object fills this specific role right now, and I can change that assignment on the fly." These models just don't have that.

Jane: And that's a big deal because language is full of that. Every sentence is a new binding of roles to fillers. If a model can't do that dynamically, it's really just doing statistical pattern matching, not understanding.

Tom: So the title is almost a warning — "relational processing limits." It's saying, here's the ceiling, and it's pretty low. But the question is, what do we do about it? That's what we're going to dig into next.

Jane: Right, because the paper doesn't just stop at showing the problem. It points toward what a solution might look like. Stay with us.

Summary: Tom: Welcome back. We're digging into "The relational processing limits of classic and contemporary neural network models of language processing." Last segment we set the stage — these models can't bind roles to fillers dynamically. Now let's get into the actual experiments, because that's where it gets really juicy.

Jane: So they ran three critical tests. The first was the concept violation, which we mentioned — restricted characters show up in forbidden roles. The second was a correlation violation, and this one is really sneaky. In the training data, there was a perfect rule: if a restaurant is expensive, it's always far away. But in the test, they flipped it — expensive restaurant, but it's near.

Tom: And the models just failed. They kept answering "far" even though the story clearly said "near." It's like if I told you, "I bought a brand new sports car, and it's very slow," and you kept insisting it's fast because that's what you've always seen.

Jane: Exactly. And that's the correlation violation. The models overfit to that perfect statistical regularity. The Story Gestalt model got about fifty-two percent correct on that test, which is basically a coin flip. The Seq2seq with attention did even worse, only about fifteen percent correct.

Tom: Wow, that's brutal. And the third test was shuffling the order of the propositions in the story. The content was identical, just the sequence was randomized. The Story Gestalt handled that better — about eighty-seven percent correct — but the Seq2seq model really struggled, dropping to about twenty-one percent.

Jane: Right, and the authors have a hypothesis about why. The attention mechanism in the Seq2seq model is great at finding correlations between specific words and positions. But when you shuffle the order, that positional information becomes useless, and the model doesn't have a fallback.

Tom: So it's like the attention mechanism is both a blessing and a curse. It helped on the concept violation test, where it got about eighty-seven percent correct, but it hurt on the other two tests.

Jane: Exactly. And that's a really important finding — the same architectural choice that helps in one scenario actively hurts in another. It's not that one model is simply better; they each have different weaknesses.

Tom: And the baseline, just to be clear, both models were near perfect when the test stories matched the training distribution. The Story Gestalt got about ninety-five percent and the Seq2seq got over ninety-nine percent. So they're not bad at memorizing. They're bad at generalizing.

Jane: Right, and that's the key distinction. They can recall, but they can't reason. And the authors argue this is because they never actually learn the relational structure — the roles and fillers are all tangled up in the same activation patterns.

Tom: So when you ask them to apply a known role to a new filler, they can't separate the two. It's all one blob.

Jane: That's the picture. And this connects to a broader debate in AI about whether we need explicit symbolic representations or whether neural networks can learn everything from data. This paper is firmly on the side that says, at least for relational reasoning, you need something more.

Tom: So what's the something more? That's the million-dollar question, and I think we're about to get into it.

Jane: We are. The paper hints at solutions that involve dynamic binding — using time or other mechanisms to keep roles and fillers separate. Let's talk about that next.

Improvements: Tom: We're back, still on "The relational processing limits of classic and contemporary neural network models of language processing." So Jane, we've established the problem — these models can't bind roles to fillers. What do the authors suggest we do about it?

Jane: Well, they don't propose a full new architecture, but they point to existing models that already solve this problem. They mention systems like LISA and DORA, which use temporal synchrony — essentially, they fire role units and filler units at the same time to create a binding.

Tom: So it's like a handshake in time. The role "agent" and the filler "Lois" fire together, and that creates the binding. And you can break that handshake and form a new one when the story changes.

Jane: Exactly. That's dynamic binding. And the authors are saying, if you want a neural network to actually do relational reasoning, you need to build that capability in from the start. You can't just hope it emerges from training on lots of data.

Tom: And that's a pretty radical suggestion for the deep learning community, because the whole philosophy has been, "Give us enough data and compute, and the model will figure out the structure on its own."

Jane: Right, and this paper is saying, "No, that doesn't work for relational reasoning." The models can learn statistical regularities, but they can't learn the abstract structure of roles and fillers. And the authors are pretty confident about this because they tested it directly.

Tom: So what would a better model look like? They mention SHRUTI, which is another model that uses temporal binding. And there's also work on neural blackboard architectures.

Jane: Yeah, and there's a broader trend now in deep learning research — people are starting to add memory modules, attention mechanisms that can point to specific slots, and graph networks that explicitly represent relations between entities. Those are steps in the right direction.

Tom: But the paper's point is that just adding attention isn't enough. The Seq2seq model had attention, and it still failed on two out of three tests. You need something that explicitly separates the role from the filler.

Jane: Right. And this is where I think the paper is really valuable — it provides a clear benchmark. If you're building a new model and you claim it can do relational reasoning, you should test it on these kinds of violations. Can it handle a concept it's never seen in a role? Can it handle a correlation that's inverted? Can it handle shuffled order?

Tom: That's a practical takeaway for the AI community. These tests are simple to run, and they expose fundamental weaknesses. I'd love to see every new language model tested this way.

Jane: And the authors are also careful to note that this isn't just an academic exercise. They connect it to human cognition — humans can easily handle these violations. If I tell you a story about an expensive restaurant that's near, you don't get confused. You just accept it and answer accordingly.

Tom: So the gap between human and machine reasoning is not just about scale — it's about the kind of representations. Humans have this dynamic binding ability, and these models don't.

Jane: Exactly. And that's the core message. The paper isn't saying neural networks are useless. It's saying they're incomplete. They need a mechanism for binding, or they'll always hit this ceiling.

Tom: So the next step is to build that mechanism in. And maybe that's where the future of AI is headed. Let's wrap this up in our final segment.

Conclusion: Tom: Alright, we're wrapping up our discussion of "The relational processing limits of classic and contemporary neural network models of language processing." Jane, give us the final takeaway.

Jane: The bottom line is that both the classic Story Gestalt model and the modern Seq2seq with attention model can memorize statistical patterns from training data, but they cannot perform dynamic binding of roles to fillers. That means they fail when you give them new concepts in familiar roles, when you break a perfect correlation, or when you shuffle the order of events.

Tom: And the numbers really tell the story. The Seq2seq model got only fifteen percent correct on the correlation violation — that's worse than random guessing. The Story Gestalt wasn't much better at fifty-two percent. These are not edge cases; these are fundamental failures.

Jane: Right. And the paper's contribution is twofold. First, it provides a clear experimental framework to test relational reasoning in any neural network. Second, it points to a solution — dynamic binding mechanisms, like temporal synchrony, that have been around in cognitive science for decades.

Tom: So the message to the deep learning community is, "You can't just scale your way out of this problem." You need to build in the right inductive biases.

Jane: And that's a hopeful message, actually. It means we know what the problem is, and we have ideas for how to fix it. Models like LISA and DORA have already shown it's possible in principle.

Tom: So as we say goodbye to this paper, I think the big question for our listeners is — what happens when someone actually builds a deep learning model with dynamic binding built in? That could be the breakthrough we've been waiting for.

Jane: Absolutely. And we'll be watching for that. Thanks for joining us, and we'll see you next time with a new paper to dig into.

Tom: Take care, everyone.

Guillermo Puebla, Andrea E. Martin, Leonidas A. A. Doumas

University of Edinburgh · Universidad de Tarapacá · Max Planck Institute for Psycholinguistics

cs.CL, cs.AI

Submitted: 2019-05-12

Updated: 2026-08-18

Code: https://github.com/GuillermoPuebla/RelationReasonNN

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 63/100

The gist: This paper investigates whether traditional neural network models can capture relational knowledge, specifically testing the ability of two models—the classic Story Gestalt (SG) model and a

Key concepts

Relational Reasoning
This is the ability understanding the roles objects play in a scenario. It means recognizing that one element is the chaser and another is chased, going beyond simple word processing to understand the actual relationships between entities.
Dynamic Binding
This refers to the capacity to link a specific role (like 'agent') to a specific filler (like 'Lois') instantly. It allows the system to change that assignment on the fly as new information is presented in a story.
Concept/Correlation Violation
These are tests designed expose model limitations. Concept violation occurs when restricted characters appear in forbidden roles, while correlation violation happens when a known statistical rule is inverted in the test data, causing the models to fail to adapt.
Seq2seq with Attention
A modern deep learning architecture used for language processing. It uses an attention mechanism to find correlations between specific words and their positions, though this same feature can sometimes lead to poor generalization when tested.

Terminology

Summary

This paper investigates whether traditional neural network models can capture relational knowledge, specifically testing the ability of two models—the classic Story Gestalt (SG) model and a contemporary Sequence-to-Sequence with Attention (Seq2seq+Attention) model—to perform dynamic binding of roles and fillers in a text comprehension task. The authors argue that relational reasoning requires predicate representations that maintain role-filler independence and allow dynamic binding, a capability they claim traditional PDP models lack.

The study uses a task based on St. John's original materials, where stories are generated from scripts (e.g., Restaurant, Bar, Park, Airport, Beach) with thematic roles such as agent-1, agent-2, topic, patient-theme, recipient-destination, location, manner, and attribute. Each story is a sequence of propositions, and models are trained to answer questions about the topic concepts by outputting the full proposition. The authors trained two versions of each model: one in a concept restricted condition (where certain concepts never appear in specific scripts) and one in a concept unrestricted condition (where all concepts appear in all scripts).

The authors designed three critical tests, all maintaining the relational structure of the stories while varying statistical properties relative to training data:

  1. Concept violation: Models trained in the concept restricted condition were tested with stories where restricted concepts filled agent-1, agent-2, or patient-theme roles. A correct (role-based) answer required using the restricted concepts in those roles.

  2. Correlation violation: Models trained in the concept unrestricted condition were tested with stories where a perfect statistical regularity (e.g., expensive restaurants are always far away) was inverted. A correct answer required using the input concept despite the violation.

  3. Shuffled propositions: Models trained in the concept unrestricted condition were tested with stories where the order of propositions was randomized. A correct answer required using the concepts from the proposition corresponding to each question, ignoring ordering.

The results, shown in Figure 4, are as follows:

  • Baseline test: Both models performed well. The SG model achieved 0.954 proportion correct, and the Seq2seq+Attention model achieved 0.996, despite being trained on half the number of stories.

  • Concept violation test: The SG model performed poorly (0.216), almost invariably filling restricted roles with the most common training concepts, replicating St. John's original findings. The Seq2seq+Attention model performed significantly better (0.872), slightly above its baseline, which the authors attribute to the attention mechanism allowing it to apply word representations to unseen sequences.

  • Correlation violation test: Both models performed poorly. The SG model achieved 0.523, and the Seq2seq+Attention model achieved 0.155. The authors note that the attention mechanism likely makes the Seq2seq+Attention model more prone to overfitting perfect correlations.

  • Shuffled propositions test: The SG model performed better (0.654) than the Seq2seq+Attention model (0.079). The authors hypothesize the attention mechanism is again the reason for the Seq2seq+Attention model's poor performance, though they could not test this directly because removing attention made the model fail the baseline test.

The authors conclude that both models fail to perform dynamic binding of independent roles and fillers. They state: Both models failed in our tasks demonstrating that they do not perform dynamic binding of independent roles and fillers. They argue that a model that dynamically binds roles to fillers would easily pass these tests by filling untrained concepts into trained roles. The failures are attributed to the models' reliance on statistical regularities of the training data rather than relational information.

The authors further argue that truly compositional behavior requires independent representations of objects and roles that can be bound dynamically, and that traditional PDP models (including deep learning models) do not implement this. They note that techniques like word embeddings (e.g., Word2vec) or spanning the input space are not solutions to the deeper problem of generalizing to new concepts or combinations based on abstract relations. They cite evidence from adversarial attacks on deep learning reading comprehension models and from Lake & Baroni (2018) showing sequence-to-sequence models fail at compositional generalization.

Finally, the authors acknowledge that neural network models could in principle integrate operations for symbolic dynamic binding, citing models like SHRUTI, LISA, and DORA that use time as a binding signal. They note a resurgence of interest in the binding problem in neural networks and computational neuroscience, and suggest that future research should address whether non-traditional architectures (e.g., with content-addressable memory) can achieve relational reasoning. They conclude: for a model to successfully account for all aspects of relational processing, it will need to implement a solution to the binding problem.

Improvements for AI systems

Based on the paper's findings, here are the specific improvements I can make to AI systems and what the improved system can do:

Improvement: Add a dedicated binding module that maintains independent representations of roles (e.g., agent, patient, location) and fillers (e.g., specific characters, objects). This module uses dynamic binding—allowing roles to be assigned and reassigned to fillers during processing, independent of statistical correlations in training data.

What the improved system can do:

  • Correctly answer questions about stories where a character appears in a role they never occupied during training (e.g., Lois tipped waiter big when Lois was never an agent in restaurant stories).

  • Process stories where a known statistical correlation is violated (e.g., an expensive restaurant is near, not far) by using the input proposition's actual filler rather than the learned correlation.

Improvement: Introduce a training regularizer that penalizes the model for over-relying on perfect correlations between roles and fillers. This forces the model to maintain a baseline ability to use the immediate input proposition's role-filler bindings, even when they contradict learned statistical patterns.

Improvement: Replace the standard attention mechanism (which weights encoder outputs based on learned alignment) with a structured memory buffer that stores each proposition as a discrete, indexed unit with its role-filler bindings explicitly encoded. The decoder retrieves propositions by matching the question's topic to the stored proposition's topic, not by statistical similarity.

Improvement: After the model generates an output proposition, add a verification layer that checks whether each role in the output is filled by a concept that actually appeared in the input story for that role. If a mismatch is detected (e.g., the model outputs Barbara as agent when Lois was the input agent), the layer forces a correction by re-binding the output to the input's filler.

Improvement: During training, randomly swap role-filler assignments across stories (e.g., take the agent from story A and the patient from story B) and require the model to answer correctly using the new bindings. This forces the model to learn that roles and fillers are independent and can be recombined.

Improvement: Use two parallel processing streams: (1) a statistical stream (standard LSTM/attention) that captures distributional patterns, and (2) a relational stream (using explicit role-filler binding) that processes each proposition independently. The final output is a weighted combination, where the relational stream's output is favored when the statistical stream's confidence is low (e.g., when input violates learned correlations).


The improved system can:

  • Answer questions about stories with completely novel role-filler combinations (concept violation test).

  • Correctly process stories that violate learned statistical correlations (correlation violation test).

  • Answer questions correctly regardless of proposition order (shuffled propositions test).

  • Maintain high performance on standard, statistically regular stories (baseline test).

  • Avoid the failure modes identified in both the Story Gestalt and Seq2seq+Attention models by explicitly supporting dynamic role-filler binding, rather than relying solely on statistical regularities.

Sources

Related papers