On the Role of Directionality in Structural Generalization

arXiv:2607.02307 · cs.CL, cs.LG · Submitted 2026-08-16 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "On the Role of Directionality in Structural Generalization".

Jane: The paper was written by Zichao Wei from Saarland University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone. Today we're digging into a paper that's been making the rounds on arXiv, and the title alone got me hooked: "On the Role of Directionality in Structural Generalization." Jane, what's your first read on that?

Jane: Tom, I love it because it's so specific. It's not saying "we made AI better." It's saying there's this one particular property — directionality — and we think it's the reason some systems fail and others succeed. That's a real scientific question.

Tom: And for our listeners who aren't deep in the weeds — what does "directionality" even mean here? I've got my own guess but I want to hear you say it.

Jane: Sure. Think of a sentence as a line of words. When a system parses "the dog chased the cat," it needs to know the dog is the chaser and the cat is the chased. That's a left-right distinction. Directionality means the grammar actually encodes which side an argument comes from. Some systems just don't bother with that — they treat left and right as the same.

Tom: Right, and that's the core claim. The authors rebuilt a semantic parser using a grammar system called CCG, which has slashes that literally point left or right to say where arguments go. And they show that this one design choice gives them a huge boost on certain test categories.

Jane: And not just a small boost. On the position-shift categories, their system beats the previous best by almost thirty percentage points. That's not noise. That's a structural advantage.

Tom: So the title is almost a thesis statement — directionality is the variable that matters. The authors are from Saarland University, and they're building on years of work in neuro-symbolic parsing. This feels like a paper that could change how people design these systems.

Jane: It could, because it's not saying "bigger model, better results." It's saying the representation you choose determines what the model even has to learn. That's a much more interesting story.

Tom: And that's the hook for the rest of the episode. We're going to dig into how they actually built this thing, what the results look like category by category, and why this might matter beyond just this one benchmark. Stick around.

Summary: Tom: So Jane, we've got the title unpacked. Now let's talk about what the paper actually does. It's a neuro-symbolic system — that means a neural network picks a type for each word, and then a symbolic parser does the heavy lifting. But the twist is the type system itself.

Jane: Right, and this is where I want to bring in Lu, because this is exactly the kind of design decision that gets overlooked. Lu, why does the type system matter so much here?

Lu: Because it's the interface, Tom. The neural network's only job is to output a type sequence. The symbolic parser's only input is that type sequence. If the type system doesn't encode direction, then the neural network has to figure out left versus right all by itself. That's a much harder learning problem.

Jane: And the authors show that with the same underlying BERT encoder, their directional system hits about seventy-six percent exact match, while the previous non-directional system gets about seventy-one percent. That's a real gap.

Tom: But here's the thing that blew my mind — when they swap in a bigger, better encoder, their system jumps to over ninety percent. And the gains from the better encoder land mostly on the recursive-depth categories, not the position-shift ones. So directionality and encoder strength are doing different jobs.

Lu: Exactly. Directionality fixes the left-right confusion. The bigger encoder fixes the deep nesting problem. They're complementary, not competing. That's a really clean experimental result.

Meng: Can I jump in here? From a practical standpoint, the fact that they only use thirty thousand trainable parameters — that's tiny. And training takes eight minutes on a single GPU. That's not a research toy. That's something you could actually deploy.

Jane: And that's the beautiful part, Meng. The symbolic layer is doing the reasoning, so the neural network doesn't have to memorize everything. It just has to be good at tagging. That's a much more tractable problem.

Tom: So the summary is: directionality in the grammar makes the neural network's job easier, and that pays off precisely where left-right matters. And then you can scale up the encoder to handle the other hard stuff. It's a two-part story and both parts are clean.

Lu: And I'd add that the paper is honest about its limitations. They can't just "turn off" directionality in CCG to run an ablation — it's baked into the formalism. So they use category-level patterns instead, and the pattern is exactly what you'd predict if directionality were the cause.

Tom: That's the kind of rigor I appreciate. Alright, so we know what they did and why it works. Next up — what does this mean for the field going forward? That's where it gets really interesting.

Improvements: Tom: Alright, we've covered the setup and the results. Now let's talk about what this paper actually improves — not just the numbers, but the way we think about building these systems. Jane, you've been nodding along, what's the big takeaway for you?

Jane: For me, it's that the representation is the bottleneck, not the model size. The authors show that AM-Parser — the previous state of the art — has two categories where it scores literally zero percent. No matter how good the encoder gets, it can never fix those categories because the grammar itself can't express the needed distinction.

Meng: That's the architectural ceiling, right? And their system doesn't have that ceiling. Every category is non-zero, so upgrading the encoder actually helps everywhere. That's a huge practical difference — you're not just polishing a system that's stuck.

Tom: And the improvements aren't just about fixing zeros. Look at the modifier position categories. The old system had a massive asymmetry — it could handle modifiers on the object side but failed on the subject side. The new system nearly eliminates that gap for PP modifiers.

Lu: And that's because of how the grammar works. In CCG, a modifier has the same type whether it's on the left or the right — the slash direction handles the distinction automatically. So the neural network doesn't have to learn two different patterns. It learns one type, and the grammar does the rest.

Jane: That's the engineering insight, but there's a linguistic one too. The authors argue that directionality isn't just a hack — it reflects something real about how humans process language. We produce speech in a linear sequence, and recovering the structure means figuring out which side things came from.

Meng: So you're saying this isn't just a better parser, it's a more human-like parser?

Lu: That's the claim, and I think it's a fair one. The benchmark SLOG was designed to test structural generalization — things like moving a modifier from one side of a verb to the other. Those are exactly the cases where directionality matters most. The fact that the gains land precisely there is strong evidence.

Tom: And the improvements compound. With the best encoder, nine out of seventeen categories average above ninety-seven percent. That's not incremental — that's a step change. This paper is showing a path forward that doesn't require inventing a new architecture every six months.

Jane: It's about choosing the right grammar and letting the neural network do what it's good at — pattern recognition — while the grammar does what it's good at — logical composition. That division of labor is the real improvement here.

Tom: So the improvements are threefold: no more zero-score categories, better handling of position shifts, and a clear path to scale with better encoders. That's a strong package. Let's wrap this up with what it all means.

Conclusion: Tom: Alright, we've spent the episode on "On the Role of Directionality in Structural Generalization," and I think we've got a clear picture. Jane, how do you want to sum it up?

Jane: I'd say the core message is that the grammar you choose determines what the neural network has to learn. Directionality — encoding left versus right in the type system — takes a huge burden off the neural network, and the results show it. seventy-six percent with a standard encoder, ninety-one percent with a better one. That's a massive jump.

Tom: And it's not just the numbers. It's the fact that the gains land exactly where you'd predict — on the position-shift categories — and not on the recursive-depth ones. That's the signature of a real cause, not a coincidence.

Lu: I'd add that this validates the neuro-symbolic approach in a deeper way. It's not just "symbols help." It's "the right symbols help in the right places." That's a much more refined understanding.

Meng: And from my side, the practical impact is clear. Thirty thousand trainable parameters, eight-minute training, and it beats systems that are orders of magnitude more complex. That's a blueprint for efficient, reliable semantic parsing.

Jane: And the linguistic argument — that directionality reflects how humans actually process language — that gives me hope that these systems aren't just fitting benchmarks. They're learning something structurally real.

Tom: So we're saying goodbye to this paper, but not to the ideas in it. Directionality, representation design, neuro-symbolic division of labor — these are going to stick around. Thanks to everyone who joined us today.

Jane: And to our listeners — if you're building parsers, or just curious about why some AI systems generalize better than others, this paper is worth your time. Until next time, keep asking the hard questions.

Tom: Take care, everyone. We'll see you on the next episode.

Zichao Wei

Saarland University

cs.CL, cs.LG

Submitted: 2026-08-16

Updated: 2026-08-18

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 81/100

Key concepts

Directionality
In language processing, directionality means the grammar encodes which side an argument comes from (e.g., left-right distinction). It ensures the system knows which element is doing what, rather than treating both sides equally.
Structural Generalization
This refers to a system's ability to handle novel linguistic structures, such as moving a modifier from one side of a verb to another. The benchmark SLOG tests these structural changes.
Neuro-symbolic Parsing
This approach combines two methods: a neural network (which assigns types) and a symbolic parser (which does the heavy lifting). The type system acts as the crucial interface between the two components.
Type System
The type system dictates what information is available to the symbolic parser. If it lacks directionality, the neural network must learn left-right distinctions, making the learning problem much harder.

Terminology

Summary

Summary

This paper investigates the role of directionality in structural generalization for semantic parsing, specifically within the context of the SLOG benchmark. The authors argue that while the previous state-of-the-art system, AM-Parser, uses an AM algebra whose operations do not encode direction (a design choice inherited from AMR parsing where word order is irrelevant), several SLOG test categories explicitly require directional distinctions, such as modifier position shifts and argument extraction positions.

To test this, the authors redesign the symbolic backend around Combinatory Categorial Grammar (CCG) directed types. The system uses a frozen BERT encoder, a Gumbel-Softmax discretization layer, and a single linear decoder with only 30K learnable parameters. The symbolic component uses deterministic CKY parsing with forward and backward function application rules, followed by deterministic semantic edge extraction. The type system consists of 26 CCG types, including 20 base types for COGS, 4 extension types for SLOG's relative clauses and wh-questions, and 2 disambiguation types.

Under the controlled condition with BERT-base, the CCG system achieves 75.9±6.4% LF exact match, surpassing AM-Parser's 70.8±4.3%. When the encoder is upgraded to DeBERTa-v3-large, the system reaches 90.7±4.9%, exceeding AM-Parser by nearly 20 percentage points, with 9/17 categories averaging above 97%.

The core finding is a precise directional separation when analyzing results by SLOG's own category groupings. The CCG system outperforms AM-Parser on all 5 position-shift categories (§2.2 modifier position and §2.3 extraction position) by an average of +29.9pp, while AM-Parser outperforms on all 6 recursive-depth categories (§2.1) by an average of −31.9pp. The authors note that the absolute magnitudes of these two deltas are nearly equal, indicating that the effect of directionality is not a general improvement but a directional one.

Further evidence comes from within-group analysis of wh-questions (§2.4). The 6 wh-categories are divided a priori into direction-relevant subcategories (Q dobj ditransV, Q iobj ditransV, Q long mv, Q modified NPs) and direction-irrelevant subcategories (Q subj active, Q subj passive). The CCG system leads by large margins on direction-relevant subcategories (+30.0 to +66.9pp), while AM-Parser leads on direction-irrelevant ones. This within-group comparison controls for system-level confounds.

The paper provides two explanations for why directionality works. The engineering argument states that in a neuro-symbolic system, the type system is the sole interface between the neural predictor and the symbolic reasoner. Directionality determines how information is distributed between non-learnable combinatory rules and learnable type labels. In the AM algebra, PP modification of object vs. subject requires different type labels because the modify operation does not distinguish left from right; in CCG, the PP's type is NP in both cases, and directionality is automatically distinguished by slash direction at composition time. The authors cite Weißenhorn et al. (2022a) as indirect confirmation, noting that AM-Parser required an additional +dist feature to distinguish linear position, which positional information was omitted by the AM algebra and had to be externally supplied.

The linguistic argument contends that directionality reflects a structural property of language production and comprehension. Natural language is a one-dimensional linear signal, and the process of converting conceptual structure into word sequence is called linearization (Levelt, 1989). Hawkins (1994, 2004) provides arguments that cross-linguistic word-order regularities are governed by structural-recognition efficiency. The authors state that directionality (which side of the head an argument appears on) is the property of word order most directly relevant to structural generalization.

The paper also demonstrates encoder scalability. AM-Parser scores 0.0±0.0 on two categories (RC iobj extracted and Q long mv), constituting an architectural ceiling. All categories under the CCG system are empirically non-zero. Testing three encoder tiers (BERT-base 75.9%, ModernBERT-base 79.9%, DeBERTa-v3-large 90.7%) shows that encoder improvements are not truncated by the composition layer. The largest encoder gains appear in recursive-depth categories (+22.5pp), complementary to directionality's gains on position-shift categories. The authors conclude that directionality and encoder strength are complementary, not substitutive.

The paper acknowledges limitations. Pipeline fidelity is 99.92%, with 13 inconsistencies stemming from inherent ambiguities in the CCG type system (transitive and ditransitive verbs missing one argument share the type (S)/NP). The system uses a frozen BERT-base encoder, so pre-training leakage cannot be entirely ruled out. The argumentative framework relies on CCG's concept of directionality, and generalization to broader evaluations remains open. The variance analyses are based on different numbers of seeds (10 vs. 5). The encoder replacement experiments were conducted only under the CCG system because AM-Parser's evaluation pipeline is not fully open-sourced.

The paper also includes an appendix explaining why a directionality ablation is infeasible. Removing direction would require merging types that differ only by slash direction, making the CKY merge table symmetric would collapse agent and theme roles, and restoring role assignment would require introducing a learned role predictor. The authors state that a 'non-directional CCG' is not a variant of CCG but a different formal system requiring redesign of all three components.

The main conclusion is that on structural generalization involving positional distinctions, directional representations have a consistent and substantial advantage. The effect is directional (improvements concentrate on all 5 position-shift categories, absent for all 6 recursive-depth categories) and architectural (directional representations eliminate the 0% category ceiling, shifting the bottleneck to the neural layer).

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in AI systems:

Improvement: Replace directionless AM algebra with CCG-style directed types (S vs S/NP) in the symbolic backend.

Implementation:

  • Encode directionality directly in type definitions (left/right argument positions)

  • Use deterministic CKY parsing with forward/backward function application rules

  • Extract predicate-argument edges (agent/theme/recipient) from merge rules based on positional order

Resulting capability: The system can now distinguish modifier position shifts (PP/RC moving from object side to subject side) and argument extraction positions without additional neural components. This yields +29.9pp improvement on all 5 position-shift categories in SLOG.

The improved AI system can:

  • Parse sentences with novel structural combinations (relative clauses, wh-questions, center embedding) never seen in training

  • Correctly assign agent/theme/recipient roles based on positional directionality

  • Handle modifier attachment on both subject and object sides

  • Process recursive nesting up to depth 12 with 97.8% accuracy

  • Achieve state-of-the-art on SLOG (90.7%) while being trainable in minutes on a single GPU

  • Provide fully interpretable parse trees with explicit role assignment rules

Sources

Related papers