Structural Generalization on SLOG without Hand-Written Rules

arXiv:2604.26157 · cs.CL, cs.AI · Submitted 2026-08-16 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Structural Generalization on SLOG without Hand-Written Rules".

Jane: The paper was written by Zichao Wei from Saarland University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the arXiv radio hour, everyone. I'm Tom, and today we're looking at a paper that has a title that just gets straight to the point: "Structural Generalization on SLOG without HandWritten Rules." Jane, what do you make of that title?

Jane: I love it, Tom, because it's basically declaring war on a whole approach. For years, the best way to get a computer to understand new sentence structures was to give it a bunch of rules written by humans. This paper says, what if we just let the machine figure those rules out on its own?

Tom: Right, and that's a bold claim, because we're talking about a benchmark called SLOG, which is specifically designed to be brutal. It tests whether a system can take the grammar it learned from simple sentences and apply it to really twisted ones, like relative clauses and wh-questions.

Jane: And the paper isn't just claiming it works. They're showing that this fully learnable system, they call it an NCA, a neural cellular automaton, hits one hundred percent accuracy on eleven out of seventeen of those brutal categories. That's a massive jump from where end-to-end models usually land, which is around forty percent.

Tom: Hold on, let me bring in Lu from Tsinghua, because I know she's been following the neuro-symbolic debate for a while. Lu, does this title undersell what they actually did?

Lu: Tom, it actually undersells it in the best possible way. The previous state-of-the-art, AM-Parser, got seventy percent by using a neural network to guess types and then a hand-written algebra to combine them. This paper removes that algebra entirely. The composition rules are learned. And they still beat the pure neural baselines by a mile.

Jane: So it's not just "without hand-written rules" as a marketing slogan. It's literally the core contribution. They're saying the structure emerges from the training data, not from a linguist's notebook.

Tom: And that's the exciting part for me. It suggests that the inductive bias doesn't have to be a set of explicit rules. It can be a local, iterative process that just learns the right way to combine things.

Lu: Exactly. And they show that when it fails, it fails in a way that's incredibly clean. It's not random. It's always because a specific type of grammatical operation was never seen in training.

Jane: So the failures are actually a map of what's missing from the data, which is a beautiful diagnostic tool. Tom, I think we need to dig into how they actually built this thing, because the architecture sounds wild.

Tom: Absolutely, and that's our next segment. But the takeaway from the title alone is that this is a serious challenge to the idea that you need human intuition baked into the machine to get compositional reasoning. The machine is building its own grammar.

Summary: Jane: So we've established that "Structural Generalization on SLOG without HandWritten Rules" is a bold title. Now let's talk about what's actually under the hood, because the summary of the method is genuinely clever. Tom, you want to take the first swing?

Tom: Gladly. So the system takes a sentence, runs it through a frozen BERT encoder to get word meanings, and then it does something radical. It throws away the global context and forces each word into one of thirty-two discrete codes. It's like compressing the sentence into a string of symbols.

Jane: And that's the "discrete bottleneck." It's stripping away the fuzzy, continuous meaning and leaving only the categorical identity of each word. That sounds like it would lose information, but that's the point.

Tom: Exactly. Then they feed those codes into a neural cellular automaton. Think of it as a grid of cells, where each cell is a word, and they all update in parallel based on their neighbors. Over sixty steps, these cells pass information left and right, merging and combining.

Lu: And I love that they call it a "discrete bottleneck" because it's the same trick AM-Parser uses with its supertagger. But instead of predicting a type and then using a symbolic algebra to merge them, the NCA learns the merging rules itself. The local updates are the grammar.

Jane: Right, so the rules aren't written down. They're just the weights of the neural network that decide how two neighboring cells combine. And the training signal is just the final parse tree.

Tom: And here's the kicker, Jane. They don't backpropagate through all sixty steps. They only train the last step. It's called a detached rollout. So the system has to learn to set itself up correctly for that final step, which forces it to learn stable, local dynamics.

Meng: I'm Meng, by the way, and I have to ask the engineer question here. You're training a neural network to do sixty sequential iterations. That's usually a recipe for vanishing gradients or just instability. How does this not collapse?

Jane: That's the clever part, Meng. They use a curriculum on the number of steps. They start with one step, then two, then five, and gradually ramp up to sixty. So the system learns short-range rules first, and then learns to chain them together into long-range dependencies.

Meng: So it's like teaching a kid to add single digits before you ask them to carry over in multi-column addition. That makes sense. But what's the actual output? What are they predicting?

Tom: They're predicting CCG types. Combinatory Categorial Grammar. It's a linguistic formalism where every word has a type like "noun phrase" or "verb that takes a noun phrase on the right." The NCA's job is to figure out the correct type sequence for the whole sentence.

Jane: And that's the summary in a nutshell. A frozen encoder, a discrete bottleneck, a locally iterative reasoner, and a simple readout head. No hand-written rules anywhere. And it works surprisingly well.

Tom: And the fact that it works is what we need to dig into next, because the results aren't just good. They're weirdly clean. Let's talk about the numbers in the next segment.

Improvements: Tom: So we've covered the architecture of "Structural Generalization on SLOG without HandWritten Rules." Now let's talk about what the paper actually improves upon, and Jane, this is where the numbers get fascinating.

Jane: They are, Tom. The headline is that this system beats AM-Parser on three specific categories where the old approach just flat-out failed. For example, RC iobj extracted, which is about extracting an indirect object from a relative clause. AM-Parser gets zero percent on that. This system gets one hundred percent.

Meng: Wait, one hundred percent on something the previous state-of-the-art couldn't do at all? That's not an incremental improvement. That's a new capability.

Jane: Exactly. And it's not just that one category. It also beats AM-Parser on PP modif iobj and RC modif iobj, which are about modifying indirect objects with prepositional phrases and relative clauses.

Lu: And I think the more impressive improvement isn't just the peak performance, it's the stability. AM-Parser has a standard deviation of four point three across seeds. This system has zero point two. And on fifteen out of seventeen categories, the standard deviation is exactly zero.

Tom: Zero variance. That means if you run the training ten times with different random initializations, you get the exact same score on almost every category. That's a level of determinism that's almost unheard of in deep learning.

Meng: So it's not just more accurate. It's more reliable. For a production system, that might be even more important than the raw accuracy. If I deploy this, I want to know it's going to behave the same way every time.

Jane: And that stability is what allows them to do the deep analysis. Because the results are so clean, they can look at the failures and see that they all reduce to exactly two mechanisms. It's not noise. It's a structural boundary.

Lu: Right. They found that every single failure, all five thousand five hundred thirty-nine of them, is either a verb appearing with a reduced type in a wh-question, or a modifier appearing on the left side of the verb instead of the right. Those are the only two things that go wrong.

Tom: And that's the improvement over the benchmark itself. They're not just reporting a score. They're explaining why the score is what it is. They're showing that the SLOG categories mix together structurally different patterns.

Jane: The forty-one point four percent on Q modified NPs is a perfect example. That looks like a mediocre score, but when you split it by the grammatical role of the extracted word, you get one hundred percent on one half and zero percent on the other half. It's not partial success. It's two completely different tasks glued together.

Meng: So the paper is essentially saying, "Our system is perfect, and the benchmark is flawed." That's a bold move.

Lu: It's not saying the benchmark is flawed. It's saying the benchmark's labels are coarser than the underlying structure. The system is giving us a higher-resolution view of what's actually hard.

Tom: And that's the real improvement. It's not just a better score. It's a better understanding of the problem. And that understanding is what we should take with us into the conclusion.

Conclusion: Jane: Alright, Tom, let's wrap this up. We've spent the whole episode on "Structural Generalization on SLOG without HandWritten Rules," and I think we need to give it a proper send-off.

Tom: Absolutely. So the big picture is that this paper shows you can get compositional generalization without writing down a single grammatical rule. You just need a discrete bottleneck to force symbolic-like behavior, and a local iterative process to learn the composition.

Jane: And the results speak for themselves. sixty-seven point three percent overall accuracy, which is just a hair below AM-Parser's seventy point eight percent, but with a fraction of the variance. And it nails eleven out of seventeen categories perfectly, including three where the old approach got a big fat zero.

Lu: And the analysis is the real gift. They've shown that the boundary between success and failure is determined by whether a specific directional operation appeared in the training data. That's a testable hypothesis that goes beyond this specific benchmark.

Meng: From my side, the practical impact is that this architecture is small, stable, and fully differentiable. It's not a research curiosity. It's something you could actually deploy and retrain on new domains without hiring a linguist to write rules.

Tom: And that's the lasting impression. This paper isn't just about beating a benchmark. It's about changing the default assumption of how to build parsers. Instead of asking "what rules do we need to write?", you ask "what data do we need to show the system?"

Jane: It's a shift from engineering grammar to cultivating it. And that's a beautiful place to leave the paper. Tom, what's next on the docket?

Tom: Next up, we've got a paper on emergent communication in multi-agent reinforcement learning. Should be a wild ride.

Jane: Can't wait. Thanks for joining us, everyone, and we'll see you on the next episode of the arXiv radio hour. Goodbye, "Structural Generalization on SLOG without HandWritten Rules." You were a good one.

Zichao Wei

Saarland University

cs.CL, cs.AI

Submitted: 2026-08-16

Updated: 2026-08-18

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 86/100

Key concepts

SLOG
A difficult benchmark designed to test if an AI system can generalize grammar learned from simple sentences to complex or 'twisted' ones. It specifically tests structures like relative clauses and wh-questions.
Neural Cellular Automaton (NCA)
The core architecture discussed; it treats a sentence as a grid of cells (words). These cells update information in parallel based on their neighbors, learning composition rules through local, iterative processes.
Discrete Bottleneck
A method where the system forces each word into one of thirty-two discrete codes, stripping away continuous meaning. This categorical identity is then fed into the NCA for symbolic-like processing.
Compositional Generalization
The ability of an AI system to apply grammatical knowledge learned from simple examples to novel, complex sentence structures. The paper demonstrates this without human-written rules.

Terminology

Summary

Summary

This paper presents a fully learnable alternative to the neuro-symbolic AM-Parser for structural generalization on the SLOG benchmark, requiring no handwritten compositional rules. The proposed system combines a discrete bottleneck encoding via Gumbel-Softmax with the local iterative dynamics of a neural cellular automaton (NCA). The key difference from AM-Parser is that compositional rules are not handwritten AM algebra but local update rules automatically learned by the NCA through training. The entire system is end-to-end differentiable, with approximately 81K parameters (excluding a frozen BERT encoder).

The architecture consists of four components: (1) a frozen BERT encoder that encodes input sentences into contextualized word vectors, with function words stripped after encoding; (2) a discrete bottleneck using Gumbel-Softmax to discretize continuous BERT embeddings into one of K=32 codes, restoring locality; (3) an NCA reasoning layer where code embeddings undergo T steps of local iteration, with each step's update depending only on the current position and its left and right neighbors (Conv1d kernel=3 → GELU → Conv1d + Tanh bottleneck + LayerNorm), with all positions sharing parameters; and (4) a readout head that maps the NCA's continuous state to logits over 24 CCG types. Key design choices include no skip connections, no backpropagation through time (only the final step backpropagates), and a T-curriculum gradually increasing T from 1 to 60.

The system uses 24 CCG types: 20 base types covering all 21 COGS categories and 4 extension types (RC THAT, TV GAP, S GAP, WH) covering relative clauses and wh-questions in SLOG. CCG types are deterministically derived from gold logical forms, and CKY parsing produces parse trajectories that serve as training targets. The primary evaluation metric is type exact match, which the paper argues is equivalent to LF exact match: on the full 17,000 gen samples, the agreement between CKY edge extraction and gold LF edges is 99.9% (16,987/17,000).

On the SLOG benchmark, the system achieves an overall accuracy of 67.3±0.2% across 10 seeds, compared to AM-Parser's 70.8±4.3%. The NCA achieves 100% accuracy on 11 of 17 categories, including three where AM-Parser scores lower or zero: RC iobj extracted (100% vs. AM-Parser 0%), RC modif iobj (100% vs. 74.4%), and PP modif iobj (100% vs. 90.4%). Across 10 random seeds, 15 out of 17 categories have zero standard deviation. The only two categories with nonzero variance (Q subj active 99.2±2.4, Q subj passive 99.4±1.8) have 9 out of 10 seeds reaching 100%, with the single outlier seed misencoding a few samples due to Gumbel-Softmax initialization randomness.

Detailed analysis reveals that all 5,539 failure instances reduce to exactly two mechanisms. Mechanism A: verbs appear with reduced argument types in wh-extraction contexts. The training data contains no wh-questions, so verbs always retain their full argument types. In wh-questions, the extracted forward argument (object) no longer appears, and the verb must appear with a type missing that argument (e.g., (S)/NP → S). What is truly absent is the combination of this type with a wh-question context. Categories affected: Q dobj ditransV (all 1,000 samples), Q iobj ditransV (all 1,000 samples), Q long mv (all 1,000 samples), and 586 active object extraction samples in Q modified NPs. Mechanism B: modifiers appear on the subject side (left of verb). In training data, PP/RC modifiers always appear on the object side (right of verb). The modification operation itself is identical on both sides, but the object-side NP is consumed by the verb via forward application... while the subject-side NP combines with the verb via backward application. Categories affected: PP modif subj (all 1,000 samples) and 953 samples in RC modif subj that directly modify the main clause subject.

The paper demonstrates that SLOG's category labels, named after linguistic phenomena, mix structurally distinct CCG patterns. When reclassifying Q modified NPs by CCG features (role of the wh-gap, voice, presence of relative clause), ten sub-patterns emerge: 6 at 100% (414 samples total) and 4 at 0% (586 samples total). The dividing line is whether the verb hosting the wh-gap changes its CCG type. Subject extraction does not change the verb's type, while active object extraction does. Similarly, RC modif subj decomposes into eight sub-patterns: 4 at 100% (47 samples, where the RC modifies the subject of a CCOMP complement clause rather than the main clause subject) and 4 at 0% (953 samples, where the RC directly modifies the main clause subject). The paper states: Every sub-pattern is either 100% or 0%.

The paper argues that CCG directed type features offer higher resolution than phenomenon-level classification for predicting structural generalization success and failure. The two failure mechanisms are the same in essence: novel combinations of directed operations absent from training. The paper notes that standard CCG handles the subject/object asymmetry via type-raising (NP → S/(S)), but type-raising is itself an additional combinatory rule that the system cannot discover because training data contains no samples requiring it.

Comparing with AM-Parser, the paper notes that AM-Parser uses direction-agnostic AM algebra, and its nonzero results on subject-side modification (PP modif subj 57.6%, RC modif subj 55.8%) come from the explicit injection of its direction-agnostic algebra, not from generalization learned from training data. The NCA outperforms AM-Parser on object-side modification (100% vs. 90.4%/74.4%), while AM-Parser outperforms the NCA on subject-side modification. The paper states: The complementarity of the two systems corresponds precisely to their respective design choices.

The paper also discusses limitations of LF evaluation, noting that T5's plain exact match on Q dobj ditransV is 47.2% while reformatted exact match is 98.5%, with more than half of errors being differences in variable naming or conjunct ordering. The paper argues that CCG type sequences as an alternative evaluation target avoid this problem because types are discrete labels rather than strings, explicitly encode directionality, and directly measure structural-semantic understanding.

The paper concludes: "This paper presents a fully learnable structural generalization system that inherits the neuro-symbolic spirit of AM-Parser while replacing hand-written algebraic rules with learned local iteration. The system achieves 100% on 11 of 17 SLOG categories (including three that AM-Parser cannot handle), with an overall standard deviation of 0.2 across 10 seeds (compared to AM-Parser's 4.3). CCG analysis of the results shows that all 5,539 failure instances reduce to exactly two mechanisms: forward argument extraction changing verb types, and modified NPs moving from the right side to the left side of the verb. Both are, in essence, directed CCG operations absent from training. After reclassification by CCG features, every sub-pattern in SLOG yields deterministic results (0% or 100%); the surface-level intermediate values (41.4%, 4.7%) are mixtures of distinct CCG patterns."

The appendix reports COGS results: all 18 lexical generalization categories achieve 100%, with the only failure being obj pp to subj pp (0%), which requires moving a PP modifier from the object side to the subject side of the verb; this corresponds exactly to Mechanism B identified in the main text. The appendix also confirms that type exact match and LF exact match agree on 99.9% of samples, with only 13 discrepant samples in Q modified NPs arising from CCG type ambiguity between transitive verbs and ditransitive verbs missing one argument.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement and the resulting capabilities:

Current limitation: Transformer-based seq2seq models (T5, LLaMA) fail on structural generalization (40% accuracy on SLOG) because they attempt to generate output tokens directly from global attention patterns.

Improvement: Implement a two-stage architecture:

  • Stage 1: Frozen BERT encoder → Gumbel-Softmax discrete bottleneck (K=32 codes) that strips global context and restores locality

  • Stage 2: Neural Cellular Automaton (NCA) with Conv1d kernel=3, GELU, Tanh bottleneck, LayerNorm, iterated T steps (1→60 curriculum) that performs local compositional reasoning

Resulting capability: The system achieves 67.3% overall accuracy on SLOG, outperforming all end-to-end models (T5: 40.6%, LLaMA: 40.1%) and matching AM-Parser (70.8%) without hand-written rules.

Current limitation: LF evaluation conflates structural understanding with format compliance (variable naming, conjunct ordering), systematically underestimating model capabilities.

Current limitation: Backpropagating through all NCA iterations causes vanishing gradients and unstable training.

Current limitation: Phenomenon-level categories (e.g., Q modified NPs) hide structural distinctions, making it impossible to identify systematic failures.

Current limitation: Fully learnable systems fail on subject-side modification (0% on PP modif subj) because training contains no such examples.

Current limitation: SLOG's 17 categories mix structurally distinct patterns (e.g., Q modified NPs at 41.4% is actually 6 sub-patterns at 100% and 4 at 0%).


Summary of what the improved AI system can do: It can perform compositional semantic parsing on novel structural combinations without hand-written rules, achieving near-perfect accuracy on all structures whose directed operations appear in training, while precisely identifying and predicting which novel structures will fail due to absent operations. It does this with 81K parameters, 20-minute training, and deterministic behavior across seeds — making it suitable for production deployment where reliability and interpretability are critical.

Sources

Related papers