MERGE: Minimal Expression-Replacement GEneralization Test for Natural Language Inference
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "MERGE: Minimal Expression-Replacement GEneralization Test for Natural Language Inference".
Tom: The Minimal Expression-Replacement GEneralization (MERGE) test introduces an automated methodology to evaluate how well current state-of-the-art reasoning models generalize when faced with minimal,
Jane: First, who's behind it and why it matters.
Paper summary: Tom: Well folks, we're jumping into a paper today that's tackling a really important issue in AI right now. It’s called "MERGE: Minimal Expression-Replacement GEneralization Test for Natural Language Inference," and I think this is going to get our heads turning.
Jane: Absolutely, Tom, it sounds like they are trying to figure out how robust current reasoning models actually are when things change just a little bit in their input.
Lu: This paper proposes the Minimal Expression-Replacement GEneralization Test for Natural Language Inference as a way to evaluate the generalization capacity of state-of-the-art reasoning models against non-adversarial variants of existing evaluation datasets, which is really important because designing new high quality reasoning datasets manually is costly and often unreliable for automated generation.
Meng: So, the core idea here seems to be that we can automatically create high quality variants from original instances using something called Minimal Expression Replacement or MERE generation, which uses Masked Language Models and safeguarding filters.
Lalam: I see a process where MLMs suggest contextually probable replacements for open class words that are shared between the premise and hypothesis sentences, ensuring the variants maintain the underlying reasoning of those original seeds.
Tom: That sounds like a clever way to generate test cases without deliberately breaking common heuristics or relying on adversarial strategies, which is exactly what they were looking for.
Jane: And the paper claims that this approach generates variants that keep the seed length and word overlap size intact, which those models can use as helpful heuristics when evaluating the changes.
Lu: They are applying this MERGE test specifically to Natural Language Inference, a popular task of reasoning, by automatically generating new NLI datasets from two widely used existing ones.
Meng: The evaluation involves testing multiple strong NLI models on these generated variants with respect to correctness and consistency using pattern-based accuracy metrics.
Lalam: They use two primary metrics for this assessment, seed accuracy to measure standard performance on individual seed problems, and a seed-variant accuracy metric that only counts a seed correctly classified if its variants are also correct above a minimum correctness threshold.
Tom: So, the main point is that even strong models can struggle to consistently and correctly classify variants minimally different in form and reasoning from the original instances because of this test.
Jane: It suggests that while models look good on individual problems, their ability to handle subtle, non-adversarial changes in language structure or wording is where the real weaknesses are exposed.
Lu: The MERE generation procedure involves using MLMs to suggest replacements for shared words and then filtering those suggestions based on quality criteria like maintaining the same class as the original word and having a higher probability.
Meng: They tested this method across various models, including fine-tuned NLI models and large language models like Gemma-two-9b and Llama-three point one-8B.
Paper summary: Lalam: Interestingly, the research also looked into how different aspects of variant generation affect model performance, specifically comparing scores grouped by the class replaced, like nouns versus verbs or adjectives.
Tom: That's interesting because it shows that some types of grammatical changes are harder for the AI to handle than others, which gives us some insight into where these models might be weakest.
Jane: And they also looked at whether the source MLM used for generating replacements had any influence on the model's final scores across different variant origin categories.
Lu: The findings indicated that while scores were consistently high on variants from equivalence types, other dataset origins didn't seem to influence predictions overall.
Meng: From an engineering standpoint, the results showed that models generally correct zero point seven percent SNLI and zero point zero five percent MNLI-m problems on average across all variants per seed/original problem.
Lalam: The paper also flagged a limitation in their setup, stating that the researchers found that verbs followed by nouns and adjectives are generally harder to generalize to than other structures.
Tom: So, even with these sophisticated methods for generating variants, the models still show a significant drop in performance once you push past a certain threshold on those variant problems.
Jane: And they found that while seed accuracy is maintained when the threshold is at most around fifty percent, performance starts to drop significantly beyond that level.
Lu: This suggests that the models are not robust across all minimal variations, even when those variations don't actively try to break them.
Meng: It points toward a need for better generalization testing methods because current fine-tuned NLI models still struggle with these subtle linguistic shifts.
Lalam: Thinking about the broader impact, if we can understand *why* they fail on these minimal changes, it could help us build more resilient AI systems that don't break when faced with slightly different natural language inputs in real-world applications.
Tom: That’s a big picture idea; understanding the failure modes of these models is crucial for improving how we deploy and trust them in complex tasks.
Jane: So, to summarize what we've covered about this MERGE test paper, it proposes an automatic method to generate non-adversarial variants of NLI datasets to evaluate model generalization against small linguistic changes.
Lu: And the conclusion from the study is that even strong models exhibit a drop in performance when those minimal variations exceed a certain threshold, highlighting weaknesses in their ability to handle subtle structural differences.
Meng: For practical purposes, this means we need to look beyond just getting high scores on clean data and test how well the model holds up when the language shifts slightly from what it was trained on.
Lalam: I think the implication is that future work should focus on developing models that are inherently more sensitive to those subtle structural variations, rather than just trying to patch them with more data.
Conclusion: Tom: So, we've been digging into the MERGE test paper that's checking how well current AI models generalize when faced with tiny changes in language, and now we’re wrapping up our discussion on its main points. Jane, can you give us a quick rundown of what this whole thing is about in plain English?
Jane: Sure thing, Tom. Basically, this research introduces a way to automatically create slightly altered versions of existing Natural Language Inference datasets using minimal changes called Expression Replacement. It then tests how robust powerful AI models are when they see these very subtle variations compared to the original problems.
Lu: What I find fascinating is the methodology behind generating those variants; it uses Masked Language Models to suggest replacements for shared words while keeping the core reasoning structure intact, which opens up a lot of creative avenues for testing model boundaries.
Meng: From an engineering standpoint, it’s cool that they used different MLM generators to see if those changes actually matter, and I wonder how much computational overhead that generation process adds to a standard evaluation pipeline.
Lalam: Personally, I think the most impactful vision here is realizing we can build AI systems that aren't brittle when the input phrasing shifts slightly from what was initially seen during training, which could really improve how we integrate AI into complex cultural contexts.
Tom: That’s a big picture thought, Lalam. And looking at the authors, they did a solid job setting up this rigorous test to expose where our current models are weakest without resorting to deliberately tricky adversarial examples.
Jane: Exactly! The authors show that even top-tier models struggle when those minimal linguistic tweaks push them past a certain point of consistency on the variant problems.
Lu: So, the core message is that generalization isn't just about handling massive changes; it’s about maintaining accuracy across these tiny, non-adversarial shifts in expression.
Meng: I see what they mean regarding the threshold where performance starts to dip; it tells us we need to build more resilient architectures that don't rely too heavily on exact word matching for complex reasoning.
Lalam: And it suggests that improving AI culture means making our models more adaptable and less sensitive to minute stylistic differences in human language.
Tom: It really makes you think about the robustness of everything we build, not just the big features, but how well the underlying reasoning holds up under subtle linguistic pressure.
Jane: Right, and this paper isn't just a technical exercise; it points toward a new way to rigorously measure how 'human-like' or flexible our AI understanding actually is.
Lu: It sets up a fantastic foundation for exploring more nuanced forms of model evaluation that look deeper than simple accuracy scores on clean data.
Meng: So, while the findings are interesting, I’m curious about the practical steps needed to implement this kind of variant generation pipeline reliably in a production environment where we have millions of examples.
Utrecht Institute of Linguistics · Utrecht University
cs.CL
Submitted: 2025-10-28
Updated: 2026-10-01
Importance score: 83/100
The gist: The Minimal Expression-Replacement GEneralization (MERGE) test introduces an automated methodology to evaluate how well current state-of-the-art reasoning models generalize when faced with minimal,
Key concepts
- Minimal Expression Replacement (MERE)
- This automated technique generates new NLI datasets by replacing shared words in original sentences with contextually probable alternatives suggested by Masked Language Models. This process creates variants that are minimally different from the originals, testing the model's robustness to small changes.
- Seed Accuracy (S) and Seed-Variant Accuracy (SV)
- Seed accuracy measures a model's performance on the original, unmodified problems. Seed-variant accuracy measures performance when variants are also correctly classified above a specific threshold. The matching correctness threshold (MC) finds the point where SV scores align with S scores.
- Contextual Probability Filtering
- Replacements for words are selected from Masked Language Models based on several criteria: the replacement must have a higher probability than the original word in a masked context, it must belong to the same part of speech, and it must be different from existing words. This ensures generated variants are linguistically plausible.
- Generalization Failure Rate
- The study found that models drop significantly in performance when testing on variants beyond a 50% correctness threshold. This indicates that even powerful models cannot reliably predict the correct classification for sentences that are only slightly modified from the training data.
Terminology
Summary
The Minimal Expression-Replacement GEneralization (MERGE) test introduces an automated methodology to evaluate how well current state-of-the-art reasoning models generalize when faced with minimal, non-adversarial variants of existing Natural Language Inference (NLI) datasets. This research is significant because it demonstrates that even strong Large Language Models and fine-tuned NLI models struggle to consistently and correctly classify variants minimally different in form and reasoning from the original instances.
The gist
Models fail to generalize to more than around half of the variants of two datasets, with surpassing the ≈50% CT yielding lower scores across models than their original seed-obtained scores.
How it works
The core of the MERGE methodology involves automatically generating high-quality variants from original instances using Minimal Expression Replacement (MERE) generation, which leverages Masked Language Models (MLMs) and safeguarding filters. This process is applied to Natural Language Inference (NLI), a popular reasoning task, by generating new NLI datasets from two widely used existing ones. The MERGE test then evaluates multiple strong NLI models on these generated variants.
The MERE generation procedure involves the following steps:
-
MLMs are used to suggest
contextually probable replacements for the open-class words shared between p and h.
-
Replacements are collected from a set of MLMs, where a replacement 'r' for an open-class word 'o' in sentence S is suggested if it has a higher probability than 'o' in the masked context S[si/MASK] under MLM Mj, or if it doesn’t occur in S.
-
Replacements are filtered based on quality criteria: they must be
same class as o,
havehigher probability than o,
and bedifferent from words already in p and h, including o.
-
The set of replacements for a word 'o' is defined as RM(S, oc>) = [Mj∈M Rj (S, oc>) where replacements are validated by the same MLM at each occurrence of o in S.
-
Variants are formed by replacing the original shared words oi ∈ p ∩ h with corresponding replacements rij ∈ RM(⟨p, h⟩, oc i). The total number of generated variants is k × d, where k is the number of shared words and d is the degree of inflation (inflation degree).
Evaluation Metrics
The MERGE test employs two primary metrics to assess model performance:
-
Seed accuracy (S): This measures a model's standard accuracy on individual seed problems.
-
Seed-Variant accuracy (SV): This metric, adapted from Abzianidze et al. (2023), counts the seed correctly classified only if its variants are also correctly classified to at least a minimum correctness threshold (CT). The matching correctness threshold (MC) is defined as the CT at which models’ SV scores are closest to their S scores.
Experimental Setup and Findings
The researchers generated 200 replacements for each seed problem using BERT, RoBERTa, ALBERT, Electra, and BART. Quality filtering was applied strictly based on class preservation (POS), probability (PROB), and validation by at least one MLM. The final dataset consisted of randomly sampled variants from the eligible seed problems. Models evaluated included various fine-tuned NLI models (e.g., BERT, RoBERTa, GPT-2) and two LLMs (Gemma-2-9b and Llama-3.1-8B).
The results indicate that while models maintain their S scores when the SV CT threshold is set to at most ≈50%, performance drops significantly beyond this level. Furthermore, analysis of word classes revealed that verbs, followed by nouns and adjectives, are generally harder to generalize to.
Finally, the study found that a source MLM does not affect an NLI model’s performance. Models generally correct 0.7% SNLI and 0.05% MNLI-m problems on average across all variants per seed/original problem.
Class and Source MLM Analysis
The researchers investigated how different aspects of variant generation affect model performance:
-
They compared scores of variants grouped by the class replaced (Noun, Adjective, Verb). They found that
adjectives are easier than verbs on higher thresholds,
though they had similar scores in lower thresholds. -
They examined the influence of the source MLM on model performance across different variant origin categories (Equiv, Size, Multi, One). The results showed that while scores are consistently high on Equiv variants (green lines),
the overall scores of other datasets suggest origin does not influence model predictions.
-
The study confirmed that replacements from different MLMs do not influence models’ scores.
Limitations
Methodological limitations include:
Improvements for AI systems
As a fastidious and diligent AI researcher, I have analyzed the proposed Minimal Expression-Replacement Generalization (MERGE) test and its methodology for generating robust generalization benchmarks.
Here are the specific improvements that can be made to current Natural Language Inference (NLI) models, based on the insights from this paper:
The core improvement lies in shifting model training and evaluation from simple in-distribution
performance to minimal variant robustness.
Current models often rely on memorizing heuristics or surface patterns that fail when minor lexical changes are introduced.
Here are specific improvements categorized by the research findings:
The improved AI system, leveraging these insights, will be capable of:
-
Maintain high performance even when presented with semantically similar but lexically minimal paraphrases or substitutions (e.g., replacing a noun with a synonym of the same part-of-speech).
-
Demonstrate superior generalization across diverse linguistic structures (nouns, verbs, adjectives) by learning deeper, reasoning-based relationships rather than relying on superficial word patterns.
-
Be less susceptible to biases introduced during data augmentation or fine-tuning processes when encountering novel but plausible word substitutions suggested by different language models.
-
Achieve a more stable and predictable performance profile across different model architectures (LLMs vs. traditional Transformers) when tested under the strict constraints of minimal expression replacement, indicating a more robust core reasoning capability rather than architecture-specific shortcuts.
Sources
- Abductive Commonsense Reasoning
- A large annotated corpus for learning natural language inference
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators
- Language models show human-like content effects on reasoning tasks
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- DeBERTa: Decoding-enhanced BERT with Disentangled Attention
- Learning the Difference that Makes a Difference with Counterfactually-Augmented Data
- ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- Lost in Inference: Rediscovering the Role of Natural Language Inference for Large Language Models
- Adversarial NLI: A New Benchmark for Natural Language Understanding
- From Superficial Patterns to Semantic Understanding: Fine-Tuning Language Models on Contrast Sets
- NLI under the Microscope: What Atomic Hypothesis Decomposition Reveals
- XLNet: Generalized Autoregressive Pretraining for Language Understanding
- OPT: Open Pre-trained Transformer Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering