Memorization Diagnostics for Code LLMs Should be Scale-Aware
Prateek Kumar Rajput, Abdoul Aziz Bonkoungou, Alberick Euraste Djiré, Xunzhu Tang, Yewei Song, Iyiola Emmanuel Olatunji, El Hacen Diallo, Jacques Klein, Tegawendé F. Bissyandé
University of Luxembourg
cs.SE, cs.AI
Submitted: 2026-08-13
Updated: 2026-08-14
Comments: 26 pages, 6 figures, 6 tables. Under review at EMSE
Code: https://github.com/pkrajput/isomorphic
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 75/100
The gist: The paper investigates whether standard memorization diagnostics for code LLMs remain effective as model scale increases, and proposes a new methodology to separate memorization from representational
Terminology
Summary
The paper investigates whether standard memorization diagnostics for code LLMs remain effective as model scale increases, and proposes a new methodology to separate memorization from representational robustness.
The authors find that the field's standard probes still separated seen problems from unseen ones at frontier scale
is a hypothesis that fails. Specifically, they test two types of probes:
-
Encoder-side probes (synonym fuzzing and dead-code insertion):
Synonym fuzzing remains diagnostic for small and some mid-tier models, but is nearly inert at the frontier. Frontier models absorb synonym perturbation.
For example,Gemini-2.0-Flash loses at most 5.4 points, while Gemini-2.5-Flash loses even less, with a maximum drop of 2.1 points.
-
Decoder-side probes (CoDeC, which measures token log-likelihood shifts with in-context examples):
CoDeC's seen-unseen gap shrinks from 58 points (Pythia-12B) to 5.5 points (Llama-3.1-405B), and AUC drops from 100% to 75%.
The paper's central contribution is the concept of representational load and a new protocol using I/O isomorphisms (metamorphic testing). The authors apply bijective value transforms (affine, base-conversion, cubic) to test-case integers and append an explicit encode/decode contract to the prompt. This keeps the algorithmic task provably fixed (f(x) = y ⇐⇒ Tθ(f(Tθ−1(x′))) = y′) while forcing the model to decode inputs, run the algorithm, and re-encode outputs.
Key findings from the isomorphism experiments:
-
Representational load degrades performance even in scaled models:
The Iso contract drops frontier-model pass@1 by 14–30 absolute points on a provably fixed task.
-
The degradation is primarily decoder-side:
Decomposing the contract shows that most of this penalty comes from output-side encoding rather than input-side decoding, pointing to contract compliance during generation as the dominant cost.
For Gemini-2.0-Flash on MBPP,decoding inputs alone (Iso (Enc only)) costs a modest 9.7 points (80.8→71.1), whereas encoding outputs token by token (Iso (Dec only)) costs a far larger 21.8 points (80.8→59.0).
-
Scaled models preserve the algorithm:
Despite large pass@1 drops, frontier models show near-zero opcode JSD under isomorphism, which indicates they attempt the same algorithmic core and mainly fumble the representational channel.
This is described as anarrow channel
wherethe same solution families still appear, only less often.
-
Smaller models lose the algorithm:
Smaller models, by contrast, abandon their strategies under the same load on complex tasks, a pattern more consistent with shallow and fragile task comprehension than with mere mis-serialization.
The paper also includes ablation studies showing the brittleness generalizes across three bijection families (affine, base conversion, cubic), and that prompt length is not a confound (dead-code insertion matching the token count causes only 1–7% drops vs. 24–51% for Iso).
The authors conclude: We do not claim that memorization is absent from scaled code LLMs, only that the diagnostics commonly used to detect it appear to stop discriminating at scale.
They recommend that Treat a bare pass@k drop as a starting point rather than a memorization finding, separate representational load before attributing any drop to recall, and report solution-space stability alongside correctness.
Improvements for AI systems
Improvements to AI Systems:
-
Add a representational-load stress-test module to code-generation pipelines. Before deploying a code LLM, run the Iso contract (affine/base/cubic transforms with explicit encode/decode instructions) on a held-out task set. If pass@1 drops >10 absolute points, flag the model as representationally brittle and require retraining or fine-tuning on transformed I/O pairs.
-
Replace single-number memorization scores with a two-axis evaluation report. For any benchmark result, report both (a) correctness under standard prompts and (b) solution-space stability (e.g., opcode JSD under isomorphism). This prevents false confidence from models that pass standard tests but fail under representational load.
-
Train with isomorphism-augmented data. During fine-tuning, randomly apply bijective value transforms (affine, base-64, cubic) to input integers and append the corresponding encode/decode contract to the prompt. This forces the model to learn the algorithm independent of surface representation, reducing decoder-side encoding costs.
-
Implement a decoder-side compliance regularizer. Since output-side encoding is the dominant cost (e.g., 21.8 vs. 9.7 points for Gemini-2.0-Flash), add a training objective that penalizes token-level deviations when the model must re-encode outputs under a contract. This can be done via contrastive loss between standard and transformed output sequences.
-
Create an adaptive probe selector for model scaling. Instead of using fixed memorization probes, dynamically choose probes based on model size. For models >100B parameters, use isomorphism-based probes; for smaller models, use synonym fuzzing or dead-code insertion. This yields reliable memorization signals across scales.
-
Build a
narrow channel
diagnostic tool. For any model, decompose performance drops under Iso into input-decoding vs. output-encoding components. If the output-encoding cost dominates and opcode JSD remains near zero, the model has preserved the algorithm but has a serialization bottleneck—target that bottleneck with constrained decoding or output templates. -
Add prompt-length confound checks to all evaluation harnesses. Automatically insert dead code matching the Iso prompt's token count as a control. If the drop from dead code is <10% of the Iso drop, attribute the performance loss to representational load, not prompt length. This prevents misattribution in future benchmarks.
What the improved AI system can do:
-
Reliably detect memorization vs. robust generalization across model scales, avoiding false negatives at frontier size.
-
Self-identify its own representational brittleness before deployment, allowing preemptive mitigation (e.g., adding output encoders).
-
Generate code that remains correct under arbitrary input/output format changes, not just canonical forms.
-
Provide interpretable failure modes—distinguishing
knows the algorithm but can't serialize
fromdoesn't understand the task
—enabling targeted fixes. -
Benchmark new models fairly by separating confounds (prompt length, token count) from true algorithmic capability.
Abstract
The extent to which large language models for code rely on memorization over genuine understanding remains highly debated. While current literature frequently reports widespread memorization, evaluating the underlying probing techniques across dense architectures reveals a severe breakdown in their utility at scale. Traditional encoder-style probes using perturbations such as synonym fuzzing or dead-code insertion struggle to expose memorization in scaled models, even on known-contaminated benchmarks, and decoder-style probes that rely on log probabilities show similar performance degradation. The specific mode of failure for these probes, particularly why such techniques disrupt smaller models but fail to impact larger ones, motivates us to untangle representation load from memorization rather than treating them as a single phenomenon. By applying invertible mathematical transforms to numeric problems, we isolate these two factors and reveal that scaled encoders successfully absorb substantial representation load while still converging on the correct family of solutions. In practical software engineering, this ability to adapt to varying surface forms is what truly matters for usability and generalizability in LLM and agentic applications. Whether a specific solution was seen during training becomes a much less pressing question because although memorization inflates scores on contaminated benchmarks, factoring out representation load makes it debatable how much we should truly care if a functional answer was originally memorized. Future evaluations must therefore be built around separating these phenomena rather than relying on methodologies that quietly entangle them.
Sources
- Evaluating Large Language Models Trained on Code
- Program Synthesis with Large Language Models
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- SWE-bench Goes Live!
- The SWE-Bench Illusion: When State-of-the-Art LLMs Remember Instead of Reason
- Memorization or Interpolation ? Detecting LLM Memorization through Input Perturbation Analysis
- Learned or Memorized ? Quantifying Memorization Advantage in Code LLMs
- Detecting Data Contamination in LLMs via In-Context Learning
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- Detecting Pretraining Data from Large Language Models
- Studying Large Language Model Generalization with Influence Functions
- Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets
- Progress measures for grokking via mechanistic interpretability
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM
- Scaling Laws for Neural Language Models
- Emergent Abilities of Large Language Models
- BigO(Bench): Can LLMs Generate Code with Controlled Time and Space Complexity?
- ARC Prize 2024: Technical Report
- ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties