Memorization Diagnostics for Code LLMs Should be Scale-Aware

arXiv:2608.12771 · cs.SE, cs.AI · Submitted 2026-08-13 · Read on arXiv

Prateek Kumar Rajput, Abdoul Aziz Bonkoungou, Alberick Euraste Djiré, Xunzhu Tang, Yewei Song, Iyiola Emmanuel Olatunji, El Hacen Diallo, Jacques Klein, Tegawendé F. Bissyandé

University of Luxembourg

cs.SE, cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: 26 pages, 6 figures, 6 tables. Under review at EMSE

Code: https://github.com/pkrajput/isomorphic

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 75/100

The gist: The paper investigates whether standard memorization diagnostics for code LLMs remain effective as model scale increases, and proposes a new methodology to separate memorization from representational

Terminology

Summary

The paper investigates whether standard memorization diagnostics for code LLMs remain effective as model scale increases, and proposes a new methodology to separate memorization from representational robustness.

The authors find that the field's standard probes still separated seen problems from unseen ones at frontier scale is a hypothesis that fails. Specifically, they test two types of probes:

  1. Encoder-side probes (synonym fuzzing and dead-code insertion): Synonym fuzzing remains diagnostic for small and some mid-tier models, but is nearly inert at the frontier. Frontier models absorb synonym perturbation. For example, Gemini-2.0-Flash loses at most 5.4 points, while Gemini-2.5-Flash loses even less, with a maximum drop of 2.1 points.

  2. Decoder-side probes (CoDeC, which measures token log-likelihood shifts with in-context examples): CoDeC's seen-unseen gap shrinks from 58 points (Pythia-12B) to 5.5 points (Llama-3.1-405B), and AUC drops from 100% to 75%.

The paper's central contribution is the concept of representational load and a new protocol using I/O isomorphisms (metamorphic testing). The authors apply bijective value transforms (affine, base-conversion, cubic) to test-case integers and append an explicit encode/decode contract to the prompt. This keeps the algorithmic task provably fixed (f(x) = y ⇐⇒ Tθ(f(Tθ−1(x′))) = y′) while forcing the model to decode inputs, run the algorithm, and re-encode outputs.

Key findings from the isomorphism experiments:

  • Representational load degrades performance even in scaled models: The Iso contract drops frontier-model pass@1 by 14–30 absolute points on a provably fixed task.

  • The degradation is primarily decoder-side: Decomposing the contract shows that most of this penalty comes from output-side encoding rather than input-side decoding, pointing to contract compliance during generation as the dominant cost. For Gemini-2.0-Flash on MBPP, decoding inputs alone (Iso (Enc only)) costs a modest 9.7 points (80.8→71.1), whereas encoding outputs token by token (Iso (Dec only)) costs a far larger 21.8 points (80.8→59.0).

  • Scaled models preserve the algorithm: Despite large pass@1 drops, frontier models show near-zero opcode JSD under isomorphism, which indicates they attempt the same algorithmic core and mainly fumble the representational channel. This is described as a narrow channel where the same solution families still appear, only less often.

  • Smaller models lose the algorithm: Smaller models, by contrast, abandon their strategies under the same load on complex tasks, a pattern more consistent with shallow and fragile task comprehension than with mere mis-serialization.

The paper also includes ablation studies showing the brittleness generalizes across three bijection families (affine, base conversion, cubic), and that prompt length is not a confound (dead-code insertion matching the token count causes only 1–7% drops vs. 24–51% for Iso).

The authors conclude: We do not claim that memorization is absent from scaled code LLMs, only that the diagnostics commonly used to detect it appear to stop discriminating at scale. They recommend that Treat a bare pass@k drop as a starting point rather than a memorization finding, separate representational load before attributing any drop to recall, and report solution-space stability alongside correctness.

Improvements for AI systems

Improvements to AI Systems:

  1. Add a representational-load stress-test module to code-generation pipelines. Before deploying a code LLM, run the Iso contract (affine/base/cubic transforms with explicit encode/decode instructions) on a held-out task set. If pass@1 drops >10 absolute points, flag the model as representationally brittle and require retraining or fine-tuning on transformed I/O pairs.

  2. Replace single-number memorization scores with a two-axis evaluation report. For any benchmark result, report both (a) correctness under standard prompts and (b) solution-space stability (e.g., opcode JSD under isomorphism). This prevents false confidence from models that pass standard tests but fail under representational load.

  3. Train with isomorphism-augmented data. During fine-tuning, randomly apply bijective value transforms (affine, base-64, cubic) to input integers and append the corresponding encode/decode contract to the prompt. This forces the model to learn the algorithm independent of surface representation, reducing decoder-side encoding costs.

  4. Implement a decoder-side compliance regularizer. Since output-side encoding is the dominant cost (e.g., 21.8 vs. 9.7 points for Gemini-2.0-Flash), add a training objective that penalizes token-level deviations when the model must re-encode outputs under a contract. This can be done via contrastive loss between standard and transformed output sequences.

  5. Create an adaptive probe selector for model scaling. Instead of using fixed memorization probes, dynamically choose probes based on model size. For models >100B parameters, use isomorphism-based probes; for smaller models, use synonym fuzzing or dead-code insertion. This yields reliable memorization signals across scales.

  6. Build a narrow channel diagnostic tool. For any model, decompose performance drops under Iso into input-decoding vs. output-encoding components. If the output-encoding cost dominates and opcode JSD remains near zero, the model has preserved the algorithm but has a serialization bottleneck—target that bottleneck with constrained decoding or output templates.

  7. Add prompt-length confound checks to all evaluation harnesses. Automatically insert dead code matching the Iso prompt's token count as a control. If the drop from dead code is <10% of the Iso drop, attribute the performance loss to representational load, not prompt length. This prevents misattribution in future benchmarks.

What the improved AI system can do:

  • Reliably detect memorization vs. robust generalization across model scales, avoiding false negatives at frontier size.

  • Self-identify its own representational brittleness before deployment, allowing preemptive mitigation (e.g., adding output encoders).

  • Generate code that remains correct under arbitrary input/output format changes, not just canonical forms.

  • Provide interpretable failure modes—distinguishing knows the algorithm but can't serialize from doesn't understand the task—enabling targeted fixes.

  • Benchmark new models fairly by separating confounds (prompt length, token count) from true algorithmic capability.

Abstract

The extent to which large language models for code rely on memorization over genuine understanding remains highly debated. While current literature frequently reports widespread memorization, evaluating the underlying probing techniques across dense architectures reveals a severe breakdown in their utility at scale. Traditional encoder-style probes using perturbations such as synonym fuzzing or dead-code insertion struggle to expose memorization in scaled models, even on known-contaminated benchmarks, and decoder-style probes that rely on log probabilities show similar performance degradation. The specific mode of failure for these probes, particularly why such techniques disrupt smaller models but fail to impact larger ones, motivates us to untangle representation load from memorization rather than treating them as a single phenomenon. By applying invertible mathematical transforms to numeric problems, we isolate these two factors and reveal that scaled encoders successfully absorb substantial representation load while still converging on the correct family of solutions. In practical software engineering, this ability to adapt to varying surface forms is what truly matters for usability and generalizability in LLM and agentic applications. Whether a specific solution was seen during training becomes a much less pressing question because although memorization inflates scores on contaminated benchmarks, factoring out representation load makes it debatable how much we should truly care if a functional answer was originally memorized. Future evaluations must therefore be built around separating these phenomena rather than relying on methodologies that quietly entangle them.

Sources

Related papers