LeakScale: Estimating the Causal Effect of Benchmark Exposure
cs.CL
Submitted: 2026-09-23
Updated: 2026-09-24
Terminology
Sources
- Language Models are Few-Shot Learners
- Quantifying Memorization Across Neural Language Models
- Recent Advances in Large Langauge Model Benchmarks against Data Contamination: From Static to Dynamic Evaluation
- Unveiling the Spectrum of Data Contamination in Language Models: A Survey from Detection to Remediation
- Gemma 4 Technical Report
- Time Travel in LLMs: Tracing Data Contamination in Large Language Models
- Is Your Benchmark Still Useful? Dynamic Benchmarking for Code Language Models
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- Proving Test Set Contamination in Black Box Language Models
- NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark
- Quantifying the Effect of Test Set Contamination on Generative Evaluations
- Detecting Pretraining Data from Large Language Models
- Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?
- Detecting Data Contamination from Reinforcement Learning Post-training for Large Language Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- On The Fragility of Benchmark Contamination Detection in Reasoning Models
- LiveBench: A Challenging, Contamination-Limited LLM Benchmark
- Benchmark Data Contamination of Large Language Models: A Survey
- Qwen3 Technical Report
- Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering