MemoReason: Evaluating the Effect of Parametric Memory on Contextual Reasoning in LLMs
cs.CL
Submitted: 2026-09-28
Updated: 2026-09-28
Terminology
Sources
- The Llama 3 Herd of Models
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- Beyond Consensus: Mitigating the Agreeableness Bias in LLM Judge Evaluations
- Let's Verify Step by Step
- RepLiQA: A Question-Answering Dataset for Benchmarking LLMs on Unseen Reference Content
- Capabilities of GPT-4 on Medical Challenge Problems
- Olmo 3
- gpt-oss-120b & gpt-oss-20b Model Card
- Impact of Pretraining Term Frequencies on Few-Shot Reasoning
- Mathematical Reasoning in Large Language Models: Assessing Logical and Arithmetic Errors across Wide Numerical Ranges
- Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap
- Cofca: A Step-Wise Counterfactual Multi-hop QA benchmark
- VERINA: Benchmarking Verifiable Code Generation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering