TELLME: Test-Enhanced Learning for Language Model Enrichment
Minjun Kim, Inho Won, Hyeonseok Lim, MinKyu Kim, Junghun Yuk, Wooyoung Go, Jongyoul Park, Jungyeul Park, KyungTae Lim
Korea Advanced Institute of Science and Technology · Seoul National University of Science and Technology · National Security Research Institute
cs.CL, cs.AI
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: Findings of the Association for Computational Linguistics: EACL 2026
Journal ref: Findings of the Association for Computational Linguistics: EACL 2026, pages 1655-1677
DOI: 10.18653/v1/2026.findings-eacl.84
Code: https://github.com/microsoft/LMOpshttps:
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 95/100
The gist: (1) difficulty in acquiring large-scale domain-specific datasets and (2) high computational costs.
Terminology
Summary
Summary
The paper introduces TELLME (Test-Enhanced Learning for Language Model Enrichment), a novel continual pre-training (CPT) method for large language models (LLMs) that applies the Test-Enhanced Learning (TEL) principle from educational psychology to improve domain-specific knowledge acquisition and long-term memory retention.
Problem and Motivation
Continual pre-training (CPT) is widely used for domain adaptation in LLMs but faces two major challenges: (1) difficulty in acquiring large-scale domain-specific datasets and (2) high computational costs. Existing methods like I NST PT integrate QA data into CPT, but they deviate from effective testing strategies suggested by TEL. Research shows that open-ended explanatory responses, rather than simple recall or multiple-choice formats, yield stronger long-term retention. The QA format in I NST PT is largely constrained by the given text, limiting its ability to elicit internal knowledge.
Proposed Method
TELLME extends conventional CPT by jointly training plain text with descriptive QA samples that require explanatory reasoning beyond the given context. The method:
-
Constructs 100K domain-specific TELLME samples using GPT-4o-mini in a cost-efficient manner (approximately 12 total)
-
Ensures high question diversity to stimulate the model's intrinsic knowledge
-
Uses a training framework where each sample X = (t, q, a) contains plain text (t), question (q), and answer (a)
-
Configures the loss function so that 1(xi ∈ t ∪ a) = 1 and 1(xi ∈ q) = 0, meaning the model predicts only plain text and answer tokens while excluding question tokens from loss computation
The dataset generation prompt instructs the model to: (1) avoid direct questions about the given excerpt, (2) create questions based on general domain knowledge, and (3) ensure questions can be answered independently of the excerpt. This promotes concept mapping
between related pieces of knowledge not explicitly mentioned in the text.
Experimental Setup
Experiments were conducted in financial and medical domains using:
-
Training data: 100k PubMed abstracts (medical) and 100k Bloomberg financial news articles (finance)
-
Base models: Llama-3.2-1B, Llama-3.2-3B, Llama-3.1-8B, SmolLM2-1.7B
-
Evaluation benchmarks: FOMC, NIFTY, MMLU-F (finance); HeadQA, MedMCQA, MMLU-C (medicine)
-
Comparisons: baseline, +CPT, +CPT+IT, +I NST PT, +TELLME
Key Results
-
Overall Performance: TELLME achieves the highest average performance, outperforming CPT+IT by 10.0% and I NST PT by 6.3% across both domains.
-
Domain Adaptation: In finance, TELLME consistently outperforms the baseline across all models with an average gain of approximately 9.8%. In medicine, TELLME shows an average improvement of 0.09 points over baseline and outperforms I NST PT by 2.58 points.
-
Perplexity: TELLME achieves the lowest perplexity across all sequence lengths in the medical domain, while +CPT+IT exhibits the highest perplexity (QA-focused fine-tuning dilutes plain text representational capacity).
-
Long-Term Retention: In experiments where models trained on finance data were subsequently trained on medical data, TELLME showed only a 0.94% reduction in finance performance compared to 5.72% for CPT. TELLME achieved a 9.8% higher final performance (3.15-point increase) compared to CPT-based approaches.
-
Training Efficiency: TELLME achieves the same perplexity 1.4 times faster than CPT.
Ablation Studies
-
Inv-TEL (QA before plain text): Performed approximately 2.82 points lower than TELLME, showing QA placement matters
-
PIT (QA and plain text as independent samples): Approximately 0.8 points lower than TELLME, showing integration within a single sample is more efficient
-
TEL-Q/L (loss on all tokens including questions): Approximately 0.42 points lower than TELLME, showing excluding question loss improves performance
-
Synthesizer robustness: TELLME remains effective even with less advanced synthesizers (Mistral-7B) or self-generated datasets, though GPT-4o-mini yields superior performance
Additional Findings
-
Multilingual Generalization: TELLME - KO (Korean version) achieves +8.4-point improvement in average accuracy on KoBEST benchmark for OLMo2-1B, with over +20-point gains on sentence understanding tasks
-
Scalability: TELLME shows improvements across various model sizes including GPT-2, Qwen2.5, phi-1.5, and even a 70B model (with LoRA and 4-bit quantization, achieving 1.08 point improvement)
-
Data Quality: Generated data achieved an average score of 4.03 out of 5 in LLM-as-a-judge quality evaluation
Conclusion
The paper demonstrates that applying TEL principles to CPT—specifically using descriptive, open-ended QA that requires intrinsic knowledge recall—improves both the efficiency of knowledge acquisition and the durability of learned representations in LLMs. TELLME achieved up to a 23.6% performance improvement on finance benchmarks compared to existing methods.
Improvements for AI systems
Improvements to AI Systems:
- Enhanced Domain Adaptation with Reduced Data Requirements
-
Implement TELLME's training framework (plain text + descriptive QA with excluded question-token loss) to enable domain-specific knowledge acquisition using only 100K samples, cutting data collection costs by 90% compared to traditional CPT.
-
The system can adapt to new domains (e.g., legal, scientific, technical) with minimal curated data, making it feasible for niche fields with scarce resources.
- Improved Long-Term Knowledge Retention
-
Integrate the TEL-based QA generation (open-ended, explanatory, independent of source text) into continual pre-training to reduce catastrophic forgetting.
-
The system retains 9.8% more prior-domain knowledge when sequentially trained on multiple domains, enabling lifelong learning without performance collapse (e.g., finance → medicine transition shows only 0.94% drop vs. 5.72% for standard CPT).
- Faster Training Convergence
-
Adopt TELLME's loss configuration (predicting only plain text and answer tokens) to achieve target perplexity 1.4× faster than standard CPT.
-
The system reduces compute costs for domain adaptation, allowing more frequent updates or deployment on resource-constrained hardware.
- Higher-Quality Reasoning and Knowledge Elicitation
-
Generate diverse, context-independent QA pairs (via GPT-4o-mini or self-generation) to force the model to retrieve and synthesize internal knowledge rather than copying from input.
-
The system improves performance on downstream reasoning benchmarks (e.g., +23.6% on finance tasks, +8.4% on Korean sentence understanding) by strengthening conceptual mapping between related facts.
- Scalable and Model-Agnostic Enhancement
-
Apply TELLME across model sizes (1B to 70B) and architectures (Llama, SmolLM, GPT-2, Qwen) with consistent gains, including via LoRA and 4-bit quantization for large models.
-
The system can be deployed as a plug-and-play continual pre-training module, improving any LLM's domain expertise without architectural changes.
- Robust Multilingual and Low-Resource Generalization
-
Use TELLME's data synthesis pipeline to generate descriptive QA in any language, enabling rapid domain adaptation for under-resourced languages.
-
The system achieves significant accuracy boosts (e.g., +20 points on Korean sentence understanding) with minimal additional data, supporting global deployment.
- Reduced Hallucination via Structured Knowledge Recall
-
By training on QA that requires answers independent of source text, the model learns to ground responses in its internal knowledge base, reducing reliance on surface-level text matching.
-
The system produces more factual and contextually accurate outputs in domain-specific applications (e.g., medical Q&A, financial analysis), as evidenced by lower perplexity and higher benchmark scores.
Abstract
Continual pre-training (CPT) has been widely adopted as a method for domain adaptation in large language models. However, CPT has consistently been accompanied by challenges, such as the difficulty of acquiring large-scale domain-specific datasets and high computational costs. In this study, we propose a novel method called Test-Enhanced Learning for Language Model Enrichment (TELLME) to alleviate these issues. TELLME leverages the TestEnhanced Learning (TEL) principle, whereby the model's training efficiency is improved using quizzes during training. It integrates this principle with CPT, thereby promoting efficient domain-specific knowledge acquisition and long-term memory retention. Experimental results demonstrate that TELLME outperforms existing methods by up to 23.6% in the financial domain and achieves a 9.8% improvement in long-term memory retention.
Sources
- SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
- GPT-4 Technical Report
- Towards Effective and Efficient Continual Pre-training of Large Language Models
- MEDITRON-70B: Scaling Medical Pretraining for Large Language Models
- SaulLM-7B: A pioneering Large Language Model for Law
- The Llama 3 Herd of Models
- Continual Pre-Training for Cross-Lingual LLM Adaptation: Enhancing Japanese Language Capabilities
- Don't Stop Pretraining: Adapt Language Models to Domains and Tasks
- Mistral 7B
- Instruction-tuned Language Models are Better Knowledge Learners
- Demystifying Domain-adaptive Post-training for Financial LLMs
- Learning Beyond the Surface: How Far Can Continual Pre-Training with LoRA Enhance LLMs' Domain-Specific Insight Learning?
- NIFTY Financial News Headlines Dataset
- Continual Learning for Large Language Models: A Survey
- PLLaMa: An Open-source Large Language Model for Plant Science
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering