LLMs learn different forms of metacognition when trained to predict their own accuracy
cs.CL
Submitted: 2026-09-27
Updated: 2026-09-29
Code: https://github.com/Nicolas-Yax/LLM-MetaCog
Terminology
Sources
- QLoRA: Efficient Finetuning of Quantized LLMs
- MAGELLAN: Metacognitive predictions of learning progress guide autotelic LLM agents in large goal spaces
- The Llama 3 Herd of Models
- Mistral 7B
- Why Language Models Hallucinate
- Large Language Models Must Be Taught to Know What They Don't Know
- Reported Confidence in LLMs Tracks Commitment More Than Correctness
- Ministral 3
- MedMCQA : A Large-scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering
- Qwen2.5 Technical Report
- Large Language Models are overconfident and amplify human bias
- Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
- Towards Objective Fine-tuning: How LLMs' Prior Knowledge Causes Potential Poor Calibration?
- Ethical and social risks of harm from Language Models
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
- Qwen3 Technical Report
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering