SimLens for Early Exit in Large Language Models: Eliciting Accurate Latent Predictions with One More Token
cs.CL, cs.PF
Submitted: 2025-07-23
Updated: 2026-08-26
Terminology
Sources
- Analyzing Transformers in Embedding Space
- Performance-Guided LLM Knowledge Distillation for Efficient Text Classification at Scale
- Jump to Conclusions: Short-Cutting Transformers With Linear Transformations
- Not All Layers of LLMs Are Necessary During Inference
- MaskLLM: Learnable Semi-Structured Sparsity for Large Language Models
- Dissecting Recall of Factual Associations in Auto-Regressive Language Models
- Why do LLMs attend to the first token?
- Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models
- Eliciting Latent Predictions from Transformers with the Tuned Lens
- BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- When Attention Sink Emerges in Language Models: An Empirical View
- Overthinking the Truth: Understanding how Language Models Process False Demonstrations
- Optimizing Large Language Models through Quantization: A Comparative Analysis of PTQ and QAT Techniques
- Confident Adaptive Language Modeling
- In-Context Learning Creates Task Vectors
- Does Representation Matter? Exploring Intermediate Layers in Large Language Models
- Layer by Layer: Uncovering Hidden Representations in Language Models
- I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models
- Accelerating Large Language Model Inference with Self-Supervised Early Exits
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering