The Truncation Blind Spot: How Decoding Strategies Systematically Exclude Human-Like Token Choices
cs.CL, cs.LG, stat.ML
Submitted: 2026-03-19
Updated: 2026-09-23
Code: https://github.com/EstebanGarces/human_vs_machine
Terminology
Sources
- Mirostat: A Neural Text Decoding Algorithm that Directly Controls Perplexity
- DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
- Closing the Curious Case of Neural Text Degeneration
- GLTR: Statistical Detection and Visualization of Generated Text
- The Llama 3 Herd of Models
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- Mistral 7B
- Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)
- Falcon2-11B Technical Report
- Pointer Sentinel Mixture Models
- Turning Up the Heat: Min-p Sampling for Creative and Coherent LLM Outputs
- Automatic Detection of Machine Generated Text: A Critical Survey
- Release Strategies and the Social Impacts of Language Models
- Contrastive Search Is What You Need For Neural Text Generation
- A Contrastive Framework for Neural Text Generation
- Qwen2 Technical Report
- OPT: Open Pre-trained Transformer Language Models
- Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering