The Judge Is Not Its Twin: Post-training makes a model's writing more predictable but barely moves its taste, as a judge, toward predictable writing
cs.CL, cs.LG
Submitted: 2026-09-26
Updated: 2026-09-26
Code: https://github.com/IamArmanNikkhah/judge-is-not-its-twin
Terminology
Sources
- Art or Artifice? Large Language Models and the False Promise of Creativity
- Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
- The Curious Case of Neural Text Degeneration
- Mistral 7B
- Where does output diversity collapse in post-training?
- Preference Leakage: A Contamination Problem in LLM-as-a-judge
- Creativity Has Left the Chat: The Price of Debiasing Language Models
- LLM Evaluators Recognize and Favor Their Own Generations
- Is Temperature the Creativity Parameter of Large Language Models?
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- 2 OLMo 2 Furious
- Zephyr: Direct Distillation of LM Alignment
- The Limits of Automatic Evaluation of Creativity in Large Language Models
- Large Language Models are not Fair Evaluators
- Self-Preference Bias in LLM-as-a-Judge
- Trading Off Diversity and Quality in Natural Language Generation
- Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering