Query Timing Produces Opposite Positional Biases Between LLMs and Humans
Jasin Cekinmez, Addison J. Wu, Thomas L. Griffiths
cs.CL, cs.AI
Submitted: 2026-08-01
Updated: 2026-08-14
Comments: Entropic Award (Top 3 Paper), ICBINB @ ICLR 2026
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Positional biases such as recency and primacy effects have been documented in large language models (LLMs), yet the underlying mechanism by which these models make their evaluations remains poorly
Terminology
Abstract
Positional biases such as recency and primacy effects have been documented in large language models (LLMs), yet the underlying mechanism by which these models make their evaluations remains poorly understood. Both primacy and recency biases have been observed in human judgments in response to evidence, but recent work suggest that when the listener updates their beliefs -- during the presentation of evidence or only at the end -- influences the presence of such effects. We investigate whether a similar phenomenon holds for LLMs, finding divergence from human behavior. These biases are more exacerbated in newer models compared to their predecessors.
Sources
- LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods
- JudgeLM: Fine-tuned Large Language Models are Scalable Judges
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering