Competence-Preserving Resume Perturbations Expose Presentation Sensitivity in LLM Screening
cs.CL, cs.AI
Submitted: 2026-09-15
Updated: 2026-09-15
Comments: Under Peer Review
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: Resume screeners must infer job-relevant competence from resumes whose presentation can vary substantially in wording, structure, stylistic polish, and document extraction quality.
Terminology
Abstract
Resume screeners must infer job-relevant competence from resumes whose presentation can vary substantially in wording, structure, stylistic polish, and document extraction quality. Ideally, such surface variation should not change decisions when the underlying qualification evidence is unchanged. We introduce a controlled audit of this property, constructing occupation-grounded candidate profiles at controlled competence levels and rendering each profile into multiple resume presentations. A deterministic validation gate excludes variants that alter the underlying evidence before scoring. Across six open instruction-tuned LLM conditions, we find a clear disconnect between screening validity and presentation stability. Llama-3.1-8B with its native chat template achieves the strongest validity (0.781) yet reverses 29.6% of matched pairwise decisions under competence-preserving presentation changes; Mistral-7B-v0.3 reaches validity 0.644 with a 41.4% flip rate. Native chat formatting improves validity for several chat-tuned models but does not remove this instability. These results show that resume-screening evaluations should assess not only whether a system identifies stronger candidates, but also whether those decisions remain stable when the same competence evidence is presented differently.
Sources
- Mistral 7B
- The Llama 3 Herd of Models
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- Qwen2.5 Technical Report
- Measuring Validity in LLM-based Resume Screening
- Can LLMs Hire Fairly? Racial Bias in Resume Screening
- Resume Screening, Fast and Slow: (Biased) AI Recommendations' Influence on Human Decision Making
- Gemma: Open Models Based on Gemini Research and Technology
- AI Self-preferencing in Algorithmic Hiring: Empirical Evidence and Insights
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering