Resist, Update, Reject: Preference Optimization Installs a Prior-Dependent Reliability Switch

arXiv:2609.22359 · cs.LG, cs.AI · Submitted 2026-07-26 · Read on arXiv

cs.LG, cs.AI

Submitted: 2026-07-26

Updated: 2026-07-26

License: http://creativecommons.org/licenses/by/4.0/

The gist: An aligned model asked to hold its answer against a manipulative source must still update on a reliable one and reject an unreliable one: resistance, reliable-update, and unreliable-source rejection

Terminology

Abstract

An aligned model asked to hold its answer against a manipulative source must still update on a reliable one and reject an unreliable one: resistance, reliable-update, and unreliable-source rejection are one three-way contract, not three independent behaviors. We show the objective most anti-sycophancy work optimizes is non-identifying with respect to source reliability: because no preference label depends on whether a source is actually reliable, any scalar mixture of the arms traces a single deference dial, and no point separates two same-template testimonies differing only in stated reliability. This fixation gullibility frontier is a property of the objective, not any model. We make reliability identifiable through data: a threshold benchmark where a source asserts the opposite answer while stating its reliability r, and the correct action is to flip iff r exceeds the model's prior strength p. Preference optimization over balanced coverage installs a prior-dependent reliability switch: across three seeds on Qwen2.5-7B-Instruct the threshold r rises monotonically with the prior, decision accuracy reaches 0.84 with a monotone flip curve (Spearman 0.56), and the policy generalizes to unseen reliability values and a held-out notation, following stated reliability over role prestige. Three controls localize the cause: an unmatched variant installs the switch equally (0.80), a second preference optimizer (IPO) installs it just as well (0.86), whereas supervised imitation does not (0.50), so the cause is preference optimization over reliability-labeled coverage, not pairing, loss, or imitation. A confirmatory battery replicates the switch on a fresh test draw, bounds it honestly (it keys on reliability stated in the testimony, not a separately audited record), and transfers it to Llama-3.1-8B. The frontier is empirical, not a theorem.

Related papers