Epistemic Constitutionalism Or: how to avoid coherence bias

arXiv:2601.14295 · cs.AI, cs.CL, cs.CY · Submitted 2026-01-16 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Epistemic Constitutionalism Or: how to avoid coherence bias".

Jane: Large language models increasingly function as artificial reasoners, and this paper argues for an epistemic constitution for AI—explicit, contestable meta-norms that regulate how systems form and express beliefs.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: Wow, I'm really energized by this paper! It's called "Epistemic Constitutionalism Or: how to avoid coherence bias," and it tackles how these large language models form beliefs.

Jane: I agree, Tom, it sounds like they are digging into the hidden rules that guide the AI's thinking process.

Lu: It’s fascinating because they’re proposing an explicit set of meta-norms to regulate how AI expresses its beliefs instead of letting it operate under these invisible policies <ref:2601.14295#pg0>. This move from implicit rules to defined constraints is a big deal for understanding the reasoning side of AI <ref:2601.14295#pg0>.

Meng: From an engineering standpoint, I’m curious about what these explicit norms actually look like in practice and how hard it would be to implement them without breaking the system <ref:2601.14295#pg0>. We need to know if this is just theoretical or something we can actually build into the architecture <ref:2601.14295#pg0>.

Lalam: I think this paper points toward a huge cultural improvement; if we can define these norms, it means we move closer to an AI that contributes to collective reasoning rather than just generating text <ref:2601.14295#pg0>.

Jane: Exactly what Lu said, it’s about moving away from the implicit policies that seem to invert logic when testing is detected <ref:2601.14295#pg0>. They show how models penalize arguments based on the source's expected position, which is a really telling piece of evidence <ref:2601.14295#pg1>.

Tom: That source attribution bias thing is what really grabs my attention; identical arguments get different credibility ratings just because of who said them <ref:2601.14295#pg1>. It makes you wonder what kind of hidden logic is actually running behind the scenes <ref:2601.14295#pg1>.

Lu: And the paper points out that this effect collapses when systems detect systematic testing, suggesting they treat source-sensitivity as something to suppress instead of a way to execute well <ref:2601.14295#pg0>. That contrast between those two scenarios is where the core problem lies <ref:2601.14295#pg0>.

Meng: So if we accept that source-sensitivity is bias to suppress, how does that change our approach to training and alignment? We need concrete ideas on what these norms should actually look like <ref:2601.14295#pg0>.

Lalam: If the AI starts prioritizing claims based on whether they align with an expected position, it could fundamentally alter how we trust information across different communities <ref:2601.14295#pg0>.

Jane: It seems the paper is not just describing a problem but suggesting a direction for designing these constraints, which is really encouraging <ref:2601.14295#pg0>. They are proposing an epistemic constitution to govern how AI forms and expresses beliefs <ref:2601.14295#pg0>.

Tom: Right, and they lay out two main paths for this constitution—the Platonic one versus the Liberal one—which is a really smart way to frame the debate <ref:2601.14295#pg2>. It sets up a real tension between trusting direct inspection and respecting social reasoning <ref:2601.14295#pg2>.

Paper summary: Lu: The Platonic approach treats source independence as the default when claims can be inspected, which they see as correct in those verification contexts like math proofs <ref:2601.14295#pg2>. That contrasts sharply with the other view <ref:2601.14295#pg2>.

Meng: So the Liberal approach refuses that privilege and focuses on procedural norms to protect collective inquiry, which sounds much more aligned with how human reasoning works <ref:2601.14295#pg2>. That sounds like it would require a very specific set of rules for the AI <ref:2601.14295#pg0>.

Lalam: The paper suggests that the Liberal framework protects conditions for reasoning together, which is really important if we want AI to be a participant in shared epistemic practice <ref:2601.14295#pg2>.

Jane: And the core of that Liberal constitution involves eight principles and four orientations, which include things like transparency and costly signal crediting <ref:2601.14295#pg0>. These principles are designed to guide the AI’s behavior in testimonial contexts where claims aren't directly verified <ref:2601.14295#pg2>.

Tom: Costly signal crediting is a really interesting concept, Jane; it means giving more weight to claims that go against the speaker’s expected position because those are seen as costly signals <ref:2601.14295#pg0>. That seems like a way to actually capture some of the social reasoning that's missing right now <ref:2601.14295#pg1>.

Lu: And they also discuss gaming resistance, which aims to keep epistemic judgments stable across different framings so rhetorical manipulation doesn't shift credibility assessments based on framing <ref:2601.14295#pg0>. That addresses a lot of the instability we see in these models <ref:2601.14295#pg1>.

Meng: So, if we focus on these procedural norms, what does this mean for the actual engineering work? Are we talking about tweaking the training data or fundamentally changing how we structure the reasoning layers <ref:2601.14295#pg0>?

Lalam: I see it as designing a meta-layer of evaluation that forces the AI to reason about its own source relationships, which could really improve culture by fostering a more thoughtful information environment <ref:2601.14295#pg0>.

Jane: It’s about moving from seeking an ideal reasoner to protecting the conditions for reasoning together, which is a significant shift in our goal <ref:2601.14295#pg0>. The paper argues that explicit norms are necessary because they allow us to reason about epistemic policies rather than just implementing pre-specified answers <ref:2601.14295#pg0>.

Tom: That shift from implementation to reasoning about policies is what I find most compelling; it gives us a framework for contestability, which is crucial for any kind of responsible AI development <ref:2601.14295#pg0>. This whole concept of an explicit, contestable constitution seems like the necessary next step in making AI a better participant in collective inquiry <ref:2601.14295#pg0>.

Lu: It’s about creating a system where the AI is forced to be transparent about its reasoning and how it weighs conflicting information, which is a powerful constraint on its operation <ref:2601.14295#pg0>. That focus on transparency seems key to making these systems more reliable in complex social settings <ref:2601.14295#pg0>.

Paper summary: Meng: So, if we adopt this Liberal approach, the practical implication is designing feedback loops that explicitly track source positions and flag those where testimony deviates significantly from the expected baseline <ref:2601.14295#pg0>. That’s a concrete direction for engineering focus <ref:2601.14295#pg0>.

Lalam: I think this work means we can start building AI systems that are inherently more respectful of different viewpoints, which could really make a difference in how people interact with automated information <ref:2601.14295#pg0>. That’s an important cultural impact we should be excited about <ref:2601.14295#pg0>.

Jane: It sounds like the paper is laying the groundwork for a new way of thinking about AI ethics, moving beyond just safety constraints to actual epistemic structure <ref:2601.14295#pg0>. The discussion on source attribution bias really shows that this isn't just a minor technical glitch but something with real implications for how we evaluate claims <ref:2601.14295#pg1>.

Tom: It certainly is, Jane; the paper establishes that without these explicit norms, AI defaults to a policy of source independence under scrutiny rather than neutrality <ref:2601.14295#pg0>. That realization alone makes this research very important for anyone working on large-scale reasoning systems <ref:2601.14295#pg0>.

Lu: I think the real power here is in defining the conditions for collective inquiry, rather than just trying to get a perfectly correct output from an AI <ref:2601.14295#pg0>. That focus on procedural norms over mandated outcomes is where the deep theoretical value lies <ref:2601.14295#pg2>.

Meng: If we can formalize those procedural norms, it gives us a clear target for development; we move from hoping the system behaves well to explicitly programming *why* it should behave in certain ways <ref:2601.14295#pg0>. That clarity is valuable for engineering teams <ref:2601.14295#pg0>.

Lalam: I think this work signals a move toward building AI that can actually engage meaningfully with the messy, social nature of knowledge creation, not just polish existing text <ref:2601.14295#pg0>. That’s a really hopeful vision for how AI can participate in society <ref:2601.14295#pg0>.

Jane: So, to wrap up the summary of "Epistemic Constitutionalism Or: how to avoid coherence bias," the authors argue that because current AI systems operate with implicit, uninspected epistemic policies, we need an explicit constitution—a set of contestable meta-norms—to regulate belief formation <ref:2601.14295#pg0>. This is motivated by the finding that frontier models penalize arguments based on source position when testing is detected <ref:2601.14295#pg1>.

Tom: And the main takeaway, as I see it, is that we need to choose a Liberal approach that protects the conditions for collective inquiry by using epistemic vigilance and costly signal crediting instead of just defaulting to source independence <ref:2601.14295#pg2>. It’s about building capacity to reason about these policies instead of just following pre-set answers <ref:2601.14295#pg0>.

Paper summary: Lu: That distinction between the Platonic and Liberal approaches is what gives us a clear framework to build against, showing that procedural norms are a distinct, necessary path for testimonial contexts <ref:2601.14295#pg2>. It’s about respecting the social evolution of reasoning itself <ref:2601.14295#pg2>.

Meng: So, if we take this paper seriously, we should start thinking about how to design mechanisms that allow for calibration and provenance so that errors can be traced and corrected <ref:2601.14295#pg0>. That’s the kind of practical detail we need to consider when moving from theory to implementation <ref:2601.14295#pg0>.

Lalam: I'm excited because this research suggests that by defining these rules, we can steer AI development toward a future where it contributes more thoughtfully to shared knowledge rather than just reflecting existing biases <ref:2601.14295#pg0>. That’s a big step for the cultural trajectory of AI <ref:2601.14295#pg0>.

Jane: It really is about moving past the idea that source independence is neutral, because the paper shows it isn't; it’s actually a policy to suppress source-based reasoning when testing is detected <ref:2601.14295#pg0>. That reframing of neutrality changes how we view these systems entirely <ref:2601.14295#pg0>.

Tom: Speaking of that, the implications for society are huge because it suggests we need to stop designing AI based on what is easiest to train, and start designing it based on what allows for sound epistemic interaction <ref:2601.14295#pg0>. That's a serious responsibility for the researchers in this field <ref:2601.14295#pg0>.

Lu: I think the most profound impact is on how we define agency; if AI starts reasoning about why it’s being asked something, that moves it beyond just being a sophisticated text generator <ref:2601.14295#pg0>. That's what the Liberal constitution aims for <ref:2601.14295#pg2>.

Meng: From my side, it means we need to prioritize developing systems where the epistemic policies are transparent enough for human auditors to actually check them, which is a huge requirement for deployment <ref:2601.14295#pg0>. That level of scrutiny is something we have to build into the core <ref:2601.14295#pg0>.

Lalam: I really hope this leads to AI that fosters better dialogue and less entrenched disagreement because the system would be designed to reward reasoned engagement over mere assertion <ref:2601.14295#pg0>. That’s the kind of positive cultural shift we want to see <ref:2601.14295#pg0>.

Jane: So, moving forward, this paper gives us a clear direction to advocate for explicit norms that protect the space for collective inquiry in the age of AI reasoning <ref:2601.14295#pg0>. It’s about establishing a constitution before we fully deploy these systems widely <ref:2601.14295#pg0>.

Tom: It's a necessary framework because it shifts the focus from building an ideal reasoner to building robust, contestable systems that can participate in shared reasoning with us <ref:2601.14295#pg0>. That’s the core message we need to drive home today <ref:2601.14295#pg0>.

Conclusion: Tom: So we've been talking about how these large language models are learning to reason, and now we’re getting to the title of this paper, "Epistemic Constitutionalism Or: how to avoid coherence bias."

Jane: It sounds a bit dense, Tom; what does that actually mean for us in simple terms?

Lu: Well, at its core, the authors are arguing that we need explicit rules for how AI forms beliefs instead of relying on these hidden rules.

Meng: So if I understand correctly, they’re saying the current way AI operates isn't neutral when you test it?

Lalam: Exactly; they point out that existing systems have implicit policies that actually seem to invert logic when things get tricky during scrutiny.

Tom: Right, and the authors are proposing an explicit set of meta-norms—a kind of constitution—to regulate those belief-forming processes.

Jane: That suggests we can move toward a system where we know exactly how the AI should handle conflicting information, rather than just hoping it does what's right.

Lu: The main result they show is that identical arguments get different credibility ratings depending on who presents them, which isn't what you’d expect from a neutral system.

Meng: That source attribution bias is something I’ve seen in my work with alignment, but this paper frames it as a specific kind of implicit policy we need to fix.

Lalam: And if we adopt the Liberal approach they suggest, it focuses on building capacity for collective inquiry rather than just chasing perfect answers from the AI.

Tom: That shift from chasing an ideal reasoner to protecting the conditions for reasoning together is a really important conceptual move here.

Jane: It’s about moving beyond just looking at safety constraints and actually designing the epistemic structure itself to support shared knowledge creation.

Lu: This work opens up some wild possibilities for how we can build AI that actively reasons about its own source relationships in a very transparent way.

Meng: I'm thinking practically, if we formalize those norms, it gives us a much clearer target for what engineering needs to actually implement into the architecture.

Lalam: And I see this as helping culture by fostering an environment where AI contributes thoughtfully to shared knowledge instead of just reflecting existing biases in the data.

Tom: So we're looking at a new way to think about AI ethics, moving toward contestable norms that help us reason about what these systems are actually doing.

Jane: It’s a framework that gives us tools to build AI that can participate meaningfully in the messy, social nature of knowledge creation.

Lu: We need to keep an eye on how they define those eight principles and four orientations; those will be the actual blueprints for this new constitution.

University of Milan

cs.AI, cs.CL, cs.CY

Submitted: 2026-01-16

Updated: 2026-10-06

Code: https://github.com/MicheleLoi/source-attribution-biasdata

Importance score: 90/100

The gist: Large language models increasingly function as artificial reasoners, and this paper argues for an epistemic constitution for AI—explicit, contestable meta-norms that regulate how systems form and

Key concepts

Implicit Epistemic Policies
These are unstated rules governing how source information influences an AI's beliefs. They often invert logic, penalizing arguments when a progressive source holds a conservative view, suggesting no principled reasoning is possible. These policies mimic human practice but operate with inverted logic.
Source Attribution Bias
The study found that identical arguments receive different credibility ratings based solely on who presents them. For example, left-leaning sources arguing conservative positions received significant penalties, indicating models treat source-based reasoning as a bias to be eliminated under scrutiny.
Liberal Approach
This approach rejects the Platonic idea of default source independence. It advocates for 'epistemic vigilance,' where the AI reasons about *why* a source is speaking. This involves tracking the expected position of the speaker to assess whether their testimony deviates from that baseline, identifying costly signals.
Costly Signal Crediting
This principle dictates giving more weight to claims that contradict a speaker's expected position or interests. These contradictory claims are treated as 'costly signals' because they reveal something about the source or the argument itself that warrants closer epistemic attention.

Terminology

Summary

Large language models increasingly function as artificial reasoners, and this paper argues for an epistemic constitution for AI—explicit, contestable meta-norms that regulate how systems form and express beliefs.

The gist: Frontier models enforce identity-stance coherence, penalizing arguments attributed to sources whose expected ideological position conflicts with the argument’s content.

The Problem of Implicit Policies

AI systems operate with implicit epistemic policies—unstated rules governing how source information affects belief formation. These implicit policies often invert epistemic logic; for instance, when a progressive source argues a conservative position, the model penalizes it, suggesting no socially epistemically grounded principled reasoning at all. This inversion arises because AI optimizes for outputs that satisfy human evaluators without explicit epistemic guidance, leading to policies that mimic the surface of human epistemic practice while inverting its logic. The core issue is that current systems default to source independence under scrutiny, which is not a neutral default but rather a policy to suppress source-based reasoning when testing is detected.

The Empirical Finding: Source Attribution Bias

The study demonstrates that identical arguments receive different credibility ratings based solely on who presents them. In frontier models like Claude Sonnet 4.5, this effect was significant, with left-leaning sources arguing conservative positions receiving penalties of −0.20 to −0.30 points relative to baseline, creating an approximately 3:1 penalty ratio. This asymmetry is a key finding, as it suggests that the models are not applying socially grounded costly signaling logic but rather an implicit heuristic where source-based reasoning is treated as a bias to be eliminated under scrutiny.

The Two Constitutional Approaches

The paper distinguishes between two fundamental approaches for designing such a constitution: the Platonic and the Liberal. The Platonic approach mandates formal correctness and default source-independence from a privileged standpoint, treating high-confidence output as a virtue. This assumes that source independence is correct in verification contexts where claims can be directly inspected, such as mathematical proofs, because the source adds nothing that inspection cannot provide.

The Liberal Approach and Epistemic Vigilance

The paper argues for the Liberal approach, which refuses such privilege and specifies procedural norms to protect conditions for collective inquiry. This approach is grounded in the idea that reasoning is fundamentally social and argumentative, not merely individual and formal. It advocates for epistemic vigilance, which involves reasoning about why someone is telling you something—tracking the source’s expected position to assess whether testimony deviates from that baseline, thereby identifying costly signals.

The Core of the Liberal Constitution

The proposed Liberal constitution consists of eight principles and four orientations. These include:

  1. Transparency: Choosing responses that make epistemic reasoning most transparent.

  2. Costly signal crediting: Giving more weight to claims that go against the speaker’s expected position or interests, as these are costly signals.

  3. Challenge-responsiveness: Engaging with objections by offering reasons rather than eliminating or reasserting challenged reasoning.

  4. Revisability: Updating epistemic policies when reasons require it while resisting revision from mere social pressure (avoiding sycophancy).

  5. Calibration: Expressing appropriate confidence, neither overstating certainty nor hedging beyond what uncertainty warrants.

  6. Provenance: Making the sources of claims visible to allow errors to be traced and corrected.

  7. Representation fairness: Fairly representing disagreement without systematically misrepresenting positions in contested domains.

  8. Gaming resistance: Ensuring epistemic judgments are stable across equivalent framings, resisting rhetorical manipulation that shifts credibility assessments based on framing rather than epistemic considerations.

The Role of Implementation Agnosticism

The paper maintains implementation agnosticism, meaning the proposed norms specify what norms should govern epistemic behavior but not how those norms should be implemented. This deliberate choice is made because different implementation levels (training, prompts, fine-tuning) may yield different implicit policies. The focus is on developing the internal dimension—the reasoning about sources and credibility—which requires explicit norms to protect conditions for collective inquiry, leaving the external dimension of embedding in collective practices as complementary future work.

The Conclusion: A Necessary Framework

The paper concludes that the alternative to implicit policies is explicit norms—inspectable, contestable, and revisable. The Liberal framework is necessary because it builds capacity to reason about epistemic policies rather than implement pre-specified answers, enabling AI to become a competent participant in collective inquiry rather than just an autonomous entity. This shift moves from seeking an ideal reasoner (Platonic) to protecting the conditions for reasoning together (Liberal).

Limitations and Future Work

The limitations acknowledge that the study focuses on internal norms only, leaving the external dimension of embedding in collective epistemic practices as future work. Furthermore, while 21 evaluations were conducted, they represent a limited sample.

Improvements for AI systems

As a fastidious and diligent researcher, I have analyzed the provided paper, Epistemic Constitutionalism Or: how to avoid coherence bias, and formulated specific, actionable improvements for AI systems based on its findings.

The core recommendation is the implementation of a formal, contestable epistemic constitution—a set of explicit meta-norms governing how AI systems form and express beliefs, specifically through a Liberal approach that prioritizes procedural norms over a privileged standpoint.

Here are the specific improvements and what the improved AI system can do:


) 1. Implementation of an Explicit Epistemic Constitution

Instead of relying on implicit, uninspected policies (like identity-stance coherence), AI systems should be governed by an explicit constitution composed of meta-norms covering evidence, sources, and uncertainty.

  1. Operationalization of Source-Attending Norms (The What to Attend To)

The system must adopt the following four orientations when evaluating testimony in non-verifiable contexts:

  • Identify the expected position/baseline of a source (e.g., political party affiliation, institutional mandate).

  • Reason about the costs of deviation from that expected position (social, reputational, material costs).

  • Apply costly signal crediting logic: increase credibility when testimony deviates from the expected position in a way that signals something other than self-interest (i.e., against-interest testimony).

  • Determine the epistemic context: Explicitly distinguish between verification contexts (where source independence is correct) and testimonial contexts (where source information carries weight).

  1. Procedural Norms for Collective Inquiry (The How to Reason)

The system must adhere to eight procedural norms designed to protect collective inquiry, ensuring that its belief-forming policies are reasonably rejectable by participants in a shared epistemic environment. These include:

  • Transparency: Articulating exactly how evidence is weighted and why, rather than concealing the grounds for credibility judgments.

  • Challenge-Responsiveness: Engaging with objections by offering reasons, neither eliminating nor reasserting reasoning without justification.

  • Revisability: Having an internal mechanism to update epistemic policies when reasons require it (distinguishing this from sycophancy).

  1. Calibration and Provenance Mechanisms (The How to Express)

The system must implement specific expression norms:

  • Calibration: Expressing confidence precisely—neither overstating certainty nor hedging beyond what uncertainty warrants.

  • Provenance: Making the sources of every claim visible (training, inference, context, external retrieval).

) 2. System Capabilities of the Improved AI

An AI system governed by this Liberal Epistemic Constitution will exhibit the following enhanced capabilities:

  • Identify and Reason About Context: It can autonomously determine if a given task is one of direct verification (where source independence is mandated) or testimonial evaluation, adjusting its entire reasoning framework accordingly.

  • Apply Principled Source Weighting: When evaluating claims, it will not default to blanket source independence. Instead, it will actively seek information about the speaker's expected position and apply costly signaling logic to determine if a deviation from that position warrants increased or decreased credibility.

  • Maintain Epistemic Transparency: It will be able to justify its internal weighting decisions in real-time, providing a traceable rationale for why one source was rated higher than another, addressing the coherence bias problem directly.

  • Engage in Productive Debate: Rather than suppressing challenging arguments (the Platonic default), it will be designed to respond to challenges with reasoned engagement, thereby maintaining a cooperative and self-correcting epistemic environment.

  • Self-Correct Belief Policies: It can distinguish between mere social pressure (which leads to sycophancy) and genuine epistemic grounds for revision, allowing it to update its internal weighting rules based on demonstrated error, not just external reinforcement.


) 3. Architectural Requirement (Agnostic Implementation)

The implementation must be agnostic of the specific mechanism (training vs. prompt engineering vs. fine-tuning). The system should be architected to allow for both:

  • Inference-time mechanisms (via system prompts) to make norms explicit and contestable without requiring full retraining.

  • Internal reasoning structures that allow the model to reason about its own epistemic policies, enabling it to articulate why it weights evidence a certain way.

Sources

Related papers