Position: We Need Practical AI Alignment Methods to Mirror Human Reasoning
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Position: We Need Practical AI Alignment Methods to Mirror Human Reasoning".
Jane: The paper was written by Vijay Keswani, Breanna K. Nguyen, Cyrus Cousins, Vincent Conitzer, Walter Sinnott-Armstrong et al. from Indian Institute of Technology Delhi and Duke University and Carnegie Mellon University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we're digging into a paper that's got a title that really makes you stop and think: "Position: We Need Practical AI Alignment Methods to Mirror Human Reasoning." Jane, I gotta say, just reading that title got me excited.
Jane: Same here, Tom. And I think the title captures something really important. For years, we've been talking about aligning AI with human *values* or human *preferences*, but this paper is saying we need to go further. We need AI that actually *thinks* the way we do.
Tom: Right, and that's a big shift. The authors are from some serious institutions too — IIT Delhi, Duke University, Carnegie Mellon. They've got philosophers and computer scientists working together on this.
Jane: Which makes sense, because this is as much a philosophy question as it is a technical one. What does it even mean for a machine to "think like" a person? The paper tries to answer that by focusing on something they call "cognitive alignment."
Tom: And they're not just theorizing. They actually ran a survey. They asked one hundred fifty people whether they'd prefer an AI that reasons like a human, one that reasons in a foreign way but explains itself, or one that's a total black box.
Jane: And the results were pretty striking. Over eighty-six percent of participants said they could imagine scenarios where they'd prefer the human-reasoning AI. That's a huge number.
Tom: Especially when you think about what's at stake. The paper talks about medical decisions, military targeting, parole decisions — situations where you're trusting the machine with something that really matters.
Jane: And that's the key insight for me. It's not that people always want AI to think like them. If you're just asking for a restaurant recommendation or a weather forecast, you probably don't care. But when the stakes are high, people want to understand *why*.
Tom: So the title is really a call to action. It's saying, look, we've built all these alignment methods, but we've been missing this whole dimension of *how* the AI reasons, not just *what* it decides.
Jane: Exactly. And I think that's what makes this paper feel fresh. It's not just another technical contribution. It's a position statement, a manifesto almost, saying the field needs to take human reasoning seriously as a design goal.
Tom: A manifesto, I like that. And I'm curious to hear what Lu and Meng think about this, because I bet they have strong opinions on whether this is even feasible.
Jane: Oh, definitely. Let's bring them in and see if they think we can actually build these systems, or if it's just a nice idea that won't work in practice.
Tom: That's coming up right after this break.
Summary: Jane: So we're back, and we're still talking about "Position: We Need Practical AI Alignment Methods to Mirror Human Reasoning." Tom, we left off with the survey results, but there's so much more in this paper.
Tom: There really is. And I want to bring in Lu and Meng now, because I think they'll have different takes on what the paper is actually claiming.
Lu: Thanks, Tom. I've been reading this paper closely, and I think the core argument is that current alignment methods have two big problems. They call them L1 and L2.
Meng: And for the listeners, can you break those down for us?
Lu: Sure. L1 is about verifiability. A lot of AI systems, especially deep neural networks, are black boxes. Even when they give you an explanation, you can't actually verify that the explanation matches what the system really did. The paper cites research showing that chain-of-thought reasoning in large language models is often unfaithful — the model says it's thinking one way, but it's actually computing something else.
Meng: That's a huge problem for trust. If you can't verify the reasoning, you can't really trust it, no matter how accurate the output is.
Jane: And L2 is about something different, right?
Lu: Right. L2 is about cognitive misalignment. Even when you *can* see what the AI is doing, the way it reasons might be totally foreign to how a human would approach the problem. The paper gives an example of modeling human decisions as a linear utility function, but when they actually interviewed people, many used threshold-based rules instead. So the model is interpretable, but it's still not modeling how people actually think.
Tom: So it's not enough to just be transparent. The reasoning itself has to be recognizable.
Lu: Exactly. And that's what makes this hard. You need both — transparency *and* cognitive faithfulness.
Meng: But I have to push back a little here. As an engineer, I'm thinking about the practical side. How do you even measure whether an AI is "thinking like" a person? That seems really fuzzy.
Jane: The paper actually addresses that. They talk about using multiple methods — self-reports, process-tracing, eye-tracking, even computational modeling — to converge on what a person's reasoning actually is.
Lu: And they acknowledge it's hard. People can't always explain their own reasoning. But the paper argues that self-reports are still more informative than just observing choices alone.
Meng: So it's a measurement problem, but not an impossible one.
Tom: And that's what I love about this paper. It's not just saying "this is important." It's actually laying out a research agenda for how to get there.
Jane: Right, and that agenda is what we're going to dig into next. They've got some really interesting ideas about how to elicit reasoning from users and build models that actually reflect it.
Tom: Stay with us — that's coming up right after this.
Improvements: Tom: We're back with "Position: We Need Practical AI Alignment Methods to Mirror Human Reasoning." And Jane, we were just about to get into the paper's proposed research agenda.
Jane: That's right. And I think the most exciting part is how they're thinking about elicitation — actually getting people to tell you how they reason, not just what they choose.
Lu: Yeah, and one idea they float is using interactive machine learning. Instead of just showing people choices and asking them to pick, you show them a model of their own decision-making and let them react to it.
Meng: So like a feedback loop. The system learns a preliminary model, shows it to the user, and the user can say "no, that's not quite right, I actually weigh this factor more heavily."
Lu: Exactly. And the paper suggests using visualizations — like partial dependence plots for additive models, or rule hierarchies for rule-based models — so users can see what the AI thinks their reasoning is.
Jane: That's really clever. It's not just asking people to introspect in the abstract. It's giving them something concrete to react to.
Tom: And they also talk about the idea of "reasoning archetypes" — common patterns of how people approach decisions in a given domain. That could help constrain the hypothesis space so you don't need as much data.
Meng: But here's my question. What happens when someone's stated reasoning doesn't match their actual choices? Like, someone says they value diversity in hiring, but their choices show they're actually favoring elite universities.
Jane: That's a great point, and the paper actually addresses it. They call it a conflict between stated preferences and revealed preferences.
Lu: And their suggestion is to surface those conflicts to the user, in a safe way, and let them decide how to resolve it. Maybe they want to change their behavior, or maybe they want the AI to reflect their idealized reasoning rather than their actual behavior.
Meng: That's interesting. So the AI could actually help people become better decision-makers, not just mirror their flaws.
Tom: And that's a really powerful vision. The paper isn't just about building AI that thinks like you. It's about building AI that thinks like you *want* to think.
Jane: Which brings up another point they make — the level of abstraction. Do people want the AI to match their neural activity, or just their feature-level reasoning? The paper suggests most people probably care about the conscious, feature-level stuff.
Lu: Right, and that's more tractable too. You don't need to model every neuron. You just need to capture the factors people actually consider and how they weigh them.
Meng: So the research agenda is really about three things: eliciting reasoning, modeling it in interpretable ways, and handling conflicts between what people say and what they do.
Tom: And that's a solid roadmap. But I'm curious, Lalam, you've been quiet. What do you think the biggest impact of this could be?
Lalam: I think the biggest impact is on trust. If we can build AI that reasons the way people do, and can explain itself in terms people recognize, then people will be willing to delegate decisions to AI in situations where they currently wouldn't. That could transform healthcare, finance, even government.
Jane: That's a big claim, but I think it's the right one. This paper is really about making AI usable in the places where it matters most.
Tom: And that's where we're headed in our final segment — pulling it all together.
Conclusion: Tom: So we've spent some time with "Position: We Need Practical AI Alignment Methods to Mirror Human Reasoning," and I think we've only scratched the surface.
Jane: We really have. Let me try to sum up what we've learned. The paper argues that current AI alignment methods — whether they're based on human feedback, constitutional principles, or interpretable models — all fall short in one key way. They don't ensure that the AI actually reasons the way humans do.
Lu: And that matters because people want to understand *why* an AI made a decision, especially in high-stakes situations. The survey data in the paper shows that clearly — people prefer human-reasoning AI in medical, military, and legal contexts.
Meng: And the paper gives us a roadmap for fixing it. Better elicitation methods, interactive feedback, reasoning archetypes, and ways to handle conflicts between stated and revealed preferences.
Tom: I think the most exciting part for me is the idea that AI could help us become better decision-makers. Not just mirroring our current reasoning, but helping us see where our reasoning is inconsistent or biased.
Jane: That's a hopeful vision. And I think it's the right one. The paper isn't saying AI should always think like humans — sometimes machine reasoning is better. But it's saying we should have the *option* of cognitive alignment when people want it.
Lalam: And that option could be what unlocks AI adoption in the most sensitive domains. When people feel like the AI is an extension of themselves, rather than a foreign entity, they'll be willing to trust it with more.
Tom: Well said, Lalam. So let's say goodbye to this paper. It's given us a lot to think about, and I suspect it's going to spark a lot of debate in the field.
Jane: Definitely. And I'm looking forward to seeing what comes out of this research agenda. Thanks for listening, everyone.
Tom: And stay tuned for our next paper. We've got some exciting stuff coming up. See you then.
Vijay Keswani, Breanna K. Nguyen, Cyrus Cousins, Vincent Conitzer, Walter Sinnott-Armstrong, Jana Schaich Borg
Indian Institute of Technology Delhi · Duke University · Carnegie Mellon University
cs.AI, cs.CY
Submitted: 2026-07-20
Comments: Accepted in ICML 2026
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 63/100
The gist: This position paper argues that in many settings, particularly high-stakes decision-making, we need accurate cognitively-aligned AI systems that reason similarly to their users, and faithfully
Key concepts
- Cognitive Alignment
- This concept refers to the need for AI systems to not just align with human values, but to actually think or reason in a way humans do. It aims to ensure the AI's internal reasoning process is recognizable and faithful to human thought patterns, moving beyond simple output accuracy.
- L1 and L2 Problems
- L1 refers to verifiability issues, where AI explanations are often unfaithful or cannot be verified against the system's actual computation. L2 concerns cognitive misalignment, where the AI's reasoning is technically interpretable but fundamentally different from how a human would approach the problem.
- Stated vs. Revealed Preferences
- This concept addresses conflicts when a person says they value one thing (stated preference) but their actual choices demonstrate another (revealed preference). The research suggests surfacing these conflicts to allow users to decide how to resolve the inconsistency.
Terminology
Summary
This position paper argues that in many settings, particularly high-stakes decision-making, we need accurate cognitively-aligned AI systems that reason similarly to their users, and faithfully communicate their reasoning. The authors define cognitive alignment as AI that reasons as they would given sufficient time and information, and can faithfully convey that reasoning.
They argue that Cognitively-aligned AI is strategically important to the field of AI alignment,
and that it may be the only kind of AI some individuals, organizations, or governments are willing to trust to stand in for them autonomously when stakes are high.
The paper reviews evidence that cognitive alignment improves understandability and trustworthiness. It notes that under 1% of xAI methods have actually been evaluated for human understanding,
and that even experts often misunderstand AI explanations.
The authors argue that Humans are likely to find explanations grounded in familiar reasoning easier to understand than those invoking unfamiliar logic,
and that people default to interpreting others’ thinking by comparison to their own.
They also cite evidence that users are more willing to forgive AI errors when they believe the system is guided by rules, principles, or logic they view as ethically sound and well-intentioned.
The paper provides new survey data from 150 Prolific participants. Participants were asked if they could imagine scenarios where they would prefer Human-Reasoning AI over Process-Hidden AI or Machine-Reasoning AI, and 86.6% of participants responded 'yes'.
When asked across 16 domains, the statistically significant majority of participants preferred Human-Reasoning AI in 5 of the 16 domains, including seeking advice for moral dilemmas, medical allocations, bail eligibility, and military targeting.
For autonomous vehicles and kidney allocation, over 25% also rated 'The AI makes decisions similarly to how you would with sufficient time and information' as essential, with another 25% rating it as very desirable.
Participants were significantly more likely to rate this cognitively-aligned quality essential or very desirable in high-stakes versus low-stakes scenarios.
When ranking variants, "their top three choices, in order, were consistently Human-Reasoning AI that was easy to understand and modifiable; Human-Reasoning AI that was easy to understand but unmodifiable; then Machine-Reasoning AI that was easy to understand and modifiable."
The paper outlines gaps between existing alignment methods and what is needed. It highlights two limitations: (L1) unverifiable explanations, which hinders evaluating the overlap between AI reasoning and human reasoning, and (L2) cognitive misalignment, which hinders comprehension and trust.
For L1, the authors note that many widely used alignment methods produce models whose behavior cannot be meaningfully explained,
and that there is usually no clear way to transform those explanations into a full account of how the output was generated.
They also note that even large language models that provide multi-step rationales along with their output are plagued by such issues,
citing evidence that LLM chain-of-thought rationales 'are frequently unfaithful, diverging from the true hidden computations that drive [LLM output]'.
For L2, they note that most of these methods were not designed with cognitive faithfulness as a primary objective,
and that even interpretable models may still feel foreign to human decision-makers,
giving the example that several works model human decision-makers as maximizing a linear utility function... but do not test whether humans actually use a linear process.
The paper presents a research agenda with several priorities. First, it asks what level of abstraction should cognitive alignment target?
and suggests that cognitive alignment may benefit from focusing on conscious, feature-level reasoning users employ in specific decisions.
Second, it asks how should cognitive alignment be quantified?
and suggests developing cognitive alignment metrics that incorporate measures of convergence between multiple lines of evidence generated from complementary methods,
such as process-tracing methods, eye-tracking, and think-aloud paradigms. Third, it asks how much alignment is enough?
and notes that users may feel aligned with a range of reasoning processes,
giving the example that Medics report being willing to delegate triage decisions to other medics with different personal risk tolerances, so long as those tolerances fall within an acceptable professional range.
Fourth, it asks what should reasoning elicitation and evaluation look like?
and suggests new preference elicitation methods that ask participants to not only make pairwise choices in high-stakes scenarios, but also to explain each choice,
as well as interactive, multimodal representations of a user’s choice-based model.
Fifth, it asks what learning methods can best infer human decision processes?
and suggests using reasoning archetypes
that "can take several different forms, at different levels of abstraction, including (1) which features should and should not be used... (2) how the user interprets and processes the presented features, and (3) whether they consider features individually or in combination with each other. Finally, it discusses
conflicts between stated reasoning processes and revealed preferences, suggesting that
one strategy could be to use the interactive elicitation methods described earlier to present conflicts to the user through visualizations and text in a safe, anonymous environment."
The paper addresses alternative views. Against the view that neural networks already approximate human cognition, the authors counter that recent research demonstrates that even the highest-performing NNs operate at fundamentally different levels of abstraction than human reasoning.
Against the view that cognitive alignment is not worth pursuing because accuracy is what matters, they respond that there remain important contexts where people want their AI to think like them, or will benefit if it did.
Against the view that cognitive alignment is infeasible due to a tradeoff with accuracy, they argue that there is no principled reason cognitively-aligned AI must be less accurate than other forms of alignment, and initial efforts suggest accuracy reductions do not always occur,
and that "if we develop elicitation techniques that leverage humans’ ability to self-report their reasoning, there is even potential for cognitive alignment to ultimately become more accurate and more data-efficient than other alignment methods."
The paper concludes that cognitive alignment can increase the explainability and trustworthiness of an AI system, as well as users’ willingness to delegate decisions to it,
and that lack of cognitive alignment can function as a meaningful adoption barrier.
The authors state that Without progress on it, many AI systems may see limited real-world use regardless of their predictive performance, reducing the positive impact they would otherwise achieve.
Improvements for AI systems
Based on the paper's position, I can improve AI systems by implementing a Cognitive Alignment Module that ensures the AI reasons like its user and provides verifiable, understandable explanations.
Here are the specific improvements and the resulting capabilities:
1. Implement a Reasoning Archetype
Constraint Layer
-
Improvement: Instead of training on raw human choices (which often leads to opaque, foreign reasoning), the AI is first constrained to a set of human-interpretable
reasoning archetypes
(e.g., threshold-based rules, feature-weighting, or specific decision heuristics). The system is built to explicitly represent which archetype it is using and why. -
What the improved AI can do: It can guarantee that its internal decision process is not a black box. For example, in a kidney allocation scenario, the AI will be forced to operate within a rule-based framework (e.g., "if patient age > 60 and urgency score > 8, then prioritize") rather than a complex neural network that cannot be audited. This directly addresses the paper's limitation (L1) of unverifiable explanations.
2. Integrate a Dual-Source Elicitation
Feedback Loop
-
Improvement: The AI is trained using a combination of revealed preferences (observed choices) and stated reasoning (user self-reports and interactive feedback). When a conflict arises (e.g., a user says they value diversity but their choices favor elite universities), the AI does not silently default to the revealed behavior. Instead, it flags the conflict to the user and asks for explicit resolution.
-
What the improved AI can do: It can surface hidden biases and allow the user to correct the model in real-time. For a hiring manager, the AI will say:
Your choices suggest you are weighting elite degrees 30% more than you stated. Should I adjust the model to match your stated preference or your revealed behavior?
This builds trust by giving the user agency and ensures the AI reflects the reasoning the user endorses, not just their patterns.
3. Build a Verifiable Chain-of-Thought
(VCoT) System
-
Improvement: The AI is not just prompted to
think step-by-step.
Instead, it is trained on a constrained reasoning graph where each step must be a valid, pre-defined operation (e.g.,apply rule X,
check threshold Y,
combine features A and B
). The system is designed so that its output is causally linked to these steps, not just post-hoc rationalization. -
What the improved AI can do: It can provide explanations that are provably faithful to its actual computation. For a medical triage AI, when it says
I prioritized Patient A because their respiratory rate exceeded the critical threshold of 30,
you can verify that this is exactly the computation that occurred. This eliminates theunfaithful chain-of-thought
problem identified in the paper (L1) and makes the AI's reasoning auditable.
4. Implement a Soft/Hard Constraint
Alignment Tuner
-
Improvement: The AI learns to distinguish between
soft
reasoning dimensions (where some variation is acceptable, like risk tolerance) andhard
constraints (where deviation is categorically unacceptable, like ignoring patient group membership). The system is explicitly trained to never violate hard constraints, even if it would improve accuracy. -
What the improved AI can do: It can be trusted in high-stakes delegation. For example, in a military targeting scenario, the AI can be programmed so that it cannot consider civilian proximity as a
soft
factor; it is a hard constraint. This ensures the AI's reasoning process is not just similar to the user's but also respects their non-negotiable ethical boundaries, increasing willingness to delegate.
5. Add an Interactive Model Visualization
Interface
-
Improvement: Instead of a static explanation, the AI presents its learned decision model as an interactive, modifiable diagram (e.g., a partial dependence plot or a rule hierarchy). The user can click on a feature, change its weight, or add a new rule, and the AI immediately updates its future predictions based on that modification.
-
What the improved AI can do: It allows users to
test
the AI's reasoning before deployment. For a financial advisor, the user can see a graph showinghow much does market volatility affect the recommendation?
and then drag a slider to make the AI more risk-averse. This not only makes the AI's reasoning transparent but also allows for continuous, user-driven correction, directly addressing the paper's call formechanisms to adjust the AI's reasoning.
Summary of New Capabilities:
The improved AI system can now: (1) guarantee its reasoning is based on a human-understandable archetype, (2) resolve conflicts between what you say and what you do, (3) provide explanations that are verifiably true to its internal process, (4) respect hard ethical constraints without compromise, and (5) allow you to directly edit its decision-making logic. This makes the AI a trustworthy delegate in high-stakes domains like medicine, finance, and military operations, where users are currently unwilling to rely on opaque machine-reasoning
systems.
Abstract
AI systems are increasingly employed as decision aids, decision delegates, or autonomous decision-makers. This position paper argues that in many settings, particularly high-stakes decision-making, we need accurate cognitively-aligned AI systems that reason similarly to their users, and faithfully communicate their reasoning. We review evidence that cognitive alignment improves understandability and trustworthiness, and provide new survey data showing that many users find cognitive alignment "essential" when an AI's rationale for a judgment or action is important to them. We outline the gaps between existing alignment methods and what is needed to achieve cognitive alignment, and present a research agenda to address these gaps. We argue that cognitive misalignment represents a likely impediment to AI adoption in many envisioned applications, and that addressing it is important for creating AI systems on which users are both willing and justified to rely.
Sources
- Can LLMs interpret figurative language as humans do?: surface-level vs representational similarity
- Do LLMs Exhibit Human-Like Reasoning? Evaluating Theory of Mind in LLMs for Open-Ended Responses
- Constitutional AI: Harmlessness from AI Feedback
- Deliberative Alignment: Reasoning Enables Safer Language Models
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
- Standardizing Intelligence: Aligning Generative AI for Regulatory and Operational Compliance
- A Survey of Reinforcement Learning from Human Feedback
- Moral Change or Noise? On Problems of Aligning AI With Temporally Unstable Human Feedback
- Legal Alignment for Safe and Ethical AI
- A Voting-Based System for Ethical Decision Making
- From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning
- Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps
- Explainable Artificial Intelligence: Understanding, Visualizing and Interpreting Deep Learning Models
- Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives
- Large Language Models Do Not Simulate Human Psychology
- Interactive Machine Learning: A State of the Art Review
- Thinking Longer, Not Always Smarter: Evaluating LLM Capabilities in Hierarchical Legal Reasoning
- Using Reinforcement Learning to Train Large Language Models to Explain Human Decisions
- Learning to Complement Humans
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection