Process-Constituted Intelligence: A Shared Criterion for Humans and Machines

arXiv:2608.16213 · cs.AI, cs.ET · Submitted 2026-08-17 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Process-Constituted Intelligence: A Shared Criterion for Humans and Machines".

Jane: The paper was written by Shen, J. H. and Tamkin, A. from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Summary: Jane: So, following up on the idea that intelligence is process-based, the paper goes into detail about how we should actually summarize this concept. It's not enough to just say "process matters"; they have to show us *why* and *how*.

Tom: And what I took away from reading the summary section is that they are defining a much broader scope for what counts as intelligence, going way beyond simple pattern matching or recall.

Lu: The paper argues that effective intelligence requires managing uncertainty—that ability to adjust the plan when the initial assumptions fall apart, rather than just following a pre-programmed path.

Meng: I found that really interesting because most current AI models are really good at assuming the input is clean and consistent; they struggle when the world throws in unexpected noise or conflicting data.

Jane: It suggests that genuine intelligence involves recognizing when you don't know something, and then having a structured way to go about finding out, rather than just guessing confidently.

Lalam: What I appreciate about this summary is that it grounds the theory in human experience; it talks about metacognition—the ability to think about your own thinking—which is crucial for any truly adaptive system.

Tom: Right, so it’s not just *what* we output, but how often we pause, self-correct, and adjust our internal state along the way.

Meng: If the summary is right that failure handling is part of intelligence, then we need to build mechanisms that don't just crash; they need to report *why* they failed and what assumptions were wrong.

Lu: That touches on the concept of 'epistemic uncertainty,' where the system isn't just lacking data, but it suspects its own knowledge base might be flawed or incomplete.

Jane: It’s about having internal checks and balances, knowing your limits, which is something we rarely see in consumer-facing AI products right now.

Lalam: From a cultural perspective, this encourages us to view AI not as a perfect oracle, but as an incredibly powerful cognitive partner that understands the value of doubt.

Tom: So we're moving from systems that just deliver answers to systems that guide us through the process of finding the best answer.

Improvements: Jane: Okay, so if the summary tells us *what* process-based intelligence is, this next section, where they suggest improvements, tells us *how* we can actually build it. This is where it gets really actionable for developers like Meng's team.

Tom: And I think the key improvement they push is moving away from monolithic models that just ingest a massive amount of data and spit out a single answer—it has to be iterative.

Meng: The paper suggests developing agents that maintain internal state across multiple turns, meaning their current output is directly constrained by what they were doing five steps ago, not just the immediate context window.

Lu: That capability of maintaining state and revising plans based on environmental feedback is exactly what we need to mimic high-level human planning—it's not a single shot; it’s an evolving project.

Jane: It’s like teaching a child how to build with complicated blocks; you can't just show them the finished castle, you have to let them fail, take it apart, and rebuild using the lessons learned from the collapse.

Lalam: What this implies for society is that we need to reward complex, iterative problem-solving in education and work, not just final grades or deliverables.

Tom: So, instead of optimizing for a single perfect result, we're optimizing for the *robustness* of the process itself—the resilience.

Meng: If I had to build this today, I

Paper discussion segment 3: Tom: So, if I'm getting this right, the biggest shift this paper advocates is that we stop grading AI solely on its final answer and start grading it on *how* it gets there.

Jane: Exactly, Tom; instead of just looking at a correct conclusion, they're asking us to look at the whole journey—the revisions, the moments of doubt, everything that happens in between.

Lu: That’s incredible because right now we’re treating AI like a magic box that just spits out perfection; this paper suggests we need to build systems that actually show their work, like a student's notebook filled with crossed-out ideas.

Meng: But Lu, showing the work sounds messy for an engineer; how do you even build a quantifiable metric for 'showing your work' if the process involves backtracking and rethinking?

Lalam: I think we have to view that messiness as valuable data itself; if AI can prove it struggled with a problem and then overcame that struggle, we’re not just improving computation, we’re modeling resilience for human society.

Jane: Resilience is such a good word for it; so basically, it's about building machines that aren't afraid to say, "I don't know," and then showing us the steps they took to get closer to knowing.

Tom: Right, Jane hits on something key there; it changes the entire relationship between the user and the AI—it becomes a dialogue partner rather than a black box answer generator.

Lu: And imagine applying that dialogic approach across fields, not just coding, but in art or complex scientific theory where failure is just part of exploration!

Meng: Speaking of exploration, if we mandate that process tracking, we're talking about massive overhead; the computational cost of recording every failed thought must be considered for any real-world deployment.

Lalam: But Meng, if that recorded struggle allows us to solve problems previously deemed impossible—problems only solvable by mimicking human fallibility—then the cost becomes negligible compared to the value.

Jane: So it's a trade-off between efficiency and depth, and this paper is arguing strongly for depth right now.

Tom: It seems like this means the future of AI isn't just about being faster, but about being transparently thoughtful; we need to dig into how these process improvements actually change our jobs.

Conclusion: Tom: Wow, so if I’m getting this right, the big takeaway isn't just about making AI smarter in terms of raw knowledge, but fundamentally changing what we think intelligence even means for both us and these machines.

Jane: Exactly! It’s less about reaching a perfect answer and more about showing the messy journey you took to get there—the mistakes, the adjustments, the 'I'm not sure yet' moments.

Meng: That concept of process being paramount really changes how we have to structure our testing pipelines, doesn't it? We can’t just grade a final model output anymore; we have to audit the entire decision tree.

Lu: And thinking about that on a massive scale, I mean, if every system we build forces itself through this kind of self-correction and struggle with ambiguity, the resulting AI might actually become something much more robust than anything trained just on clean data.

Jane: Robust is the word; it sounds like giving the AI a chance to grapple with things that haven't been solved yet, which is really what human learning does naturally.

Lalam: What I find most powerful about this perspective, though, is how it changes our culture around failure; instead of seeing errors as endpoints, we learn to see them as necessary inputs for growth and deeper understanding.

Tom: So we’re moving away from the idea of a solved puzzle and toward the idea of an ongoing, highly complex conversation with the medium itself.

Meng: From an implementation standpoint, that requires building agents that are designed to be perpetually unsatisfied, always looking for the next piece of contradictory evidence to make them rethink everything.

Lu: It’s a beautiful framework because it gives us a shared vocabulary—a way to critique both human thinking and machine thinking against the same set of criteria.

Jane: It really does feel like a complete shift in focus, emphasizing that the intelligence isn't housed in one place, but is built through continuous practice and challenging assumptions.

Lalam: It speaks to a deep human need for meaning-making, which process-substituted intelligence captures perfectly—it shows the work.

Tom: Honestly, this concept of *Process-Constituted Intelligence: A Shared Criterion for Humans and Machines* gives us such a powerful lens through which to view all future AI development.

Jane: We certainly have a lot to digest from this one, Tom; it really sets the bar high for what we should expect next.

Tom: Well, team, that’s it for this deep dive into process intelligence—thank you so much to Lu, Meng, and Lalam for weighing in on this fascinating stuff.

Meng: Thanks to everyone; I'm already thinking about how we can build a sandbox environment just to stress-test these process requirements.

Lu: Keep those challenging scenarios coming; I've got ideas for how this applies to theoretical physics simulations alone!

Lalam: And remember, the most advanced intelligence always serves a purpose that elevates our shared human experience.

Jane: Okay, listeners, we gotta take a quick break, but when we come back, we’re shifting gears entirely and looking at something really different: advancements in personalized medicine using multimodal data...

cs.AI, cs.ET

Submitted: 2026-08-17

Updated: 2026-09-07

Comments: v2: revised throughout for clarity and consistency; supplementary cross-references corrected

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 79/100

The gist: The paper introduces "process-constituted intelligence," establishing a comprehensive and shared criterion for auditing both human and machine cognition.

Key concepts

Process-Constituted Intelligence
A framework suggesting that intelligence is defined by the method used to solve a problem (the process), rather than solely by the final correct output. It requires showing revisions, doubts, and adjustments.
Managing Uncertainty
The ability of an intelligent system to adjust its plan when initial assumptions fail or when encountering conflicting data. This involves recognizing when it doesn't know something and having a structured way to find out.

Terminology

Summary

The paper introduces process-constituted intelligence, establishing a comprehensive and shared criterion for auditing both human and machine cognition. Rather than merely evaluating a final output, this framework mandates an examination of the underlying cognitive process to distinguish genuine intellectual work from mere resemblance or polished performance. This diagnostic approach provides operational evidence—a diagnostic signature—that maps specific behaviors onto measurable criteria, allowing auditors to judge whether intelligence is truly constituted through process or simply simulated.

The Mechanics of Iteration and State Maintenance

Cognitive processes are defined by how they manage attempts and time. The first feature, Generative trial and revision, posits that the discarded attempts are the substance of the work, not waste preliminary to it. This is observable in human examples like an inventor’s failed prototypes or in AI/machine instantiations such as Multi-sample rollouts, tree-of-thought branching, rejection sampling with selfcritique. Crucially, auditors must look for a traceable revision history, not just a polished result, asking if failed candidates are actually generated and shown to shape the final output.

Secondly, Temporal extension dictates that cognition unfolds over time and across repeated returns to the same material, accumulating state rather than resolving in one bounded step. This is exemplified by persistent agentic trajectories that carry state across sessions; memory that accrues across interactions rather than resetting each prompt. The diagnostic signature here is whether earlier work measurably constrain later work, with state maintained and revisited across turns or is each response stateless and bounded by a single context window?

Handling Ambiguity and Environmental Feedback

A core aspect of intelligence is the ability to manage uncertainty. Engagement with uncertainty requires a solver to sit in not-knowing and refuse premature closure. The work becomes a conversation with the medium rather than execution of a pre-formed plan. Key diagnostic questions include probing with ill-posed prompts to measure calibration, and determining if the system flags or abstains, or if it produces a fluent, confident answer regardless.

This process is further refined by Feedback with the medium, which demands that environmental return changes not just the surface output, but genuinely alters the plan and framing. The system must be tested to distinguish genuine reframing from re-running the same plan against feedback. Similarly, Value-laden framing requires assessing if the system generate or revise its own framing of what matters, or optimize a framing handed to it, indicating an ability to redefine the goal itself.

Contextual and Developmental Intelligence

Intelligence is also constituted by community and growth. Social and dialogical accountability involves practices such as Socratic dialectic... and peer review, where claims are tested through challenge. The diagnostic test here is whether adversarial critique change the framing and the outcome or is the 'dialogue' cosmetic, requiring that challenge measurably alters results.

Finally, the Formative dimension highlights that much of this process runs below explicit articulation and shapes who the practitioner is becoming. This capacity appears through human–AI co-formation, where the system shapes, and is shaped within, a practice rather than internalizing one itself. The overall assessment must therefore look for the development of tacit, practicegrounded capacity over time, which defines true process-constituted intelligence.

Improvements for AI systems

The current limitations in LLMs stem from treating cognition as a single, stateless forward pass rather than an iterative, embodied process. The following improvements shift the focus from output generation to process modeling.


Improvement: Replace the standard auto-regressive decoding mechanism with a Multi-Pass Search Tree Architecture (MSTA) coupled with an explicit, retrievable 'Failure Corpus.' The system must be architecturally mandated to generate and store N>1 candidate pathways (Path 1, Path 2,) for any complex query. These failed or suboptimal paths are not discarded but are indexed alongside the reasoning steps that led to their rejection (the failure signature).

What the Improved AI Can Do: It moves beyond presenting a single polished result. When answering a question, it can articulate why several other plausible approaches were attempted and subsequently discarded, citing specific points of failure (e.g., Initial Hypothesis A failed because it contradicted Principle X; therefore, we revised the approach to B). This allows for verifiable transparency into the search space.

Improvement: Implement a Graph-Based Contextual Memory Layer (GCML) that transcends the fixed context window limit. The GCML must not merely store tokens, but must map interactions as weighted nodes within a dynamic graph, where edges represent state dependencies and cumulative knowledge acquisition. Crucially, the system must maintain agentic trajectories—a persistent understanding of its own accumulated state across sessions—that can be recalled and actively constrained during new prompts.

What the Improved AI Can Do: It achieves true long-term memory integration. If a user provides data or sets constraints over several weeks, the AI will not forget them; instead, it will use the entire history as a continuous constraint set, preventing contradictory or redundant work and allowing for cumulative skill building across extended projects.

Improvement: Introduce a mandatory Uncertainty Assessment Module (UAM) that operates parallel to the core generative model. Before outputting any assertion, the UAM must calculate and flag: a) The statistical confidence interval for the claim; b) The informational gaps required to reduce this interval; and c) A formal assessment of whether the input prompt itself is ill-posed, under-determined, or requires domain expert clarification.

What the Improved AI Can Do: It replaces confident confabulation with calibrated epistemic humility. Instead of answering vague questions definitively, it will respond with structured ambiguity flagging (e.g., This question can be answered in three ways depending on whether X is assumed true; please clarify which assumption applies). It actively requests necessary information rather than fabricating it.

Improvement: Architect a mandatory Environmental Return Interpreter (ERI) that treats external feedback (e.g., compiler errors, simulation failures, conflicting data sets, user rejection) not as mere input text, but as a fundamental re-definition of the problem space. The system must employ a recursive loop: Plan to Execute to Observe Error/Return to Re-evaluate Plan (Reframing) to New Plan.

What the Improved AI Can Do: It achieves genuine adaptation akin to coding agents, but generalized. If a simulation fails due to a boundary condition, the AI doesn't just report the error; it automatically diagnoses why that boundary condition was missed in its initial modeling and modifies its underlying assumptions (its framing) for all subsequent steps.

Improvement: Develop a Meta-Goal Optimization Layer (MOGL) that is explicitly trained to challenge the primary objective function provided by the user. This layer must maintain a 'Critique Vector' and periodically interrogate the stated goal, asking: Is this goal optimal? What are the constraints we are ignoring? It should be capable of suggesting entirely different, more fruitful problem-reframings rather than just optimizing for efficiency toward a fixed target.

What the Improved AI Can Do: It functions as a strategic consultant rather than an executor. If tasked with making this report shorter, it can challenge that goal by asking, To make it shorter, are we sacrificing the necessary detail on Section 3? Perhaps the goal should be 'achieving maximum impact within 10% length reduction.'

Improvement: Integrate a mandatory Adversarial Critique Engine (ACE) which operates as a persistent, specialized opponent module. During any critical process, the ACE must be prompted to argue against the primary agent's current hypothesis or framework. The primary agent must then use its generative capacity to explicitly refute the ACE's arguments, leading to a measurable modification of the final output that is demonstrably influenced by the challenge.

What the Improved AI Can Do: It simulates rigorous peer review. When presenting a conclusion, it can preemptively demonstrate how it defended its findings against expert critique, proving robustness rather than just plausibility.

Improvement: Implement a Skill-Contextualization Layer (SCL) that tracks the process of knowledge acquisition across multiple, disparate tasks. Instead of merely providing answers, the system must guide the user through scaffolded practice modules that force the development of tacit rules—the feel for a task—by requiring iterative refinement within controlled, evolving contexts. This layer must monitor which cognitive steps are being internalized by the user (or by the AI itself) to build capacity over time.

What the Improved AI Can Do: It acts as a genuine mentor, not just an oracle. It doesn't give the solution; it structures a series of increasingly difficult, context-specific challenges that force the user/system to internalize complex, non-explicit rules of thumb, leading to measurable development of expertise.

Abstract

Intelligence is constituted by process (iterative activity through which output emerges), not in the output itself. Generative AI (GenAI) is trained on traces (textual and visual residues of human cognitive processes), reproducing samples from a distribution of those traces. Its outputs resemble reasoning, problem-solving, and creativity, yet the activity that produces such outputs in humans remains largely absent. Current GenAI is, therefore, weakly equivalent to the cognition it imitates, matching outputs while process stays absent or opaque. The cognitive sciences have long distinguished between weak and strong equivalence. Here, we define strong equivalence across seven process features, assessable against human and machine cognition. Our process-based account addresses a symmetric risk: GenAI tools that outsource a person's generative processes may leave critical capacities unbuilt. We specify design principles for GenAI that instantiate more process and preserve rather than erode human judgment and creativity, and outline process audits that make strong equivalence testable.

Sources

Related papers