Innovating with Generative AI: A Human Bottleneck Framework
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Innovating with Generative AI: A Human Bottleneck Framework".
Jane: The paper was written by Julian De Freitas, Ayelet Israeli, Gideon Nave, Artem Timoshenko and Olivier Toubia from.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Welcome back to the channel, everybody. Today we’re digging into a paper that’s been making the rounds called “Innovating with Generative AI: A Human Bottleneck Framework.” Jane, I gotta say, just that title alone got me excited — it’s not another “AI will save innovation” hype piece, is it?
Jane: Not at all, Tom. And that’s exactly why I wanted to bring it on. The authors — Julian De Freitas, Ayelet Israeli, Gideon Nave, Artem Timoshenko, and Olivier Toubia — they’re asking a much sharper question. Instead of asking what AI can do, they’re asking what humans can’t do, and where AI actually helps versus where it just makes things worse.
Tom: So it’s flipping the whole conversation on its head. We usually hear about AI’s capabilities, but they’re saying the real bottleneck is us — human cognition, group dynamics, organizational habits.
Jane: Right. And they call these “human bottlenecks.” For example, in ideation, we suffer from cognitive fixation — we get stuck in our mental ruts. The paper argues that if you just hand an AI tool to someone and let them use it naively, it can actually deepen that fixation, because the AI outputs anchor you even further into familiar territory.
Tom: That’s a punch in the gut, honestly. So the tool that’s supposed to make us more creative can make us less creative if we’re not careful.
Jane: Exactly. But the good news is they don’t stop at identifying the problem. They lay out interventions for each bottleneck. For cognitive fixation, they suggest chain-of-thought prompting — getting the AI to first sketch out bold directions before elaborating. And that works for the AI, but interestingly, it doesn’t work for humans. You can’t just tell a person “think more diversely” and expect it to happen.
Tom: Wait, really? So the same prompt that helps the model doesn’t help us?
Jane: Nope. There’s a whole literature on ironic process theory — trying to suppress a thought keeps it active. So the fix has to be on the AI side, not on the human side. That’s a really clean example of their whole framework: understand the mechanism, then decide where to intervene.
Tom: I love that. It’s so specific. And it makes me wonder — are there bottlenecks where AI genuinely can’t help at all?
Jane: Oh, absolutely. And that’s coming up in the next segment. But let me just tease it — they point to lead users, people with extreme lived experience of a problem, as one of the few natural ways humans escape fixation. And that kind of tacit knowledge, built through hands-on practice, is probably not something AI trained on text can ever fully substitute.
Tom: Okay, I’m hooked. Let’s keep going — I want to hear more about where AI hits its limits.
Abstract: Tom: So we’re back with “Innovating with Generative AI: A Human Bottleneck Framework,” and Jane, the abstract really sets up this whole roadmap. Can you walk us through it?
Jane: Sure. The abstract lays out the central premise: many constraints in innovation are cognitive and social in origin — how we generate ideas, how we evaluate novelty, how we communicate through social systems. And generative AI doesn’t act uniformly on these constraints. At each stage, it can deepen some bottlenecks while alleviating others.
Tom: So it’s not a blanket “AI helps” or “AI hurts.” It depends on the specific mechanism.
Jane: Exactly. And they organize the paper around four stages of innovation: ideation, screening and testing, preference measurement and consumer insight, and then diffusion and market learning. For each stage, they identify the human bottleneck, then ask how generative AI interacts with it.
Tom: And I remember from the abstract they also make a really important distinction — between bottlenecks that will narrow as AI capabilities improve, and those rooted in enduring human constraints that won’t go away no matter how good the models get.
Jane: That’s the key insight for me. For example, in screening and testing, they talk about novelty-averse evaluation. People say they want creative ideas, but implicitly they associate novelty with uncertainty and risk, so they systematically discount original ideas. That’s a deep human bias — loss aversion, status quo bias. AI can’t just train that away.
Tom: But AI can make it worse, right? Because AI-generated ideas are fluent and polished, and that fluency gets mistaken for quality.
Jane: Right. The paper calls that a “false positive” trigger. Evaluators see a well-structured, smooth-sounding idea and assume it’s better, even if it’s actually mediocre. So the intervention has to be structural — strip ideas to a common format before scoring, blind evaluators to whether a human or an AI wrote it.
Tom: That’s so practical. It’s not about telling people to try harder; it’s about redesigning the evaluation process so the bias doesn’t get triggered in the first place.
Jane: And that’s the real contribution of this paper. It gives you a diagnostic framework. When a new AI capability comes out, you can ask: which bottleneck is it operating on? Is it addressing the mechanism, or is it just adding more fuel to the fire?
Tom: I also noticed they mention digital twins in the abstract — simulating individual consumers. That sounds super futuristic. What do they say about that?
Jane: That’s in the preference measurement section, and it’s sobering. They found that digital twins — even ones built from rich behavioral profiles — fail to reproduce non-normative behaviors like the sunk cost fallacy or loss aversion. So the twins choose rationally when real consumers wouldn’t. And that matters because adoption of genuinely new products is governed by those irrational behaviors.
Tom: So the simulation looks great on paper but misses exactly the behaviors that predict whether a novel product takes off.
Jane: Precisely. And that’s why the paper argues some bottlenecks — like the articulability gap, where consumers can’t articulate preferences for products they’ve never encountered — still require ethnographic and observational methods. No AI fix for that, at least not yet.
Tom: Okay, so we’ve got a framework, we’ve got specific bottlenecks. What are the actual interventions they’re proposing? I feel like that’s where the rubber meets the road.
Improvements: Tom: We’re back with “Innovating with Generative AI: A Human Bottleneck Framework,” and Jane, I want to dig into the interventions. The paper doesn’t just diagnose problems — it actually proposes fixes. What stood out to you?
Jane: One of my favorites is in the ideation section, specifically for design fixation. That’s when creators converge on the surface features of successful examples — like copying the look of a winning design without understanding why it works. The paper suggests using AI to generate variants of successful exemplars that keep the core elements but change the visual execution.
Tom: So you get the essence without the copy-paste trap. That’s clever.
Jane: And it’s grounded in evidence. They cite work by Wu, Aridor, and Timoshenko showing that AI intermediation can separate surface features from core design elements. Designers still learn from the exemplar, but they’re not anchored to its exact appearance.
Tom: What about the group dynamics bottleneck? That psychological safety thing?
Jane: That’s a big one. In organizational settings, people self-censor because they fear judgment — especially junior employees in front of seniors. The paper proposes an AI-mediated hybrid nominal process: people generate ideas privately with the AI, then submit them to an anonymized pool where the model removes source cues and clusters near-duplicates. Group discussion only starts after that pool exists.
Tom: So you separate idea generation from social display. That’s really smart — it reduces production blocking and evaluation apprehension.
Jane: Exactly. And they also talk about separating divergent and convergent modes. During divergence, the AI acts as a “yes-and” facilitator, building on ideas. During convergence, it switches to critical mode, surfacing assumptions and failure modes. The point is that psychological safety and critical evaluation are complements — they belong at different stages of the process.
Tom: Now, in the screening section, I remember they had some really concrete numbers. Something about a study with one hundred fifty-three ideation contests?
Jane: Yes — Kireyev and colleagues. They combined LLM-generated ratings with historical human expert ratings across one hundred fifty-three contests and over seventy-four thousand submissions. The model could identify all sponsor finalist choices while reviewing twenty-eight point four percent fewer submissions compared to sorting by average expert scores. And here’s the kicker — most of that gain came from re-weighting the historical human ratings, not from the LLM signal itself.
Tom: So the AI’s real value was in optimizing which submissions needed careful human review, not in replacing the human judgment.
Jane: Exactly. That’s a beautiful example of human-AI complementarity. The AI handles the volume problem — triaging thousands of submissions — but the substantive judgment stays with humans.
Tom: And what about the rationales? I recall they found something surprising about AI explanations.
Jane: Oh, that’s fascinating. Lane and colleagues ran a field experiment where AI screening recommendations came with written rationales. The rationales increased compliance with rejection recommendations, but they didn’t improve evaluators’ ability to discriminate between correct and incorrect AI judgments. In fact, evaluators who got rationales engaged less with the underlying submissions — the rationale substituted for independent verification.
Tom: So the explanation made people more likely to go along with the AI, even when the AI was wrong.
Jane: Right. And the effect was strongest for borderline cases — exactly where human judgment matters most. So the paper suggests withholding rationales in screening contexts. Black-box recommendations actually improved decision quality compared to human-only evaluation, while the same recommendations with narrative justifications did not.
Tom: That’s counterintuitive and super important. We assume explanations help, but they can actually make us lazier. Okay, so we’ve covered ideation and screening. What about the consumer insight stage? I know there were some tricky issues there.
Jane: That’s where the authenticity gap comes in — and it’s coming up next. Let’s get into that.
First Page: Tom: We’re still on “Innovating with Generative AI: A Human Bottleneck Framework,” and now we’re looking at the opening pages. Jane, the introduction sets up this whole problem — firms are using generative AI at every stage of innovation, but the empirical findings are accumulating faster than frameworks for interpreting them.
Jane: That’s the gap they’re filling. And I love how they put it — the answer to which developments matter, which human constraints they operate against, and what the division of labor should be — all of that still depends on features of human involvement that predate generative AI.
Tom: So the framework is built on stable human bottlenecks, not on rapidly changing AI capabilities. That’s why it’s going to stay relevant even as the models improve.
Jane: Exactly. And they’re explicit about that — they say they’re grounding the analysis in human behavior rather than AI capabilities, so you can assess whether new developments alleviate or intensify the bottlenecks that matter most.
Tom: Now, I remember from the first page they also mention this idea of “digital twins” and consumer simulation. That’s in the preference measurement section. What’s the big issue there?
Jane: The authenticity gap. Preferences aren’t stored — they’re constructed in the moment, sensitive to context, framing, comparison sets. And LLMs trained on text tend to capture the central tendency of expressed preferences, not the distribution of individual preferences. That’s what they call heterogeneity loss.
Tom: So you get the average consumer, but you lose the variation that drives segmentation and targeting.
Jane: Right. And there’s also prompt-induced confounding. If you vary price in a blind prompt, the model treats it as informative of context — higher price signals premium quality — rather than as an experimental manipulation. Gui and Toubia showed this: human data produced downward-sloping demand curves, but GPT-four produced flat or even upward-sloping curves.
Tom: That’s a fundamental failure. The model reads price as a signal, not as a variable you’re manipulating.
Jane: And the fix is to unblind the experimental design — tell the model that a variable is being experimentally varied across conditions. That consistently improves simulation accuracy. But there’s a trap: if you unblind the contextual variables instead, you get focalism — the model over-attends to them. So you have to unblind the design, not the variables.
Tom: So there’s a right way and a wrong way to prompt for simulation. That’s really actionable.
Jane: And then there’s the articulability gap — preferences consumers can’t articulate because they’ve never encountered a product that would make the preference salient. The canonical example is the touchscreen smartphone. Nobody in two thousand six could articulate a preference for swipe navigation.
Tom: Because the possibility hadn’t entered their experience yet.
Jane: Exactly. And AI trained on language can only surface preferences that have been expressed somewhere in the training corpus. So a preference that’s never been articulated is, by construction, absent. That’s why ethnographic and observational methods remain the only route to those “hidden gems.”
Tom: So some bottlenecks are structural — no amount of scaling fixes them.
Jane: Right. And that’s the core message of the framework. Some constraints are informational or volumetric — AI can help. Others are rooted in tacit lived experience or non-normative behavior — AI can’t help, and might even make things worse.
Tom: Okay, so we’ve got the framework, the bottlenecks, the interventions. What about the bigger picture? What does this mean for how we think about innovation as a whole?
Jane: That’s the final section of the paper — and it’s coming up in our conclusion.
Conclusion: Tom: We’ve covered a lot of ground on “Innovating with Generative AI: A Human Bottleneck Framework.” Let’s wrap this up, Jane. What’s the big takeaway for our listeners?
Jane: The big takeaway is that the question isn’t whether humans stay in the innovation process — it’s when and how. The framework gives you a way to figure that out by looking at the mechanism behind each bottleneck. If the constraint is informational, volumetric, or methodological, AI is the right intervention. If it’s rooted in tacit lived experience, non-normative behavior, or motivated interpretation, the constraint sits with humans, and no foreseeable AI capability dissolves it.
Tom: And that’s a really useful diagnostic. When a new AI tool comes out, you can ask — which bottleneck is it addressing, and is it addressing the mechanism or just adding more fuel?
Jane: Exactly. And the paper also raises systemic concerns that go beyond individual stages. There’s the deskilling risk — as AI takes over ideation, screening, and consumer interviews, junior practitioners lose the formative work that builds tacit judgment. And there’s the automation trap — if all the feedback signals are AI-generated, the learning loop gradually loses contact with actual consumer behavior.
Tom: That’s scary. You could have a system that’s internally consistent but completely disconnected from reality.
Jane: Right. And they also talk about inequality and representation. Because AI systems are trained on majority populations, they implicitly optimize for those groups. Digital twins are more accurate for affluent, highly educated consumers. LLM-based conjoint recovers population averages but not between-group differences. So segments whose preferences are hardest for AI to reproduce are the ones whose unmet needs are least likely to surface.
Tom: So the tools we’re building might systematically under-serve minority segments.
Jane: Unless we’re deliberate about it. The paper suggests augmenting rather than replacing human surveys, and being careful about where AI adoption deepens versus alleviates representation gaps.
Tom: And there’s that whole section on rethinking the innovation process itself — the inverted development curve, AI agents as end users. That’s pretty wild.
Jane: Yeah — prototyping is now cheap, but reliable deployment is hard. So the screening bottleneck shifts from “should we build this?” to “can we make this reliable enough to deploy?” And they even ask what innovation looks like when AI agents are the customers. That’s a whole new frontier.
Tom: I love that this paper doesn’t just analyze the present — it pushes us to think about the future. So, final thoughts, Jane?
Jane: I think the most important message is that the human capabilities the framework relies on — lead-user immersion, ethnographic observation, willingness to challenge organizational narrative — are exactly the ones current practice is least equipped to protect. They’re load-bearing, they erode under conditions that look like progress, and they require deliberate organizational protection.
Tom: That’s a powerful way to end. Thanks for joining us on this deep dive into “Innovating with Generative AI: A Human Bottleneck Framework.” We’ll be back next time with another paper. Until then, keep questioning what the tools are actually doing to us — and for us.
Jane: See you next episode, everyone.
Julian De Freitas, Ayelet Israeli, Gideon Nave, Artem Timoshenko, Olivier Toubia
cs.HC, cs.AI, econ.GN, q-fin.EC
Submitted: 2026-06-23
Updated: 2026-08-11
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 70/100
The gist: ideation, screening and testing, preference measurement and consumer insight, and diffusion and market learning.
Key concepts
- Human Bottlenecks
- These are the cognitive and social limitations humans face in innovation, such as cognitive fixation or group dynamics issues. The paper argues that naive use of AI can sometimes deepen these bottlenecks instead of alleviating them.
- Chain-of-Thought Prompting
- This intervention suggests getting an AI to first sketch out bold directions before elaborating. While this works for the AI, it does not work for humans, highlighting where human cognitive limitations persist even with AI assistance.
- Authenticity Gap
- This occurs in consumer insight when preferences are constructed in the moment and are sensitive to context. LLMs trained on text capture central tendencies but lose the distribution of individual preferences, leading to a loss of variation that drives segmentation.
Terminology
Summary
Summary
The paper proposes a human bottleneck framework
for understanding how generative AI (GenAI) transforms the innovation process. The central premise is that many constraints traditionally plaguing the innovation process are cognitive and social in origin, rooted in how people generate ideas, evaluate novelty, and communicate through social systems.
The authors argue that "Generative AI does not act uniformly on these constraints. At each stage, it can deepen some bottlenecks while alleviating others, and predicting these outcomes requires understanding the underlying mechanisms of the constraint itself." The framework identifies bottlenecks in four stages of the innovation process: ideation, screening and testing, preference measurement and consumer insight, and diffusion and market learning. The paper's goal is to provide a principled basis for asking which developments are likely to matter, which human constraints they are operating against, and what the division of labor between generative AI and human capability should be,
concluding that the answer to all three questions still depends, in large part, on features of human involvement in the innovation process that predate generative AI.
Ideation Stage. The paper identifies three human bottlenecks. First, cognitive fixation at the individual level, where existing mental models–built from experience, knowledge, and habit–make some ideas accessible while rendering others effectively elusive.
Naive use of LLMs can deepen this fixation: LLMs are trained to generate text sequences by predicting the next token given the preceding context, so early outputs in an ideation session anchor the model toward related concepts,
and when a human ideator reads the model's outputs, those outputs may dominate the retrieval process in human memory, suppressing access to alternatives.
The intervention is chain-of-thought prompting, which works for LLMs but not for humans,
as it substantially increases within-session idea diversity in these models... but has essentially no effect on the diversity of ideas from human ideators.
Second, design fixation at the social level, where exemplar exposure makes specific surface features highly accessible in memory, disproportionately shaping what designers retrieve during ideation.
The AI analog is mode collapse,
where independent generative AI sessions sample from the same centralized distribution, producing outputs that cluster regardless of superficial variations in the prompt,
leading to a tragedy-of-the-commons dynamic: widespread generative AI adoption... narrows the collective solution space explored across an industry even as individual productivity rises.
Interventions include persona modifiers
that push the model into different regions of its training distribution
and AI intermediation that can separate surface features of designs from their more core elements by showing designers AI-generated variants of successful exemplars instead of the originals.
Third, psychological safety in group settings, where ideators frequently self-censor not because a novel idea is cognitively inaccessible but because they anticipate social judgment.
AI helps by removing the immediate risk of social judgment,
but the benefit is easily displaced rather than eliminated: if the AI-generated ideas are then presented and evaluated in a hierarchical group setting, status dynamics re-enter at the selection stage.
The intervention is an AI-mediated hybrid nominal process
where participants first use the model privately to generate and elaborate on ideas, then submit outputs to an anonymized pool,
separating idea generation from social display.
Screening and Testing Stage. The paper identifies three bottlenecks. First, novelty-averse evaluation, where people explicitly endorse creative thinking, they implicitly associate novelty with impracticality and uncertainty, triggering heuristic responses that disfavor genuinely original ideas.
AI exacerbates this "through its cognitive-fluency channel. Because LLM-generated ideas are fluent and well-structured, their polish can be mistaken for quality... Evaluators relying on fluency as a proxy for merit may favor AI-produced ideas not because they are better but because they are easier to process. Interventions include
stripping ideas to a common format before scoring, blinding evaluators to authorship (human versus AI), or having a model rewrite all submissions into a uniform style before review." Second, cognitive load, where assessing highly novel ideas requires more cognitive resources than familiar ones,
and cognitive load shifts evaluators from systematic to heuristic processing.
AI exacerbates the cognitive-load bottleneck directly, by dramatically increasing the number of ideas flowing into the evaluation pipeline.
The intervention is architectural: combining LLM-generated ratings with historical human expert ratings in a prediction model
to optimally select the items that need careful review, lowering effective cognitive load without delegating substantive judgment to the model.
Third, shared mental model lock-in, where screening panels that evaluate ideas together over time converge on implicit criteria for what counts as 'promising.'
AI exacerbates
this by changing both its substrate and its reach,
as a screening AI trained on historical panel decisions learns and replicates the panel's implicit evaluative criteria at scale, institutionalizing those criteria without anyone having articulated or endorsed them.
The paper cites a field experiment showing that "pairing AI pass/fail screening recommendations with a written rationale increased evaluators' compliance with rejection recommendations more than with acceptance recommendations, without improving their ability to discriminate between correct and incorrect AI judgments, and that
confident AI rationales thus furnish evaluators with ready-made justifications for dismissing the ideas the innovation process most needs to advance. Interventions include training models
on post-launch outcomes rather than on screening decisions and
the cleanest design choice is to withhold rationales altogether in screening contexts."
Preference Measurement and Consumer Insight Stage. The paper identifies two bottlenecks. First, the authenticity gap, which is the difficulty of capturing what consumers actually prefer and feel, rather than what they articulate when asked in a decontextualized research setting.
The paper notes that preferences are not stored, but are constructed in the act of elicitation,
and that consumer adoption of genuinely new products is disproportionately governed by nonnormative behaviors
such as the sunk cost fallacy, status quo bias, and loss framing. AI interacts with this gap through three mechanisms: (1) heterogeneity loss, where LLMs are trained on population-level text data, they tend to reveal the central tendency of expressed preferences rather than the distribution of individual preferences within the population,
and even when fine-tuned on demographically stratified human conjoint data, the model recovers population averages but differences across groups are incoherent
; (2) prompt-induced confounding, where when a variable like price is varied in a blind prompt, the model does not treat that variation as an exogenous manipulation; it treats it as informative of other contextual variables,
leading to implausibly flat or upward-sloping
demand curves; and (3) failure to reproduce non-normative behavior in synthetic data, where digital twins choose normative options far more often than human participants, do not exhibit the sunk cost fallacy, and fail to violate the independence axiom of utility theory in the way real consumers systematically do.
Interventions include fine-tuning on historical human conjoint data for aggregate accuracy, using LLM-generated responses as auxiliary data disciplined by a smaller human benchmark sample,
and unblinding the experimental design rather than the contextual variables
by telling the model that a variable is being experimentally varied across conditions.
For the non-normative behavior failure, no current intervention yet restores the irrational patterns that govern adoption of novel products,
so decisions whose outcomes depend for example on the sunk cost fallacy, omission bias, or loss framing remain the province of human research.
Second, the articulability gap, which concerns preferences consumers cannot articulate, often because they have never encountered a solution that would make the preference salient.
AI exacerbates this because AI systems trained on language can only surface preferences that have been expressed in language somewhere in the training corpus, so a preference that has never been articulated is by construction absent.
The intervention is that ethnographic and observational methods that watch consumers work around the failure of existing solutions, remaining the only route to hidden gems,
while AI serves supporting roles: it can activate the researcher's own tacit understanding by serving as a probe rather than a measurement,
and it provides scalable access to expressed customer needs
through fine-tuned models that can perform this transformation at least as well as professional analysts.
Diffusion and Market Learning Stage. The paper identifies three bottlenecks. First, social diffusion modeling, where adoption is not the sum of independent individual responses. It is a social process in which early adopters generate the proof that makes the product credible to mainstream adopters.
AI-based agent-based simulation inherits the same problem
of individual simulation failures, and likely amplifies it as errors propagate through interaction effects.
The paper states that no AI-side intervention currently resolves the bottleneck,
and the appropriate role for AI is the same one identified for preference measurement—screening plausible adoption paths rather than adjudicating which will materialize.
Second, attention and aggregation, where a product is in the market, it emits signals at volumes no human team can easily absorb,
and the signals that matter most for innovation—emerging use cases or silent non-adoption—are precisely the ones least likely to surface.
AI dramatically alleviates the volume constraint,
but the bottleneck is not fully resolved but shifted... Product teams now face a prioritization problem.
The intervention is scalable extraction of structured needs from unstructured text,
with the open question of whether it can be tuned to surface weak signals about emerging use cases or silent non-adoption—rather than the loudest themes in the corpus.
Third, motivated interpretation, where organizations have incentives to read [signals] charitably toward decisions already made,
and providing better data does not resolve the constraint if the underlying incentives to charitably interpret that data persist.
AI has a peculiar two-sided relationship
with this bottleneck: an AI system has no ego investment in past decisions and will surface a pattern of complaints about a championed feature with the same priority as a pattern of praise,
but the fluent synthesis it produces is easier to selectively quote in support of conclusions the team was already inclined toward.
The paper notes that whether these forces net out toward more honest interpretation or more sophisticated rationalization is, we suspect, the most important open empirical question in this part of the framework.
Broader Systemic Implications. The paper discusses three cross-cutting issues. First, organizational issues: the technical barrier to building AI-enabled tools has fallen dramatically,
allowing functional experts closest to the innovation problem... to participate directly in producing the logic of the tools they use,
but this widespread ability to generate code creates a dangerous illusion that anyone can deploy enterprise software.
The proposed solution couples a centralized technical infrastructure with distributed collaborative building.
Second, deskilling and automation trap: as generative AI takes over adjacent tasks in the innovation process, the human capabilities those tasks used to develop may quietly erode,
and the organization of the future may contain sophisticated users of AI-generated outputs who lack the judgment to recognize when those outputs are wrong.
The systemic risk is that the learning loop gradually loses contact with actual consumer behavior, producing generic outputs that lack real sensibility about real consumers.
Third, inequality and representation: because AI systems are trained predominantly on data from majority populations, they implicitly optimize for those groups,
and at the level of an innovation system that uses these methods to allocate research effort across segments, it compounds into market-level under-representation.
Re-thinking the Innovation Process. The paper identifies three structural changes. First, the inverted development curve: Reaching a prototype is now fast... but moving from a plausible response to a reliable deployment is substantially harder.
This means firms accumulate stalled pilots,
stage-gate processes lose informational value,
and competitive advantage shifts
to post-prototype infrastructure.
Second, the changing architecture of innovation, applying Von Hippel's sticky information
framework to distinguish three modes: AI as integrated innovator
(limited by the bottlenecks identified), AI-moderated needs elicitation
(using AI to interview human consumers), and AI as user innovation toolkit
(making solutions information accessible to users, empowering them to become innovators themselves
). The paper suggests the focus on Mode 1 may be disproportionate to its likely long-term value.
Third, innovating for AI agents: the internet as shifting from a human-centric content economy to a machine-centric agentic economy,
raising questions about Agent Experience (AX)
and whether preference measurement
methods developed for humans may not transfer.
The paper notes that AI agent 'preferences' are objective functions set by the humans or organizations that configure them,
creating a principal-agent problem where optimizing for AI agent criteria may therefore diverge systematically from optimizing for human satisfaction.
Conclusion. The paper reframes the dominant question whether humans will remain in the process
as when and how.
It concludes that "where the constraint is informational, volumetric, or methodological—cognitive fixation in ideation, cognitive load in screening, attention and aggregation in market learning—the appropriate intervention is on the AI system itself, while
where the constraint is in tacit lived experience, non-normative behavior, or motivated interpretation—design fixation across an industry, the articulability gap, the failure of digital twins to reproduce irrational adoption dynamics, the limits of social diffusion modeling, and the organizational incentives that distort post-launch signal—the appropriate intervention remains (for now) on the human side. The paper warns that
the human contributions the framework relies on — lead-user immersion, ethnographic observation, willingness to challenge organizational narrative — are exactly the contributions that current organizational practice is least equipped to protect, and that
those capabilities are load-bearing, they erode under conditions that look like progress, and they require deliberate organizational protection."
Improvements for AI systems
Based on the paper, here are specific improvements I can make to AI systems:
Improvement: Implement chain-of-thought prompting that forces the model to first generate brief idea summaries, then explicitly revise them to be bolder and more distinct before expanding. This counteracts the typicality bias from alignment training.
What the improved system can do: Generate significantly more diverse ideas within a single session, avoiding the clustering around typical outputs that current LLMs exhibit.
Abstract
We propose a human bottleneck perspective for understanding how generative AI transforms the innovation process. The central premise is that many constraints traditionally plaguing the innovation process are cognitive and social in origin, rooted in how people generate ideas, evaluate novelty, and communicate through social systems. Generative AI does not act uniformly on these constraints. At each stage, it can deepen some bottlenecks while alleviating others, and predicting these outcomes requires understanding the underlying mechanisms of the constraint itself. We identify bottlenecks in four stages of the innovation process: ideation, screening and testing, preference measurement and consumer insight, diffusion, and market learning. By grounding analysis in human behavior rather than rapidly changing AI capabilities, we offer a framework for assessing whether new developments alleviate or intensify the bottlenecks that matter most at each stage. We also distinguish bottlenecks likely to narrow as capabilities improve from those rooted in enduring human constraints. We further discuss AI-related issues that cut across the entire innovation pipeline, challenging the very existence and structure of the traditional innovation process.
Sources
- The Challenge of Using LLMs to Simulate Human Behavior: A Causal Inference Perspective
- TextBO: Bayesian Optimization in Language Space for Eval-Efficient Self-Improving AI
- A402: Binding Cryptocurrency Payments to Service Execution for Agentic Commerce
- Prompting Diverse Ideas: Increasing AI Idea Variance
- When Large Language Models contradict humans? Large Language Models' Sycophantic Behaviour
- Confucius Code Agent: Scalable Agent Scaffolding for Real-World Codebases
Related papers
- EduGage: A Multimodal Dataset and Benchmark for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning
- EvoDesign: Agentic Editable Diagram Creation via Design Expertise Evolution
- HAGI++: Head-Assisted Gaze Imputation and Generation
- Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
- Review of Explainable Decision Support and Adaptive Human-Machine Interfaces for Automation Transparency in Maritime Autonomous Surface Ships
- Towards Cognitive Process-Aware Proactive Writing Support