The Scaling Paradox in Human–AI Collaboration

summary

Video file (mp4)

The gist

This paper develops an analytical model to examine whether the empirical scaling benefits of AI—where larger models, more data, and more compute lead to predictable performance

In short

This episode discusses 'The Scaling Paradox in Human–AI Collaboration,' a paper by Anyan Qi and Mengxin Wang. The hosts explore how AI performance scaling breaks down when humans are involved. They conclude that overestimating AI capabilities leads to reduced human effort and poor outcomes, while practical solutions like cost internalization and perception alignment can help manage this paradox.

Key concepts

Scaling Law
This is the foundational idea that AI performance improves predictably as more compute and data are added. The paper uses this law as a baseline assumption, but tests whether it holds true when a human worker is integrated into the system.
The Scaling Paradox
This occurs when scaling up a powerful AI model causes the entire human-AI system to perform worse than expected. It happens because human overconfidence leads to reduced effort, causing under-investment that outweighs the gains from larger AI models.
Cost Internalization
This strategy involves making the worker bear some of the cost of using the AI, such as paying a share of its usage fees. This changes incentives and helps counteract overconfidence by making workers more careful about their effort and time.
Perception Alignment
This refers to correcting a worker's beliefs about what the AI can actually do. By providing transparent reporting or training, this alignment prevents workers from slacking off too much due to an inaccurate view of the AI's capabilities.

Terminology used across episodes

This episode discusses

The paper

The Scaling Paradox in Human-AI Collaboration · Read on arXiv

Anyan Qi, Mengxin Wang

University of Texas at Dallas

The discovery of scaling laws has highlighted the extraordinary potential of AI systems with a striking empirical pattern: as AI systems scale, their capabilities tend to improve predictably. Yet, in real-world applications, AI rarely operates in isolation; instead, it often works alongside humans, raising the question of whether these gains persist in human-AI collaboration. In this work, we develop an analytical model to examine when the empirical scaling benefits of AI translate into improved human-AI joint system performance. We demonstrate that the performance of a human-AI system can scale positively as the AI scales up-provided that humans have an accurate perception of the AI's capabilities. Human misperception, however, can fundamentally alter this relationship: i) when humans over-perceive the AI's capabilities, a scaling paradox may arise, in which greater AI scale reduces overall system performance and amplifies firm-level profit losses, and (ii) when humans under-perceive the AI's capabilities, performance still improves with scale but at a substantially slower rate. We further show that firms can actively manage these distortions through operational policies such as cost internalization and perception alignment, whose effectiveness depends on the economics of AI deployment and the direction of human misperception. These findings suggest that organizations may benefit more from managing the human-AI interface than from simply investing in larger, more expensive AI systems. More broadly, our results suggest that AI scaling should be viewed not only as a technological challenge, but also as a behavioral and operational one, and caution against the view that larger AI systems will automatically lead to better operational outcomes. Whether AI scaling creates value ultimately depends on how increased AI capabilities shape human beliefs and collaborative efforts.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "The Scaling Paradox in Human–AI Collaboration".

Jane: The paper was written by Anyan Qi and Mengxin Wang from University of Texas at Dallas.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the arXiv Radio Hour, everyone. I'm Tom, and today we're digging into a paper that's been making the rounds — it's called "The Scaling Paradox in Human–AI Collaboration."

Jane: And I'm Jane. Tom, this title grabbed me immediately. We keep hearing that bigger AI models are automatically better, and this paper is basically asking, "Is that actually true when humans are in the loop?"

Tom: Exactly. The authors are Anyan Qi and Mengxin Wang from UT Dallas. They're taking the scaling law — you know, the idea that AI performance improves predictably as you throw more compute and data at it — and they're asking whether that holds up when a human worker is part of the system.

Jane: So it's not just about the AI. It's about the whole team — the human plus the machine working together.

Tom: Right. And the title says "paradox," which is a hint that the answer isn't a simple yes. The paper shows that under some conditions, scaling up the AI can actually make the whole system worse.

Jane: That sounds counterintuitive. If the AI gets better at its job, how could the team get worse?

Tom: Because the human changes their behavior. If the worker thinks the AI is more capable than it really is, they'll slack off more than they should. And if they're wrong about how good the AI is, that overconfidence can tank the whole project.

Jane: So it's not the AI that breaks. It's the human's perception of the AI that breaks the system.

Tom: Precisely. And that's why this paper is so important. It's not just a technical result — it's a behavioral one. It's saying that the value of AI scaling depends on what people believe about the AI, not just what the AI can actually do.

Jane: And that's a huge deal for companies pouring billions into bigger models. They might be getting less than they think.

Tom: Or worse — they might be actively hurting their own productivity. That's the paradox. And we're going to spend the next few segments unpacking exactly how that happens and what firms can do about it.

Jane: I'm hooked. Let's get into the details.

Summary: Tom: So, Jane, we've set the stage. Let's talk about what this paper actually does. The core setup is a human and an AI working together on a project. The human sets up the task, the AI takes a crack at it, and then the human reviews and corrects the output.

Jane: And the AI's success rate depends on its scale — bigger AI, better odds of getting it right on the first try.

Tom: Right. The paper models that with a scaling law, where the failure rate drops exponentially as the AI gets bigger. But here's the twist — the human has to decide how much time to spend reviewing each project.

Jane: So there's a trade-off. Spend more time on each project to make sure it's right, or spend less time and get through more projects.

Tom: Exactly. And the human makes that call based on what they think the AI can do. If they think the AI is great, they'll review less. If they think it's weak, they'll review more.

Jane: And the paper shows that when the human's perception is accurate, everything works beautifully. The human adjusts their effort perfectly as the AI scales up, and the whole system gets better.

Tom: But when the human overestimates the AI — thinks it's better than it really is — they cut their effort too aggressively. And at some point, that under-investment outweighs the gains from the bigger AI.

Jane: So you get this weird situation where a bigger, more expensive AI actually makes the team perform worse.

Tom: That's the scaling paradox. And the paper shows it can get so bad that the human-AI team does worse than a human working alone.

Jane: And what about the opposite? What if the human underestimates the AI?

Tom: That's the safer mistake. The human works harder than they need to, so the system still improves as the AI scales — just not as fast as it could. It's a slower climb, but at least you're not falling.

Jane: So overconfidence is the dangerous one. That's a really important distinction for managers to understand.

Tom: And it gets even more interesting when you bring the firm's bottom line into it. But we'll save that for the next segment.

Improvements: Jane: Tom, we've talked about the problem. Now let's talk about what the paper says firms can actually do about it. Because just telling managers "your workers are overconfident" isn't very helpful without a fix.

Tom: Right. And the paper proposes two concrete levers. The first is what they call cost internalization — basically, making the worker bear some of the cost of using the AI.

Jane: So if the AI costs money per use, the worker has to pay a share of that out of their own pocket.

Tom: Exactly. And that changes their incentives. If the worker has to pay for each AI-assisted project, they'll be more careful about how many projects they take on and how much effort they put into each one.

Jane: So it's a way to counteract the overconfidence. If the worker thinks the AI is amazing and wants to rush through everything, making them pay for it slows them down.

Tom: But here's the catch — the paper shows that the optimal amount of cost sharing depends on how expensive the AI is. If the AI is cheap, you can shift almost the whole cost to the worker and it works great. But if the AI is expensive, you need to absorb more of the cost yourself, or the worker will just stop using the AI altogether.

Jane: So it's a balancing act. Too much cost shifting and the worker abandons the tool. Too little and they overuse it.

Tom: And the second lever is what they call perception alignment. That's about correcting the worker's beliefs about what the AI can actually do.

Jane: So training, transparent reporting of AI performance, showing workers where the AI fails.

Tom: Right. And the paper shows this is really valuable when workers overestimate the AI. Aligning their beliefs with reality prevents them from slacking off too much.

Jane: But what about under-perception? If the worker thinks the AI is worse than it is, wouldn't aligning their beliefs also help?

Tom: Here's the surprise — not necessarily. The paper shows that when a worker underestimates the AI, they work harder than they otherwise would. And sometimes that extra effort is actually good for the firm, because it compensates for the fact that the worker doesn't internalize the AI's cost.

Jane: So a little bit of skepticism can be a feature, not a bug?

Tom: Exactly. The paper shows that aggressively correcting under-perception can actually reduce firm profit in some cases. It's only when the under-perception is really severe that alignment helps.

Jane: That's a really nuanced result. It's not just "fix all misperceptions." It's about understanding which direction the misperception goes and responding appropriately.

Tom: And that's what makes this paper so practically useful. It gives managers a framework for thinking about when and how to intervene.

First Page: Jane: Tom, let's go back to the very beginning of the paper for a moment. The first page sets up this really compelling motivation.

Tom: It does. The authors start with the scaling law — the idea that bigger AI models get predictably better. And they point out that this is usually measured in isolation, with the AI working alone on a benchmark.

Jane: But in the real world, AI doesn't work alone. It works with doctors, lawyers, engineers, customer service agents.

Tom: And the paper gives some striking examples. There's the METR study where experienced developers thought AI made them twenty percent faster, but they actually got nineteen percent slower. They believed the AI was helping, but it wasn't.

Jane: And there's Klarna — they touted their AI chatbot as replacing hundreds of human agents, and then quietly brought humans back because customers wanted empathy.

Tom: And the paper mentions that despite billions in enterprise AI investment, most organizations report zero return. That's a massive gap between expectation and reality.

Jane: So the first page is really making the case that this is a real, widespread problem. It's not just one company making a mistake — it's a systemic issue.

Tom: And the authors argue that the reason is behavioral. Humans change how they work when AI is introduced, and if their beliefs about the AI are wrong, the whole system suffers.

Jane: So the scaling law holds for the AI in isolation, but it breaks down when you put a human in the loop.

Tom: That's the core insight. And it's why the paper is so important. It's telling us that AI scaling isn't just a technical challenge — it's a human challenge.

Jane: And that's a message that should resonate with anyone deploying AI in their organization.

Tom: Absolutely. And it sets up the rest of the paper, where they build the model and show exactly when and why scaling fails.

Conclusion: Tom: Well, Jane, we've covered a lot of ground on "The Scaling Paradox in Human–AI Collaboration."

Jane: We have. And I think the biggest takeaway is that bigger AI doesn't automatically mean better outcomes. It depends on what the humans in the system believe and how they respond.

Tom: Right. When workers accurately perceive the AI's capabilities, scaling works beautifully. When they overestimate it, you can get the paradox — bigger AI, worse performance, and amplified profit losses.

Jane: And when they underestimate it, you get slower gains, but at least you're not falling backward.

Tom: And the paper gives us practical tools to manage this. Cost internalization and perception alignment can help, but they need to be tailored to the specific situation — the cost of the AI and the direction of the misperception.

Jane: So for managers, the message is clear. Don't just invest in bigger models. Invest in understanding how your workers perceive those models and manage that perception actively.

Tom: Exactly. And that's a much more nuanced — and I'd argue more useful — message than "AI will save us all" or "AI is overhyped."

Jane: It's a thoughtful, balanced take on a topic that's usually discussed in extremes.

Tom: Well said. That's all the time we have for this paper. Thanks for joining us, and we'll be back soon with the next one.

Jane: Take care, everyone.

More episodes

← Home