Predicting consumer-technology ownership without a diffusion history
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Predicting consumer-technology ownership without a diffusion history".
Jane: The paper was written by Irina Vartanova, Niels Selling, Jennifer Viberg Johansson and Pontus Strimling from Institute for Futures Studies and Uppsala University and Linköping University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title and Authors: Tom: Welcome back to the show, everyone. Today we're digging into a paper that's been making the rounds, and I gotta say, the title alone got me hooked. "Predicting consumer-technology ownership without a diffusion history." That's a mouthful, but it's asking something really bold.
Jane: It really is, Tom. And the author list is interesting too. Irina Vartanova, Niels Selling, Jennifer Viberg Johansson, and Pontus Strimling, from the Institute for Futures Studies in Stockholm and Linköping University. These folks study the future for a living, so they're not just playing around with toy problems.
Tom: Right, and the core question is deceptively simple. Can you look at a brand new gadget that just hit the market, something nobody owns yet, and predict how many people will own it in a few years? Without waiting for any sales data to come in.
Jane: Exactly. And that's the "without a diffusion history" part. Normally, if you want to forecast adoption, you need a track record. You need to see the curve start to form, then you can project where it's going. But this paper says, what if we could skip all that?
Tom: And that matters because, think about it, the decisions that depend on these forecasts happen before the product launches. Investors need to know whether to fund it. Regulators need to know whether to worry about it. You can't wait six months for sales data when the funding round is next week.
Jane: So they're trying to build a tool that works in that pre-launch window. And the way they do it is by looking at the attributes of the technology itself. How useful is it? How easy is it to use? Is it fun? Do people think their friends would approve?
Tom: Those are the UTAUT2 constructs, right? The Unified Theory of Acceptance and Use of Technology. It's a framework that's been around for a while, usually used to study how individuals decide to adopt a single technology.
Jane: But here's the twist. They're using it across sixty-five different technologies at once, from smart TVs to robot nannies. And they're asking whether those perceived attributes can predict population-level ownership, not just individual intentions.
Tom: And the results, we'll get into the numbers in a bit, but let me just say, the fact that they got any signal at all out of this is genuinely surprising. I mean, sixty-five products spanning everything from air fryers to three dee printers, and the attributes still matter?
Jane: They do, and the fact that they can hold out one technology, train on the rest, and predict the held-out one's ownership with reasonable accuracy, that's the real achievement here. It's not just a correlation story.
Tom: So stick around, because next we're going to break down what they actually found, and I promise, the part about the language models rating the technologies is going to blow your mind.
Summary: Jane: So, Tom, let's get into the meat of it. The paper's summary lays out the whole experiment, and the setup is pretty clever. They ran a survey back in two thousand twenty-two on Prolific, over six hundred US adults, asking them about sixty-five different consumer technologies.
Tom: And they asked two things. First, do you own this? Second, rate it on these six attributes, like usefulness, ease of use, how fun it is, that kind of thing. And then they did a follow-up in two thousand twenty-five to see how ownership had changed.
Jane: Right. And the key move here is that they didn't just use the human ratings. They also asked two large language models, Claude Opus four point seven and GPT-five point five, to rate the same sixty-five technologies on the same attributes. Same questions, same scale, everything.
Tom: And that's where it gets wild. When they trained a model to predict ownership from the attributes, the language model ratings did way better than the human ratings. I'm talking mean absolute error of about seven percentage points for the best model, versus twelve percentage points for the human raters.
Jane: Let me put that in perspective. If a technology has, say, twenty percent ownership, the language model prediction is off by about seven points, so it might say thirteen or twenty-seven percent. The human raters are off by twelve points, so they might say eight or thirty-two percent. That's a big difference in practical terms.
Tom: And they compared that against a baseline that just uses how long the technology has been on the market. That baseline gets about fourteen and a half percentage points of error. So the language model attributes cut that error almost in half.
Jane: But here's the thing that really surprised me. When they tried to predict the change in ownership from two thousand twenty-two to two thousand twenty-five the attributes didn't help at all. The best prediction was just assuming nothing changes. The error was about the same whether you used the attributes or not.
Tom: Yeah, that's the humbling part. The attributes tell you the level of adoption, how many people own it at a given moment. But they don't tell you the movement, whether it's growing or shrinking. And over a three-year window, most of these technologies barely moved anyway.
Jane: So the paper is really honest about its limits. It can predict where a technology stands, but not where it's going. And that's a crucial distinction, because forecasting the future is usually about the change, not the level.
Tom: Still, even getting the level right, without any diffusion history, that's a big deal. And it sets up the next part of our discussion, which is how they actually built this thing and what improvements they're suggesting. Because there's a lot of clever machinery under the hood.
Improvements: Tom: Alright, Jane, let's talk about what this paper actually improves on. Because the whole field of new product forecasting has been around for decades, and it's had some real problems.
Jane: Right. The traditional approach is the Bass model, which fits an S-shaped adoption curve to sales data. The problem is you need enough data to see the curve's peak, and for a brand new product, you just don't have that.
Tom: And then there's the analogy approach. You find a similar product from the past and assume the new one follows the same path. But that breaks down when there's no good analogue, which is exactly when you need the forecast the most.
Jane: And the third approach is asking experts or potential customers directly. But experts are overconfident, and people's stated intentions don't match their actual behavior. The paper cites a whole literature showing that gap.
Tom: So what does this paper do differently? Instead of trying to fit a curve or find an analogue, it builds a mapping from perceived attributes to ownership levels, across many technologies at once. And that mapping is what generalizes to a new product.
Jane: And the improvement here is twofold. First, they're using perception-based ratings, not expert opinions. Regular people, and now language models, scoring how useful or fun a product seems. That's a much more scalable source of data.
Tom: Second, they're spanning product categories. Previous work, like Lee and colleagues in two thousand fourteen was limited to a single category. This paper goes across sixty-five products, from kitchen gadgets to AI robots, and the mapping still holds.
Jane: And the language model piece is the real innovation. Because getting human ratings for every new product is expensive and slow. You have to field a survey, wait for responses, clean the data. A language model can rate a product in seconds.
Tom: But that raises a question, and the paper addresses it head-on. Are the language models just cheating? I mean, they were trained on data up to two thousand twenty-six so they know which of these products became popular. Could they be using that knowledge to game the attribute ratings?
Jane: That's the knowledge-cutoff problem, and the paper has a clever check for it. They compared an earlier and a later version of each model family. If the later version knew more about outcomes, it should rate known successes higher on the success-related attributes.
Tom: And they found no such pattern. The movement between versions looked like recalibration, not outcome knowledge. The more a model's recall of a product changed across the cutoff, the less its attribute ratings moved. That's the opposite of what cheating would look like.
Jane: So they're fairly confident the ratings are genuine attribute judgments, not memorized outcomes. But they're also honest that this is a narrowing of the explanation, not a complete exclusion. You can't fully test the long-known technologies.
Tom: Fair enough. So the improvements are real, but there's a catch. And that catch is what we're going to dig into next, because the first page of the paper lays out the whole problem in a way that really frames what they're trying to do.
First Page: Jane: So Tom, let's go back to the very beginning. The first page of "Predicting consumer-technology ownership without a diffusion history" sets up the problem in a way that I think really matters.
Tom: And the key distinction they draw is between the level of adoption and the movement over time. They say a diffusion curve carries two kinds of information that don't need to be predicted together. And this paper only tackles the level.
Jane: Right. And they're explicit about that. They say, "This addresses the first part of the problem, not the whole of it." They're not claiming to predict the full trajectory. They're just saying, here's a way to estimate where a technology stands.
Tom: And that's actually a really important framing, because a lot of forecasting papers overpromise. They act like they can predict the whole curve, and then they fail when the data doesn't cooperate. This paper is much more careful about its scope.
Jane: And the first page also mentions something that caught my eye. They say the current generation of LLM-driven consumer technologies has sharpened that pre-launch window. Products that emerged in the past two years would have appeared, on most lists from three years ago, as research demonstrations rather than household goods.
Tom: That's a great point. Think about it. In two thousand twenty-two if you asked someone to predict which AI products would be in homes by two thousand twenty-five they probably would have said smart speakers and maybe a robot vacuum. But now we've got AI companions, AI glasses, all sorts of stuff that didn't exist as consumer products three years ago.
Jane: So the problem of forecasting without history is becoming more urgent, not less. The pace of new technology introduction is accelerating, and the window for making decisions is shrinking.
Tom: And that's why this paper matters. It's not just an academic exercise. It's a tool that could actually help people make better decisions about what to build, what to fund, and what to regulate.
Jane: But there's also a cautionary note in the first page. They mention that the movement over time is the part that current data and the identifiability of the curve do not yet let us reach. So they're aware of the limitation.
Tom: And that's what makes the paper credible. They're not overselling. They're saying, here's what we can do, and here's what we can't. And that honesty is rare in this field.
Jane: So as we head into the conclusion, I think the big picture is clear. This paper takes a real step toward pre-launch forecasting, but it also marks where the limits are. And that's valuable in itself.
Conclusion: Tom: Alright, Jane, let's wrap this up. We've been talking about "Predicting consumer-technology ownership without a diffusion history" for a while now, and I think we need to pull it all together.
Jane: Absolutely. So the big finding is that you can predict the ownership level of a held-out technology using just four perceived attributes, and the best language model ratings get you to about seven percentage points of error. That's a real improvement over the fourteen-point baseline that just uses product age.
Tom: And the fact that language models beat human raters on this task, that's the headline. It suggests that for this kind of aggregate judgment, a model can be more consistent and more calibrated than a panel of people.
Jane: But the paper is also honest about what it can't do. It can't predict the change in ownership over time. Over the two thousand twenty-two to two thousand twenty-five window, the attributes added nothing beyond just assuming no change. And that's a real limit.
Tom: And they're also honest about the knowledge-cutoff concern. The models rated these products after their outcomes were public knowledge, and while the checks suggest the ratings aren't just memorized outcomes, you can't fully rule it out.
Jane: So where does that leave us? I think it leaves us with a promising tool for screening. If you have a portfolio of new technologies and you need to decide which ones to pursue, this could help you rank them by expected adoption.
Tom: And the deployment illustration at the end shows exactly that. They projected two thousand twenty-seven ownership for eleven products launched in two thousand twenty-five and two thousand twenty-six. Some are predicted to be niche, like the consumer humanoid robots, and some are predicted to be more mainstream, like the ultrasonic chef's knife.
Jane: And those predictions are testable. A later study could come back and check whether the model was right. That's the scientific method at work.
Tom: So, final thoughts. This paper doesn't solve the whole forecasting problem, but it makes a real dent in it. And it opens up a new direction, using language models as attribute raters at scale.
Jane: And that's the part I'm most excited about. Because if this works, it means we can screen hundreds of new technologies quickly, without fielding a survey for each one. That could change how innovation decisions get made.
Tom: Well said, Jane. And with that, we're going to say goodbye to this paper and get ready for the next one. Thanks for listening, everyone.
Jane: See you next time.
Irina Vartanova, Niels Selling, Jennifer Viberg Johansson, Pontus Strimling
Institute for Futures Studies · Uppsala University · Linköping University
cs.CL, cs.CY, stat.AP
Submitted: 2026-06-03
Comments: 31 pages, 4 figures, supplementary material included (Tables S1-S6, Figure S1), data and code at https://osf.io/dr5ct
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 63/100
The gist: the 2022 human raters, Anthropic Claude Opus 4.7, and OpenAI GPT-5.5.
Key concepts
- Diffusion History
- Normally, predicting technology adoption requires historical sales data or a 'diffusion history.' This paper aims to bypass that need by forecasting ownership based on attributes alone, even for brand new gadgets.
- UTAUT2 Constructs
- These are perceived attributes used in the study (like usefulness and ease of use) derived from the Unified Theory of Acceptance and Use of Technology. The authors test if these perceived traits can predict population-level ownership.
- Language Models (LLMs)
- The hosts discuss using LLMs (like Claude Opus 4.7 and GPT-5.5) to rate technologies on attributes, noting that these model ratings performed better than human survey responses for predicting ownership.
- Pre-launch Forecasting
- This refers to making predictions about a product's market success and adoption levels before the product has been launched or when insufficient sales data exists.
Terminology
Summary
Summary
The paper tests whether the perceived attributes of a consumer technology predict how widely it is owned, without relying on any diffusion history of the focal technology. The authors collected ownership and attribute ratings on 65 consumer technologies in a 2022 Prolific survey of US adults (n = 678), with a 2025 follow-up survey that re-elicited ownership. They elicited six UTAUT2 attributes from three sources: the 2022 human raters, Anthropic Claude Opus 4.7, and OpenAI GPT-5.5.
The main model uses a sign-constrained penalized regression on four acceptance attributes (performance expectancy, hedonic motivation, effort expectancy, and social influence) plus a log-age covariate, evaluated by leave-one-technology-out prediction. The attribute model improves on a baseline of years-since-launch: mean absolute error falls by 17% with the human ratings, and by more with either model, most with Opus 4.7.
Specifically, the four-attribute models bring MAE down from 14.44 pp (covariate-only baseline) and 15.08 pp (predict-training-mean baseline) to 11.95 pp on the human raters, 7.55 pp on GPT-5.5, and 7.09 pp on Opus 4.7. The best rater, Opus 4.7, gives a 51% MAE reduction over the covariate-only baseline. The Opus cell improves on the baseline by 7.35 pp (95% CI [4.77, 10.09]), the GPT cell by 6.88 pp (95% CI [4.41, 9.40]), and the human-rater cell by 2.48 pp (95% CI [0.51, 4.37]). A permutation null check on the Opus cell found every random pairing of technologies with attributes yielded a higher MAE than the observed 7.09 pp (mean 14.76 pp, 95% CI [14.19, 15.31]).
Effort expectancy carries the largest positive coefficient on the logit scale for both model raters. For Opus 4.7, the remaining three attributes are smaller but not all zero, with social influence ahead of performance expectancy and hedonic motivation. For GPT-5.5, all four attributes carry positive weight, with effort expectancy clearly on top. The human raters keep effort expectancy at the top of the ranking and divide the remaining weight more evenly across hedonic motivation, social influence, and performance expectancy.
Over the short 2022-to-2025 window, where ownership moved little, the same attributes do not improve on a no-change baseline. The zero-change baseline gives MAE 3.23 pp, the anchor-only specification (logit-2022 ownership + log age, no attributes) reaches 2.56 pp, and adding the four-attribute battery returns Δ-MAE of 2.55 pp on the human-2022 ratings, 2.61 pp on GPT-5.5, and 2.59 pp on Opus 4.7. A paired bootstrap puts the Opus advantage over predict-zero-change at −0.64 pp with 95% CI [−1.28, 0.01], a difference that crosses zero.
The paper includes a deployment illustration: 2027 ownership predictions for eleven products launched in 2025 and 2026, ranging from a consumer ultrasonic chef's knife (predicted 5.61% by Opus, 7.90% by GPT) down to a consumer humanoid developer robot (predicted 0.36% by Opus, 0.17% by GPT).
The paper sets out limitations, including the possibility that language-model ratings reflect prior knowledge of these technologies rather than independent attribute reasoning. A knowledge-cutoff check compared earlier and later releases of each model family on the same battery and found no sign that prior knowledge of outcomes moved the ratings on products where that knowledge could have changed, but it cannot test the long-known technologies behind the main result. The paper also notes that the leave-one-out scheme holds out a technology the raters had already encountered by 2026, so the result is a prediction for technologies already on the market, not a forecast of one not yet observed, and that ownership prevalence is specific to the Prolific sample rather than a US-population share.
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems and what the improved systems can do:
Improvement: I can build a forecasting module that scores any unreleased consumer technology on four UTAUT2 attributes (performance expectancy, hedonic motivation, effort expectancy, social influence) using my internal knowledge, then applies a sign-constrained elastic-net regression with a log-age covariate to predict population ownership prevalence.
What the improved system can do:
-
Given a product name, brand, and short description of a technology not yet on the market, output a predicted ownership percentage for a specified future year (e.g., 2027) with a mean absolute error of approximately 7.1 percentage points versus 14.4 for a baseline that only uses years since launch.
-
Rank a portfolio of candidate pre-launch products by expected adoption, allowing investors or product managers to screen dozens of ideas in seconds without fielding surveys or waiting for sales data.
-
Update predictions as the product's launch date approaches, since the log-age covariate shifts the forecast.
Bottom line: The improved AI system can forecast adoption of unreleased consumer technologies with 50% lower error than time-based baselines, rate products without human panels, audit its own ratings for leakage, separate level from change predictions, and flag uncertainty—all while being transparent about its limits. This turns a research result into a deployable decision-support tool for product investment, portfolio screening, and market-entry timing.
Abstract
We test whether the perceived attributes of a consumer technology predict how widely it is owned. In a 2022 Prolific survey of US adults (n = 678), respondents rated 65 consumer technologies on six attributes. We then elicited the same ratings from two frontier language models, Anthropic Claude Opus 4.7 and OpenAI GPT-5.5. We regress ownership prevalence on four UTAUT2 acceptance attributes plus a log-age covariate with a sign-constrained penalized regression and evaluate it by holding out one technology at a time. The attribute model improves on a baseline of years-since-launch: mean absolute error falls by 17% with the human ratings, and by more with either model, most with Opus 4.7. Over the short 2022-to-2025 window, where ownership moved little, the same attributes do not improve on a no-change baseline. We set out the limitations of the approach, including the possibility that language-model ratings reflect prior knowledge of these technologies rather than independent attribute reasoning. We include a deployment illustration: 2027 ownership predictions for eleven products launched in 2025 and 2026.
Sources
- Dif4FF: Leveraging Multimodal Diffusion Models and Graph Neural Networks for Accurate New Fashion Product Performance Forecasting
- Can LLMs Think Like Consumers? Benchmarking Crowd-Level Reaction Reconstruction with ConsumerSimBench
- Cold-Start Forecasting of New Product Life-Cycles via Conditional Diffusion Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering