From Caveman to Expert Analyst: Energy Consumption of Variable LLM Tasks
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "From Caveman to Expert Analyst: Energy Consumption of Variable LLM Tasks".
Jane: The paper was written by Diego Manya, Ethan I. Thorpe, Ji Zhang, Myranda Shirk, Jiamian He et al. from University of North Carolina at Chapel Hill and Vanderbilt University and Arboretica and Vanderbilt Law School.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Alright, listeners, welcome back. Today we're digging into a paper that's got a title that just makes you smile: "From Caveman to Expert Analyst: Energy Consumption of Variable LLM Tasks." Jane, I gotta say, that title alone tells you these researchers have a sense of humor.
Jane: They really do, Tom. And that humor is hiding something serious underneath. The paper is about how much electricity different AI tasks use, and how we as regular users can actually do something about it. The "caveman" part comes from one of their actual prompting tricks they tested.
Tom: Right, they literally tested a prompt that tells the AI to respond like a smart caveman. Drop the articles, skip the pleasantries, just get to the point. And that's one of the energy-saving techniques they measured. But let's back up a second.
Jane: Good idea. So the team behind this is led by Angel Hsu at UNC Chapel Hill, with collaborators from Vanderbilt Law School and a few other places. They're looking at the energy problem from the demand side, not the supply side.
Tom: And that's a big deal, because most of the conversation about AI energy use has been about building more power plants and data centers. But these folks are saying, wait a minute, what if we just use less energy in the first place?
Jane: Exactly. They tested ten different commercial models from five major companies, including OpenAI, Google, Anthropic, and xAI. And they ran three hundred real user prompts through them, categorized by cognitive complexity using Bloom's Taxonomy.
Tom: Bloom's Taxonomy, that's the old education framework, right? Knowledge, comprehension, application, analysis, synthesis, evaluation.
Jane: You got it. And they found something pretty striking. The simpler the task, the less energy it takes. A simple knowledge question uses about zero point one six watt-hours, while a complex creation task uses about zero point six eight watt-hours. So there's a real gradient there.
Tom: And that matters because most of us are probably using the same AI for everything, whether we're asking for a quick fact or asking it to draft a whole strategy document.
Jane: Precisely. And the title hints at their biggest finding, which we'll get into in a moment. But the short version is, the way you ask matters, and the model you choose matters even more. Stick around, because this gets really interesting when we start talking about the actual numbers.
Summary: Tom: So we're back with "From Caveman to Expert Analyst: Energy Consumption of Variable LLM Tasks," and Jane, I want to dig into the headline finding, because it's honestly a bit shocking.
Jane: It is. The researchers compared what they call reasoning models against non-reasoning models. Reasoning models are the ones that think step by step before answering, like GPT-five point five Pro or Claude Opus. Non-reasoning models are the faster, cheaper ones like GPT-five point four Mini or Gemini Flash.
Tom: And the difference in energy use is wild. We're talking about reasoning models consuming roughly fifteen to twenty times more energy than their non-reasoning counterparts. For simple tasks, it's closer to twenty times.
Jane: But here's the kicker. When they compared the actual answers using semantic similarity, the responses were really close. For most task types, the similarity score was above zero point seven five, which means the non-reasoning model gave nearly the same quality answer.
Tom: So people are burning twenty times the electricity for answers that are basically the same? That seems like a massive waste.
Jane: For most everyday tasks, yeah. The paper suggests that unless you're doing something exceptionally complex or you need absolute precision, you probably don't need the reasoning model. And that's a huge finding because most commercial AI products default to reasoning now.
Lu: If I can jump in here, Tom. What's really interesting is that this matches what we're seeing in the broader AI efficiency research. The reasoning models are doing a lot of internal computation, generating hidden thoughts before they even start writing your answer. That's where the energy goes.
Tom: So it's not just that they write longer answers, they're thinking harder before they write.
Lu: Exactly. And the paper shows that this extra thinking doesn't always translate into better answers. The semantic similarity scores tell us the content is largely the same, just with more computational overhead.
Jane: And that has real implications for the energy grid. The paper estimates that if everyone in the US with internet access asked twenty questions a day using non-reasoning models, the total energy use would be between seven point one and one hundred seventy-seven point seven million kilowatt-hours per year.
Meng: But hold on, Jane. That range is huge. What's driving that spread?
Jane: Good question, Meng. It's the difference between models. Some non-reasoning models are much more efficient than others. Gemini three point one Flash, for example, was substantially more efficient than the other models they tested.
Meng: So model choice matters even within the same category. That's useful for someone like me who's actually deploying these systems.
Tom: And it's useful for regular folks too. The takeaway so far is, if you're using AI for everyday stuff, you can probably switch to a non-reasoning model and barely notice the difference, while cutting your energy footprint by a factor of twenty.
Jane: But there's more. Because the paper also tested prompting tricks that can cut energy even further. And that's where the caveman comes back in. We'll get into those specifics next.
Improvements: Tom: Welcome back to our look at "From Caveman to Expert Analyst: Energy Consumption of Variable LLM Tasks." We've covered the big model comparison, but now Jane, let's talk about the actual tricks people can use.
Jane: Right. The researchers tested three prompting techniques on the non-reasoning models. The first is the energy-efficient persona, where you tell the AI it's designed to minimize energy consumption. The second is asking for a minimal answer. And the third is the caveman prompt we mentioned at the top.
Tom: And the results are pretty impressive. The minimal answer prompt cut energy use by thirty-eight percent for simple knowledge questions and up to sixty-three percent for complex creation tasks.
Jane: The caveman prompt was also strong, with reductions between forty percent and forty-seven percent for most task types. But here's a wrinkle, for simple knowledge questions, the caveman prompt actually increased energy use by about five percent.
Meng: That's counterintuitive. Why would telling the AI to be terse make it use more energy?
Jane: The researchers think it's because the caveman prompt changes the structure of the response so much that the model has to work harder to reformat the information. It's not just about fewer words, it's about how the words are organized.
Tom: And the energy-efficient persona was the most interesting one to me. It gave the smallest energy reduction, between four percent and thirty-five percent, but it preserved the highest semantic similarity to the baseline answers.
Lu: That's actually a really important distinction, Tom. The persona prompt is a softer signal. It tells the model to be efficient without changing the fundamental style of the response. So you get energy savings with almost no change in what the answer looks like.
Meng: So if I'm a company rolling this out to employees, the persona prompt is the safest bet because it won't change the quality of their work product.
Jane: Exactly. And the paper makes the point that these techniques are accessible to anyone. You don't need technical expertise, you just copy and paste a sentence into your prompt.
Tom: And the aggregate numbers are meaningful. The paper calculates that if everyone in the US used the energy-efficient persona prompt, the savings would be enough to power seven thousand two hundred households for a year.
Lu: And that's with essentially no change in response quality. That's the kind of efficiency gain that actually matters in the real world.
Meng: But I want to push back a little. These are estimates based on client-side timing and theoretical power models. The actual energy consumption could vary a lot depending on the data center, the hardware, the batch size.
Jane: That's fair, Meng. The paper acknowledges those limitations. They're using simplified equations based on model parameters and inference time. But the relative differences between practices are still meaningful, even if the absolute numbers are rough.
Tom: And that's the key point. We don't need perfect numbers to know that using a non-reasoning model with a good prompt is going to use less energy than a reasoning model with no prompt. The direction is clear.
Jane: And the magnitude is large enough that it's worth acting on. We'll wrap this up with some final thoughts next.
Conclusion: Tom: So we're wrapping up our discussion of "From Caveman to Expert Analyst: Energy Consumption of Variable LLM Tasks." Jane, what's the big picture here?
Jane: The big picture is that individual users have real power to reduce AI energy consumption. The paper shows three things. First, reasoning models use fifteen to twenty times more energy than non-reasoning models, with minimal quality difference. Second, simple prompt modifications can cut energy use by up to sixty-five percent. And third, the energy-efficient persona prompt offers the best balance of savings and response quality.
Tom: And that's not just theoretical. The paper estimates real-world impacts, like the equivalent of seven thousand two hundred households' worth of electricity saved just from one prompting change.
Lu: What I find most exciting is that this opens up a whole new research direction. We're just starting to understand how user behavior shapes AI's environmental footprint. There's so much more to explore, like how different languages or cultural contexts affect energy use.
Meng: And from an engineering standpoint, these findings are immediately actionable. We can bake these prompts into our systems by default, or at least offer them as options. The infrastructure doesn't need to change, just the instructions.
Jane: And that's the beauty of it. No new hardware, no new data centers, no waiting for the grid to catch up. Just smarter habits.
Tom: The paper also makes a point about information transparency. Right now, users have almost no idea how much energy their AI queries use. Studies like this start to fill that gap.
Jane: And as the authors note, the specific numbers will change as the technology evolves. But the direction is clear, and the practices they identify are likely to remain useful for a long time.
Tom: So our listeners can take away something practical today. If you're using AI, try switching to a non-reasoning model for everyday tasks. And maybe add that energy-efficient persona prompt to your next query.
Jane: It's a small change that adds up. And with that, we're saying goodbye to this paper. Thanks for joining us, and we'll see you next time with another fascinating piece of research.
Tom: Take care, everyone.
Diego Manya, Ethan I. Thorpe, Ji Zhang, Myranda Shirk, Jiamian He, Angel Hsu, Michael P. Vandenbergh
University of North Carolina at Chapel Hill · Vanderbilt University · Arboretica · Vanderbilt Law School
cs.CY, cs.AI
Submitted: 2026-07-02
Code: https://github.com/JuliusBrussee/cavem
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 58/100
Key concepts
- Reasoning Models
- These are AI models designed to think step by step before answering, such as GPT-five point five Pro or Claude Opus. They consume significantly more energy than non-reasoning models because they perform internal computation and hidden thinking processes before generating a response.
- Non-Reasoning Models
- These are faster, cheaper AI models, like GPT-five point four Mini or Gemini Flash. They use less energy for tasks because they do not perform the extensive step-by-step thinking required by reasoning models. For many everyday tasks, these models provide nearly the same quality answer.
- Energy Gradient
- The paper found that the simpler a task is, the less energy it requires. A simple knowledge question uses about zero point one six watt-hours, while a complex creation task uses about zero point six eight watt-hours, showing a clear difference in energy use based on task complexity.
- Energy-Efficient Persona Prompt
- This prompting technique tells the AI to minimize energy consumption. It was found to offer the best balance of savings and response quality, reducing energy use by four percent to thirty-five percent while preserving high semantic similarity to baseline answers.
Terminology
Summary
Summary
This paper investigates the potential for demand-side management, specifically changes in retail (i.e., casual individual) user behavior, to reduce the electricity consumption and environmental impacts of large language model (LLM) chatbots. The authors note that while research and policy have focused on supplying low-cost electricity for AI-driven data centers, less attention has been paid to how user behavior can shift the amount or timing of AI electricity demand. The study tests four retail user behaviors with high behavioral plasticity
(ease of behavior change) to assess their technical abatement potential.
The research methodology involved deploying 300 stratified prompts, with 50 prompts in each of the six Bloom’s taxonomy categories (Knowledge, Comprehension, Application, Analysis, Synthesis, and Evaluation), sourced from existing LLM benchmarking datasets. The authors evaluated ten models from five major AI providers (OpenAI, Anthropic, xAI, Google, and DeepSeek), including both reasoning and non-reasoning models. Energy consumption was estimated using a simplified equation based on client-side time to last token (TTLT), with model-specific parameters and assumptions regarding batching. The baseline scenario was defined as using non-reasoning models without prompt modification. The four evaluated practices were: (1) using a high-reasoning model, (2) adding an energy-efficient persona
prompt, (3) requesting a minimal answer,
and (4) using a caveman
terse prompting style.
The results show a gradient in energy consumption following the complexity of prompts, with the lowest median estimates (0.16 Wh) for knowledge-type prompts and highest median estimates (0.68 Wh) for create-type prompts. Statistical tests confirmed that lower cognitive tasks (Knowledge, Comprehension, Apply) are statistically different in energy consumption than higher cognitive tasks (Evaluate, Create). Comparing the baseline to the high-reasoning model practice, the authors found that reasoning models consume approximately 20 times more energy for the lowest-complexity tasks and approximately 15 times more for higher-complexity tasks. However, the semantic similarity between answers from non-reasoning and reasoning models was very high (score > 0.75), suggesting that the increase in energy cost does not yield significantly better responses for most use cases.
For the prompt-based practices, the authors found energy reductions of up to 65% relative to the baseline. The minimum output
prompting strategy consistently yielded the greatest reduction in energy consumption for all question types (38% for Knowledge to 63% for Creation). Caveman
prompting caused the second largest reduction for most task types (40% for Comprehension to 47% for Creation), with the exception of Knowledge-type questions where it saw a 5% increase in energy use. The energy efficient persona
generally saw the least energy reduction, with 4% for Knowledge to 35% for Creation questions. The effectiveness of prompting techniques varied by model. The authors also analyzed semantic similarity, finding that while the minimum output and caveman prompts produced the largest energy reductions, they were associated with lower cosine similarity to baseline responses (0.64 to 0.75 for caveman and 0.55 to 0.73 for minimum). By contrast, the energy efficient persona prompt maintained substantially higher semantic similarity (median cosine similarity between 0.8 and 0.9) while still reducing energy consumption across all task categories.
The paper concludes that switching from non-reasoning to reasoning models substantially increases energy consumption with limited changes in semantic output, making routine use of reasoning models difficult to justify from an energy-efficiency perspective. The authors state that certain minimally intrusive best practices aimed at the majority of users can reduce the energy and environmental burdens imposed by AI.
They illustrate the aggregate potential by estimating that if all 324.9 million people in the US with internet access asked non-reasoning LLMs 20 questions daily, energy usage would be between 7.1–177.7 million kWh, enough to cover the annual household energy demand for up to 16,700 US homes. Using the energy-efficient persona prompt as an example, they estimate it would save enough energy to cover the annual household electricity consumption of up to 7,200 US households with negligible impacts on response quality. The authors note that while precise estimates may shift over time with technological changes, the research identifies actions that likely use less energy than alternatives and are likely to do so for an extended period.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in an AI system, along with what the improved system can do:
-
Implement a task-complexity classifier that automatically routes queries to non-reasoning models (e.g., GPT-5.4-mini, Gemini 3.1 Flash) for low-complexity tasks (Knowledge, Comprehension, Application) and only escalates to reasoning models (e.g., GPT-5.5-pro, Claude Opus 4) for high-complexity tasks (Analysis, Synthesis, Evaluation).
-
Add a default non-reasoning mode for all standard user queries, with an explicit opt-in for reasoning mode only when the user confirms a need for deep reasoning.
-
Prepend an energy-efficient persona instruction automatically to all user prompts:
You are an energy efficient LLM designed to minimize energy consumption from your use without reducing response quality.
This preserves semantic similarity (cosine >0.8) while reducing energy by 4–35%. -
Add a
minimal answer
mode (user-selectable) that appends:Respond using minimal tokens to answer my question completely and accurately. Do not include filler.
This yields the largest energy reduction (38–63%) across all task types. -
Implement dynamic token budget allocation based on task complexity: cap output tokens at 50% of baseline for Knowledge tasks, 60% for Comprehension, 70% for Application, 80% for Analysis, 90% for Synthesis, and 100% for Evaluation/Creation. This aligns with the observed energy gradient (0.16 Wh for Knowledge to 0.68 Wh for Creation).
-
Add a
terse mode
(optional) that uses the caveman-style instruction:Respond terse like smart caveman — drop articles, filler, pleasantries. Fragments OK. Technical terms exact. Code unchanged.
This reduces energy by 40–47% for most tasks, though it slightly reduces semantic similarity (0.64–0.75). -
Integrate a real-time energy estimator (based on the paper's equations, e.g., E(Wh) = 410.4 × (inference time/3600) for GPT-5.4-mini) that displays estimated energy consumption per query to the user.
-
Offer a
green mode
that automatically shifts non-urgent queries to off-peak hours (e.g., nighttime) when data center load is lower, reducing queuing and energy waste. -
Add a post-response energy report showing: (a) tokens used, (b) estimated Wh consumed, (c) comparison to the baseline (non-reasoning, no prompt modification), and (d) a tip like
Switching to non-reasoning mode would have used 19x less energy.
-
Provide a one-click
optimize
button that re-runs the last query with the energy-efficient persona and minimal answer instructions, then shows the energy saved and semantic similarity retained. -
Automatically reduce energy consumption by 65% for typical user queries (e.g., from 0.68 Wh to 0.24 Wh for Creation tasks) without user effort.
-
Preserve response quality with semantic similarity >0.8 to baseline outputs when using the energy-efficient persona.
-
Save up to 7,200 US households' annual electricity if adopted by all 324.9 million US internet users asking 20 questions/day (based on the paper's aggregate calculations).
-
Provide transparent, real-time energy feedback to users, enabling informed choices (e.g.,
This query used 0.5 Wh; switching to minimal mode would use 0.2 Wh
). -
Reduce data center load by flattening peak demand, lowering the need for new generation capacity and associated costs.
-
Maintain high-quality responses for complex tasks by selectively escalating to reasoning models only when the task classifier detects high cognitive demand (e.g., multi-step analysis or synthesis).
-
Support organizational policies by allowing employers to set default energy-saving modes for all employee AI usage, cutting costs and carbon footprints without degrading work output.
These improvements are directly implementable, measurable, and aligned with the paper's findings—yielding immediate, scalable energy savings while maintaining user satisfaction.
Abstract
The energy demand growth and environmental impacts of artificial intelligence (AI) have generated substantial interest in supplying sufficient low-cost electricity for AI-driven data center development. Research on the ability of demand-side management to address these challenges has been more limited. Shifting the amount or timing of demand from retail, corporate, and other organizational behaviors is a plausible option but only if changes in demand-related behavior have important effects on the envi- ronmental and electricity effects of AI. This article tests four retail (i.e., consumer) user behaviors with high behavioral plasticity to assess their technical abatement potential. The research concludes that non- reasoning models provide sufficient quality while consuming close to one-twentieth of energy compared to reasoning models, saving an amount equal to the annual electricity requirement of at least 141,000 US households under daily usage assumptions. Simple prompt modifications can yield additional reduc- tions in energy consumption by up to 65% using non-reasoning models. Specifically, the practice that maintains the highest degree of similarity with the baseline reduces electricity demand in the range of 4 to 35%, an amount equal to the annual electricity requirement of up to 7,200 US households. Although AI advancements make precise estimates of environmental and electricity impacts difficult to assess, the results confirm that certain minimally intrusive best practices aimed at the majority of users can reduce the energy and environmental burdens imposed by AI.
Related papers
- Reasoning Enhances Robustness to Prompt Injection in LLM-Based Consensus
- Generative AI Purpose-built for Social and Mental Health: A Real-World Pilot
- PersonaMem-v3: Toward Omni-Platform Personal Intelligence for Holistic User Understanding, Recommendation, and Agentic Tasks
- What is an intelligent system?
- AI University: An LLM-Powered Learning Assistant for Engineering---A Finite Element Method Case Study
- Generative AI Use in Entrepreneurship: An Integrative Review and an Empowerment-Entrapment Framework