From Caveman to Expert Analyst: Energy Consumption of Variable LLM Tasks

summary

Video file (mp4)

In short

This episode discusses a paper on how different Large Language Model (LLM) tasks consume energy, focusing on user actions rather than infrastructure. Researchers found that reasoning models use fifteen to twenty times more energy than non-reasoning models for similar answers. The hosts conclude that users can reduce AI energy use by switching to non-reasoning models and using specific prompting techniques like the 'energy-efficient persona' prompt.

Key concepts

Reasoning Models
These are AI models designed to think step by step before answering, such as GPT-five point five Pro or Claude Opus. They consume significantly more energy than non-reasoning models because they perform internal computation and hidden thinking processes before generating a response.
Non-Reasoning Models
These are faster, cheaper AI models, like GPT-five point four Mini or Gemini Flash. They use less energy for tasks because they do not perform the extensive step-by-step thinking required by reasoning models. For many everyday tasks, these models provide nearly the same quality answer.
Energy Gradient
The paper found that the simpler a task is, the less energy it requires. A simple knowledge question uses about zero point one six watt-hours, while a complex creation task uses about zero point six eight watt-hours, showing a clear difference in energy use based on task complexity.
Energy-Efficient Persona Prompt
This prompting technique tells the AI to minimize energy consumption. It was found to offer the best balance of savings and response quality, reducing energy use by four percent to thirty-five percent while preserving high semantic similarity to baseline answers.

Terminology used across episodes

This episode discusses

The paper

From Caveman to Expert Analyst: Energy Consumption of Variable LLM Tasks · Read on arXiv

Diego Manya, Ethan I. Thorpe, Ji Zhang, Myranda Shirk, Jiamian He, Angel Hsu, Michael P. Vandenbergh

University of North Carolina at Chapel Hill · Vanderbilt University · Arboretica · Vanderbilt Law School

The energy demand growth and environmental impacts of artificial intelligence (AI) have generated substantial interest in supplying sufficient low-cost electricity for AI-driven data center development. Research on the ability of demand-side management to address these challenges has been more limited. Shifting the amount or timing of demand from retail, corporate, and other organizational behaviors is a plausible option but only if changes in demand-related behavior have important effects on the envi- ronmental and electricity effects of AI. This article tests four retail (i.e., consumer) user behaviors with high behavioral plasticity to assess their technical abatement potential. The research concludes that non- reasoning models provide sufficient quality while consuming close to one-twentieth of energy compared to reasoning models, saving an amount equal to the annual electricity requirement of at least 141,000 US households under daily usage assumptions. Simple prompt modifications can yield additional reduc- tions in energy consumption by up to 65% using non-reasoning models. Specifically, the practice that maintains the highest degree of similarity with the baseline reduces electricity demand in the range of 4 to 35%, an amount equal to the annual electricity requirement of up to 7,200 US households. Although AI advancements make precise estimates of environmental and electricity impacts difficult to assess, the results confirm that certain minimally intrusive best practices aimed at the majority of users can reduce the energy and environmental burdens imposed by AI.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "From Caveman to Expert Analyst: Energy Consumption of Variable LLM Tasks".

Jane: The paper was written by Diego Manya, Ethan I. Thorpe, Ji Zhang, Myranda Shirk, Jiamian He et al. from University of North Carolina at Chapel Hill and Vanderbilt University and Arboretica and Vanderbilt Law School.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Alright, listeners, welcome back. Today we're digging into a paper that's got a title that just makes you smile: "From Caveman to Expert Analyst: Energy Consumption of Variable LLM Tasks." Jane, I gotta say, that title alone tells you these researchers have a sense of humor.

Jane: They really do, Tom. And that humor is hiding something serious underneath. The paper is about how much electricity different AI tasks use, and how we as regular users can actually do something about it. The "caveman" part comes from one of their actual prompting tricks they tested.

Tom: Right, they literally tested a prompt that tells the AI to respond like a smart caveman. Drop the articles, skip the pleasantries, just get to the point. And that's one of the energy-saving techniques they measured. But let's back up a second.

Jane: Good idea. So the team behind this is led by Angel Hsu at UNC Chapel Hill, with collaborators from Vanderbilt Law School and a few other places. They're looking at the energy problem from the demand side, not the supply side.

Tom: And that's a big deal, because most of the conversation about AI energy use has been about building more power plants and data centers. But these folks are saying, wait a minute, what if we just use less energy in the first place?

Jane: Exactly. They tested ten different commercial models from five major companies, including OpenAI, Google, Anthropic, and xAI. And they ran three hundred real user prompts through them, categorized by cognitive complexity using Bloom's Taxonomy.

Tom: Bloom's Taxonomy, that's the old education framework, right? Knowledge, comprehension, application, analysis, synthesis, evaluation.

Jane: You got it. And they found something pretty striking. The simpler the task, the less energy it takes. A simple knowledge question uses about zero point one six watt-hours, while a complex creation task uses about zero point six eight watt-hours. So there's a real gradient there.

Tom: And that matters because most of us are probably using the same AI for everything, whether we're asking for a quick fact or asking it to draft a whole strategy document.

Jane: Precisely. And the title hints at their biggest finding, which we'll get into in a moment. But the short version is, the way you ask matters, and the model you choose matters even more. Stick around, because this gets really interesting when we start talking about the actual numbers.

Summary: Tom: So we're back with "From Caveman to Expert Analyst: Energy Consumption of Variable LLM Tasks," and Jane, I want to dig into the headline finding, because it's honestly a bit shocking.

Jane: It is. The researchers compared what they call reasoning models against non-reasoning models. Reasoning models are the ones that think step by step before answering, like GPT-five point five Pro or Claude Opus. Non-reasoning models are the faster, cheaper ones like GPT-five point four Mini or Gemini Flash.

Tom: And the difference in energy use is wild. We're talking about reasoning models consuming roughly fifteen to twenty times more energy than their non-reasoning counterparts. For simple tasks, it's closer to twenty times.

Jane: But here's the kicker. When they compared the actual answers using semantic similarity, the responses were really close. For most task types, the similarity score was above zero point seven five, which means the non-reasoning model gave nearly the same quality answer.

Tom: So people are burning twenty times the electricity for answers that are basically the same? That seems like a massive waste.

Jane: For most everyday tasks, yeah. The paper suggests that unless you're doing something exceptionally complex or you need absolute precision, you probably don't need the reasoning model. And that's a huge finding because most commercial AI products default to reasoning now.

Lu: If I can jump in here, Tom. What's really interesting is that this matches what we're seeing in the broader AI efficiency research. The reasoning models are doing a lot of internal computation, generating hidden thoughts before they even start writing your answer. That's where the energy goes.

Tom: So it's not just that they write longer answers, they're thinking harder before they write.

Lu: Exactly. And the paper shows that this extra thinking doesn't always translate into better answers. The semantic similarity scores tell us the content is largely the same, just with more computational overhead.

Jane: And that has real implications for the energy grid. The paper estimates that if everyone in the US with internet access asked twenty questions a day using non-reasoning models, the total energy use would be between seven point one and one hundred seventy-seven point seven million kilowatt-hours per year.

Meng: But hold on, Jane. That range is huge. What's driving that spread?

Jane: Good question, Meng. It's the difference between models. Some non-reasoning models are much more efficient than others. Gemini three point one Flash, for example, was substantially more efficient than the other models they tested.

Meng: So model choice matters even within the same category. That's useful for someone like me who's actually deploying these systems.

Tom: And it's useful for regular folks too. The takeaway so far is, if you're using AI for everyday stuff, you can probably switch to a non-reasoning model and barely notice the difference, while cutting your energy footprint by a factor of twenty.

Jane: But there's more. Because the paper also tested prompting tricks that can cut energy even further. And that's where the caveman comes back in. We'll get into those specifics next.

Improvements: Tom: Welcome back to our look at "From Caveman to Expert Analyst: Energy Consumption of Variable LLM Tasks." We've covered the big model comparison, but now Jane, let's talk about the actual tricks people can use.

Jane: Right. The researchers tested three prompting techniques on the non-reasoning models. The first is the energy-efficient persona, where you tell the AI it's designed to minimize energy consumption. The second is asking for a minimal answer. And the third is the caveman prompt we mentioned at the top.

Tom: And the results are pretty impressive. The minimal answer prompt cut energy use by thirty-eight percent for simple knowledge questions and up to sixty-three percent for complex creation tasks.

Jane: The caveman prompt was also strong, with reductions between forty percent and forty-seven percent for most task types. But here's a wrinkle, for simple knowledge questions, the caveman prompt actually increased energy use by about five percent.

Meng: That's counterintuitive. Why would telling the AI to be terse make it use more energy?

Jane: The researchers think it's because the caveman prompt changes the structure of the response so much that the model has to work harder to reformat the information. It's not just about fewer words, it's about how the words are organized.

Tom: And the energy-efficient persona was the most interesting one to me. It gave the smallest energy reduction, between four percent and thirty-five percent, but it preserved the highest semantic similarity to the baseline answers.

Lu: That's actually a really important distinction, Tom. The persona prompt is a softer signal. It tells the model to be efficient without changing the fundamental style of the response. So you get energy savings with almost no change in what the answer looks like.

Meng: So if I'm a company rolling this out to employees, the persona prompt is the safest bet because it won't change the quality of their work product.

Jane: Exactly. And the paper makes the point that these techniques are accessible to anyone. You don't need technical expertise, you just copy and paste a sentence into your prompt.

Tom: And the aggregate numbers are meaningful. The paper calculates that if everyone in the US used the energy-efficient persona prompt, the savings would be enough to power seven thousand two hundred households for a year.

Lu: And that's with essentially no change in response quality. That's the kind of efficiency gain that actually matters in the real world.

Meng: But I want to push back a little. These are estimates based on client-side timing and theoretical power models. The actual energy consumption could vary a lot depending on the data center, the hardware, the batch size.

Jane: That's fair, Meng. The paper acknowledges those limitations. They're using simplified equations based on model parameters and inference time. But the relative differences between practices are still meaningful, even if the absolute numbers are rough.

Tom: And that's the key point. We don't need perfect numbers to know that using a non-reasoning model with a good prompt is going to use less energy than a reasoning model with no prompt. The direction is clear.

Jane: And the magnitude is large enough that it's worth acting on. We'll wrap this up with some final thoughts next.

Conclusion: Tom: So we're wrapping up our discussion of "From Caveman to Expert Analyst: Energy Consumption of Variable LLM Tasks." Jane, what's the big picture here?

Jane: The big picture is that individual users have real power to reduce AI energy consumption. The paper shows three things. First, reasoning models use fifteen to twenty times more energy than non-reasoning models, with minimal quality difference. Second, simple prompt modifications can cut energy use by up to sixty-five percent. And third, the energy-efficient persona prompt offers the best balance of savings and response quality.

Tom: And that's not just theoretical. The paper estimates real-world impacts, like the equivalent of seven thousand two hundred households' worth of electricity saved just from one prompting change.

Lu: What I find most exciting is that this opens up a whole new research direction. We're just starting to understand how user behavior shapes AI's environmental footprint. There's so much more to explore, like how different languages or cultural contexts affect energy use.

Meng: And from an engineering standpoint, these findings are immediately actionable. We can bake these prompts into our systems by default, or at least offer them as options. The infrastructure doesn't need to change, just the instructions.

Jane: And that's the beauty of it. No new hardware, no new data centers, no waiting for the grid to catch up. Just smarter habits.

Tom: The paper also makes a point about information transparency. Right now, users have almost no idea how much energy their AI queries use. Studies like this start to fill that gap.

Jane: And as the authors note, the specific numbers will change as the technology evolves. But the direction is clear, and the practices they identify are likely to remain useful for a long time.

Tom: So our listeners can take away something practical today. If you're using AI, try switching to a non-reasoning model for everyday tasks. And maybe add that energy-efficient persona prompt to your next query.

Jane: It's a small change that adds up. And with that, we're saying goodbye to this paper. Thanks for joining us, and we'll see you next time with another fascinating piece of research.

Tom: Take care, everyone.

More episodes

← Home