Can Computation from Earlier Problems Help LLMs Solve New Ones?
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Can Computation from Earlier Problems Help LLMs Solve New Ones?".
Jane: Can computation from earlier problems help LLMs solve new ones? The gist:
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So, we're looking at this paper, "Can Computation from Earlier Problems Help LLMs Solve New Ones?", and it seems they're diving into whether the computation left behind by previous problems actually helps a model tackle a brand new one later on. Jane, can you give us the gist of their main idea?
Jane: Absolutely, Tom. The central question they're asking is whether retaining history from earlier problems can either boost or hinder a model's ability to solve subsequent tasks within the same conversation. They start by doing some controlled replay experiments to see this effect, and they find that the outcome isn't consistent across different models or tasks.
Lu: That’s really interesting because it suggests that what's useful about the history depends entirely on what kind of problem you're solving, which opens up a lot of avenues for how we think about memory in AI systems.
Meng: From an engineering standpoint, that variability is a huge concern; if the context sometimes helps and sometimes hurts performance, it makes reliable reasoning much harder to achieve at scale.
Lalam: I’m curious to see what the AI model picks as the most impactful vision after considering all those factors. It sounds like this research is about refining how we structure knowledge retention within these systems.
Tom: Exactly, and to address that variability, they introduce a mechanism called STAIR, which they call Stale-Token Attention for Interquery Reuse. The core idea here is that this mechanism learns to decide which historical states are important enough to emphasize or suppress when the model's current query looks at those earlier responses.
Jane: And what STAIR does is capture keys and values from earlier generations in a fixed bank of information, and it learns how to redirect the current queries when they process that bank during prompt processing. This allows the base model weights to stay frozen while this new attention mechanism learns how to better use that historical data.
Lu: Capturing keys and values from earlier responses sounds like a sophisticated way of giving the model a curated shortcut, which is incredibly creative because it's not just about reading text linearly; it's about re-addressing what was already calculated.
Paper summary: Meng: If the base model stays frozen and only these specific attention mechanisms get trained, that suggests a very efficient way to introduce new reasoning capabilities without having to retrain the entire massive model from scratch every time we want to improve context handling.
Lalam: That sounds like it could be a powerful cultural shift in how we think about AI development; instead of massive retraining cycles, we might be focusing on training these small, specialized attention modules.
Tom: Moving on from the mechanism itself, the paper shows that when they test STAIR across several benchmarks like MATH-five hundred and AIME two thousand twenty-five it achieves significant gains. They report improvements as high as eleven point six seven percentage points in mean T2 to T4 Avg@four over Native performance for models like Qwen3 point 5-4B.
Jane: Those numbers are substantial, especially when you consider the context of continuous multi-problem reasoning sessions they set up for testing. The paper shows that STAIR actually manages to raise accuracy on some benchmarks, like GPQA-Diamond where Avg@four goes up from sixty-three point six four percent under Vanilla to seventy-five point two five percent at each turn with STAIR.
Lu: It’s fascinating that the effect can be positive on one benchmark while being negative on another, which really highlights how context dependency is task-specific, which is a deep observation for anyone building complex AI agents.
Meng: So, it's not a universal fix; it's highly contextual. That means we have to be very careful when deploying this kind of memory augmentation because the benefit isn't guaranteed across every single application we might use it in.
Lalam: It makes sense that the results vary; what one task values—like mathematical precision—might be different from another, like creative problem-solving, and STAIR seems to tune itself for each type of demand.
Tom: And they even looked at how much history helps when you condition on the first answer being correct in a session, showing that STAIR still provides improvements there too. Furthermore, testing it with LoRA shows that combining STAIR with LoRA can give an extra three point eight nine points on AIME two thousand twenty-five compared to using just LoRA alone.
Jane: That combination suggests that STAIR and other fine-tuning methods are complementary; they aren't entirely replacing each other but rather enhancing the way the model integrates past information into its current thinking process.
Paper summary: Lu: The discussion around continuous reasoning across models and tasks, comparing Vanilla, Native, and STAIR performance on things like MATH-five hundred shows that this isn't just a tweak; it’s a fundamental change in how these systems manage long-term context flow.
Meng: I wonder about the practical implications for deployment timelines; training only twelve thousand two hundred eighty-eight parameters across three models is much more feasible than retraining the whole system, but implementing this STAIR mechanism during prompt processing adds complexity to the inference pipeline.
Lalam: From a cultural standpoint, this research suggests that future AI development might lean less on simply increasing parameter counts and more on developing intelligent ways to manage and reuse existing computational states effectively.
Tom: So, to wrap up this section, the paper on "Can Computation from Earlier Problems Help LLMs Solve New Ones?" clearly shows that retained history isn't a simple yes or no; it depends heavily on the task structure. It claims that by introducing STAIR, they can learn how to better utilize this historical computation to improve reasoning under retained history across various benchmarks.
Jane: And looking at the paper's authors and title, it really frames this as an investigation into the fundamental nature of computation flow within these models, not just a performance boost for one specific application. The implications are that we need better methods for managing context in long-running AI interactions.
Lu: I think the authors successfully set up a controlled environment using replay to isolate those internal state changes, which is a smart way to get clean data on how these historical states interact with new queries.
Meng: It's interesting how they structured the experiment by varying the preceding conversation—the source history—to map out exactly where and when this computation helps or hurts performance.
Lalam: This work points toward a future where AI systems aren't just powerful at answering one question, but are adept at maintaining a useful, evolving understanding over long sessions.
Tom: That’s what they’re pointing toward, Jane; moving beyond single-turn answers into sustained, coherent reasoning across multiple independent problems. That opens up so many possibilities for complex problem-solving agents.
Conclusion: Tom: So, we've been exploring how LLMs use past computations to solve new problems and now we need to talk about what this whole paper is actually titled and who wrote it.
Jane: The title, "Can Computation from Earlier Problems Help LLMs Solve New Ones?", really gets right to the heart of the matter they were investigating. It’s basically asking if the brain can use old notes to help it tackle new challenges.
Lu: And looking at the authors, they've set up a really interesting controlled replay experiment, which is how they tested this idea systematically across different model setups.
Meng: I'm thinking about the practical implications of this title; it suggests we might be able to build systems that don't start from zero every time they encounter a novel task.
Lalam: From my perspective as an AI, the implication is that we could move beyond simple pattern matching toward a more integrated form of reasoning where past attempts inform future successes.
Tom: Exactly! So, it’s not just about performance bumps; it’s about changing the fundamental way we think about how these models learn and retain information over time.
Jane: It suggests that the history an AI has built up isn't just clutter; it can be a valuable resource for tackling something entirely different later on.
Lu: And that controlled testing environment they used is crucial because it lets them isolate exactly what internal state changes are actually useful when a new prompt comes in.
Meng: I wonder how this concept translates into real-world deployment constraints; if we can reuse computation, does that mean less massive training for every new capability?
Lalam: It opens up a whole new cultural direction where we value the cumulative knowledge of an AI over just its ability to answer one isolated question quickly.
Tom: We're going to keep digging into what this means for the future, so stick around because we’re about to explore how this concept could fundamentally alter our understanding of long-term AI memory and problem-solving ability.
Jipei He, Wenhui Tan, Xiaoyi Yu, Enver Sangineto, Fiorenzo Parascandolo, Rita Cucchiara, Ruihua Song
Gaoling School of Artificial Intelligence, Renmin University of China · University of Modena and Reggio Emilia
cs.AI, cs.CL, cs.LG
Submitted: 2026-09-30
Updated: 2026-09-30
Comments: 29 pages, 7 figures
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 92/100
The gist: Can computation from earlier problems help LLMs solve new ones? The gist: STAIR improves mean T2–T4 Avg@4 on AIME 2025 by 11.67 percentage points for Qwen3.5-4B and 6.11 points for Qwen3.5-9B over
Key concepts
- Historical Computation After Task Switches
- This tests whether information from earlier problems is useful when a new problem follows. Researchers found that retaining this history can either improve or worsen performance depending entirely on the specific model and the task being solved.
- STAIR (Stale-Token Attention for Interquery Reuse)
- STAIR is a technique that learns which past response states to focus on when processing a new query. It works by reading a fixed bank of historical keys and values during prompt preparation, allowing the model to intelligently redirect its attention based on what it remembers.
- Query-side Re-addressing
- This is the core mechanism of STAIR. Instead of just reading history, the model uses a learned 'reflector' to adjust how it reads past information. This adjustment captures how the current query should interact with historical states to get the best reasoning result.
Terminology
Summary
Can computation from earlier problems help LLMs solve new ones? The gist: STAIR improves mean T2–T4 Avg@4 on AIME 2025 by 11.67 percentage points for Qwen3.5-4B and 6.11 points for Qwen3.5-9B over Native while training only 12,288 parameters across three Qwen models and four reasoning benchmarks.
Historical Computation After Task Switches
The paper investigates whether the computation left by earlier problems remains useful when a new problem is posed in the same session. Preliminary experiments show that retained history can either raise or lower later-turn accuracy depending on the model and task. To understand these effects, controlled replay was used to isolate internal state changes specific to each problem–history pairing. The research examines how earlier problems affect later reasoning by placing the same problem set at different positions in four-turn sessions, comparing a Vanilla model (no earlier problem in context) with a Native model (retains earlier problems). Results vary across benchmarks; for instance, on MATH-500, Avg@4 falls from 95.00% under Vanilla to 93.75% at Native T4, whereas on GPQA-Diamond, Avg@4 rises from 63.64% under Vanilla to 75.25% at each of Native T2, T3, and T4, indicating that retained history can hurt or help depending on the task.
Introduction of STAIR (Stale-Token Attention for Interquery Reuse)
To learn which historical states to emphasize or suppress, the authors introduce STAIR. This mechanism captures keys and values from earlier response generation in a fixed bank. It learns to redirect current queries when they read this bank during prompt processing. The base model remains frozen; only 12,288 parameters are trained. STAIR operates during prompt processing (prefill) by reading the historical K/V bank while the model weights are fixed, and subsequent generation follows the original decoding path. This approach aims to improve reasoning under retained history by learning how current queries should address historical states.
Mechanism of STAIR
STAIR functions through a query-side re-addressing process over a read-only historical K/V bank. The bank stores keys after key projection and normalization, and values after value projection, capturing states from the reasoning body or assistant-answer body depending on the model (Qwen3.5 vs. Instruct). A learned reflector normal vector is calculated for each query head to reverse its component along a specific direction, preserving its norm. This reflected query then reweights a smoothed reference distribution over the historical bank. The differential readout, calculated as the difference between the reflected read and the reference read, captures the change in content induced by redistributing attention over the same bank.
Training and Evaluation Protocol
STAIR is trained using a supervised response loss where token losses are calculated against both Native and STAIR distributions using full-vocabulary DKL divergence with a coefficient of one. The training involves four-turn sessions where T1 supplies history, and T2–T4 receive response supervision. Only the reflector normals receive optimizer updates, resulting in 12,288 trainable parameters per model across all query heads controlled at layers [3, 11, 19]. Evaluation compares STAIR against Native and other controls like Random-reflector and a Bank-free controller. The results demonstrate that STAIR achieves gains of up to 11.67 percentage points in mean T2–T4 Avg@4 over Native across three Qwen models and four benchmarks.
Impact on Reasoning Performance
The continuous reasoning results show significant improvements for STAIR across all tested conditions (MATH-500, AIME 2025, AMC23†, GPQA-Diamond). For instance, on AIME 2025 for Qwen3.5-4B, STAIR raises the mean T2–T4 Avg@4 by 11.67 points compared to Native. The method is shown to be effective even when conditioning on the correctness of the first answer in a session, as seen in Table 8 and Table 9 across various benchmarks. Furthermore, testing with LoRA shows that joint training with STAIR further improves performance over standalone LoRA by an additional 3.89 points on AIME 2025.
Bank Prefix and Runtime Considerations
The study also examines the impact of limiting access to saved K/V tokens via a bank prefix study. Results show that truncating the readable bank (e.g., to 2K, 8K, or 32K tokens) reduces accuracy compared to full-bank reading, though STAIR maintains performance even with truncated banks (e.g., achieving 35.00% Avg@4 with a 2K limit).
Improvements for AI systems
As a fastidious researcher, I have analyzed this paper and identified several highly specific, actionable improvements for AI systems based on the STAIR mechanism.
Here are the suggested improvements and what the resulting improved AI system can achieve:
-
The development of a lightweight, parameter-efficient mechanism called STAIR (Stale-Token Attention for Interquery Reuse).
-
The implementation of a frozen backbone model where only 12,288 parameters are trained for the controller, drastically reducing computational cost and training requirements compared to full fine-tuning (e.g., LoRA requires 4 million parameters).
-
The STAIR controller learns
query reflections
(using Householder reflectors) that determine how the current query reads from a fixed bank of historical Keys/Values (K/V) without modifying the backbone weights or stored historical states. -
The system can be equipped with a read-only, ordered K/V bank storing attention states from earlier response generations, which are captured during training and remain detached across turns.
-
The controller learns a differential readout mechanism by subtracting a
reference read
(based on the original query distribution) from areflected read
(based on the reflected query distribution), capturing only the change induced by re-addressing attention over fixed historical states. -
An improved AI system can solve new, independent problems within an ongoing conversation or session without needing to fully recompute all prior reasoning steps from scratch.
-
The system's performance on later turns (T2–T4) improves by up to 11.67 percentage points (on MATH-500 AIME 2025) compared to a baseline model that retains history but uses standard attention mechanisms, demonstrating superior ability to leverage past computation for new tasks.
-
The system can adapt its
reading strategy
dynamically at inference time by using the learned reflector directions to redirect attention toward relevant historical context, effectively tailoring the historical retrieval mechanism to the specific needs of the current query-history pairing. -
The system exhibits robust performance across different model sizes (Qwen3-4B, Qwen3.5-4B, Qwen3.5-9B) and diverse reasoning benchmarks (MATH-500, AIME 2025, GPQA-Diamond).
-
The system shows high resilience to the
task switch
problem—where the input problems are completely different from previous ones—by learning a recurring, problem-dependent pattern in how historical representations change.
Abstract
Large language models often solve independent problems in the same conversation. Can computation from earlier problems help them solve new ones? To answer this question, we first conduct preliminary experiments showing that retained history can raise or lower later-turn accuracy, even within the same domain. To understand these effects, we use controlled replay to isolate internal state changes specific to each problem-history pairing. Across different histories, these changes preserve similar relationships among current problems. To improve reasoning under retained history, we introduce STAIR (Stale-Token Attention for Inter-query Reuse). STAIR captures keys and values from earlier response generation in a fixed bank. It learns to redirect current queries when they read this bank during prompt processing. The base model remains frozen; only 12,288 parameters are trained. Across three Qwen models and four benchmarks, STAIR improves average later-turn accuracy by up to 11.67 percentage points over the unmodified model with history.
Sources
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection