Think Right: Learning to Mitigate Under-Over Thinking via Adaptive, Attentive Compression

arXiv:2510.01581 · cs.LG, cs.AI, cs.CL · Submitted 2025-10-02 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Think Right: Learning to Mitigate Under-Over Thinking via Adaptive, Attentive Compression".

Jane: Recent advancements in language models have enabled complex reasoning tasks, but their performance is often limited by an inability to regulate their reasoning length appropriately,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So we’re looking at the paper titled "Think Right: Learning to Mitigate Under-Over Thinking via Adaptive, Attentive Compression," written by Joykirat Singh, Justin Chih-Yao Chen, Archiki Prasad, Elias Stengel-Eskin, Akshay Nambi, and Mohit Bansal. These authors seem to have really dug into the nuances of reasoning length control.

Jane: Those names sound like a solid team tackling a complex problem; I think their work is focused on showing how we can train models to be more nuanced in their thinking process rather than just brute-forcing computation.

Lu: The authors are clearly looking at the tension between producing accurate, deep reasoning and being computationally efficient, which is a crucial balance for any large system we build.

Meng: I wonder what specific architectural tweaks they made to get this adaptive compression working online during post-training RL, because that sounds tricky to implement without introducing instability.

Lalam: I think the authors’ goal is really about making the AI's internal process more reflective of human problem-solving—knowing when to stop exploring and when to keep digging deeper.

The paper's summary: Tom: To summarize, this paper introduces TRAAC, which is an online post-training Reinforcement Learning method specifically designed to fix that under-adaptivity issue by dynamically allocating a reasoning budget based on how hard the problem seems.

Jane: It’s smart because it learns to prune redundant steps during the model's long reasoning trajectory by using self-attention, which means it identifies what parts of its thinking are actually important and keeps those while cutting out the noise.

Lu: The mechanism involves estimating task difficulty from N rollouts, segmenting the trajectory with control tokens like "wait," and then calculating an importance score for each token based on attention from the final delimiter.

Meng: So, if I understand correctly, they are essentially teaching the model to be selective about which steps it commits its computational resources to during inference by learning this compression strategy beforehand.

Lalam: That ability to compress reasoning while maintaining high accuracy sounds like a massive win for efficiency, and it could really mean our AI systems run much faster without losing quality on tough tasks.

The paper's improvements: Tom: The improvements they propose center around using a three-part reward system during Reinforcement Learning training: correctness reward, a format reward for those structural tokens, and a length reward that smooths out how verbose the model gets based on the problem's estimated difficulty.

Jane: That multi-objective approach is clever because it forces the model not just to get the right answer but also to follow good reasoning structure while respecting its allotted thinking time.

Lu: The difficulty-adaptive compression strategy itself is key; for easy problems, it aggressively compresses once the answer is found, while for harder problems, it maintains a lower compression rate to allow more exploration.

Meng: I see how that dynamic rate adjustment directly addresses the trade-off they identified: sacrificing accuracy on hard problems versus wasting tokens on easy ones. That fine-tuning of the compression degree sounds like where the real practical gains happen for efficiency.

Lalam: If this works as described, it means we can expect models to be much more efficient across different types of tasks, which is something we need when deploying these systems widely.

Conclusion: Tom: So, to wrap up on "Think Right: Learning to Mitigate Under-Over Thinking via Adaptive, Attentive Compression," the paper successfully proposes TRAAC as an online RL method that uses attention-based compression guided by difficulty estimation to balance accuracy and efficiency.

Jane: It really shows how we can engineer models to be more contextually aware of the complexity of a problem they are facing, which is a big step forward in controlling their internal cognitive process.

Lu: The implication here for the field is that we can move toward reasoning systems that aren't just scaling up raw compute but are optimizing the *way* they use that compute based on task demands.

Meng: From an engineering standpoint, achieving an eight point four percent absolute improvement in accuracy while cutting reasoning length by nearly thirty-seven percent across benchmarks like AIME and AMC is a very solid metric for deployment planning.

Lalam: I think this work suggests a future where AI agents can genuinely adjust their level of depth instantly, making them incredibly versatile tools rather than just one-size-fits-all processors.

UNC Chapel Hill University of Texas at Austin Microsoft Research

cs.LG, cs.AI, cs.CL

Submitted: 2025-10-02

Updated: 2026-09-30

Code: https://github.com/joykirat18/TRAAC

Importance score: 90/100

The gist: Recent advancements in language models have enabled complex reasoning tasks, but their performance is often limited by an inability to regulate their reasoning length appropriately, leading to either

Key concepts

Under-adaptivity
This is when language models fail to change how much they think based on whether a problem is easy or hard. They either overthink simple tasks, wasting time, or underthink complex ones, leading to errors.
Adaptive Compression Module
This core mechanism uses the model's attention scores across its reasoning steps to score each step's importance. Steps with low importance are then removed (pruned) to create a shorter, more focused reasoning path.
Difficulty-Level Calibration
TRAAC estimates problem difficulty by checking the proportion of correct answers during training rollouts. This estimate is then used to set the compression rate: high compression for easy problems and low compression for hard ones.

Terminology

Summary

Recent advancements in language models have enabled complex reasoning tasks, but their performance is often limited by an inability to regulate their reasoning length appropriately, leading to either underthinking on hard problems or overthinking on easy ones. This paper addresses this issue by proposing TRAAC (Think Right with Adaptive, Attentive Compression), an online post-training Reinforcement Learning method designed to mitigate under-adaptivity by dynamically allocating a reasoning budget commensurate with task difficulty. By leveraging the model’s self-attention over a long reasoning trajectory, TRAAC learns to prune redundant steps adaptively, resulting in improved accuracy and significant efficiency gains across diverse benchmarks.

TRAAC Framework and Motivation

The core problem addressed is under-adaptivity, where models fail to modulate their response length appropriately given problems of varying difficulty. This manifests as underthinking on hard problems, causing errors, or overthinking on simple tasks, which inflates test-time computation and reduces efficiency. TRAAC aims to strike a balance between these two extremes by learning to allocate reasoning budget based on problem difficulty. The paper highlights the trade-off: underthinking for hard problems leads to accuracy loss, while overthinking for easy problems wastes tokens. TRAAC is introduced as an online post-training RL method that achieves this balance by adapting compression based on estimated task difficulty during training.

Adaptive, Attentive Compression Module

The central mechanism of TRAAC is its attention-based compression module designed to identify and remove redundant reasoning steps. The process involves several key steps:

  1. The model generates N rollouts, and the task difficulty d is estimated from these rollouts as the proportion of correct answers among the N samples.

  2. The reasoning trajectory is segmented into steps using special control tokens (e.g., “wait”, “alternative”).

  3. For each token, an importance score is defined as the aggregated attention from the delimiter across all layers and heads to that token: sj = 1/LH X Ll=1 X H h=1 α (l,h) →tj.

  4. Steps with lower importance scores are pruned, yielding the compressed reasoning trajectory, where Steps with lower importance scores are pruned, yielding the compressed reasoning trajectory rcomp.

Difficulty-Level Calibration and Reward Shaping

To ensure adaptivity, TRAAC incorporates difficulty estimation into both compression strategy and reward calculation. The method employs three distinct reward signals during GRPO training:

  1. Correctness Reward (CR): A high-weight reward for producing the correct final answer.

  2. Format Reward: Ensures the presence of special delimiter tokens like " and ".

  3. Length Reward (LR): This reward regulates verbosity by penalizing unnecessary length while adapting to difficulty using a sigmoid-based smoothing mechanism to provide a soft bonus for rollouts beyond the median length.

Difficulty-Adaptive Compression Strategy

The degree of compression applied is dynamically adapted based on the estimated task difficulty, which is categorized as easy, medium, or hard based on the pass rate during rollout. The strategy dictates:

  1. For easier problems, a higher compression rate to aggressively compress once the correct final answer is reached.

  2. For harder problems, TRAAC maintains a low compression rate, allowing the model to extend its reasoning trajectory.

  3. To maintain stability, the system calculates the uniformity of attention score distribution; if it is close to uniform (indicating no step stands out), the compression rate is reduced to avoid removing potentially useful steps.

Performance and Generalization

TRAAC was evaluated on a variety of benchmarks, including AIME, AMC, GPQA-D, BBEH, and OptimalThinkingBench (OTB). The results demonstrate that TRAAC consistently adapts to problem difficulty: yielding improvements in efficiency on simple tasks and stronger accuracy on complex tasks. Across various OOD tasks like GPQA-D and BBEH, TRAAC shows a generalizable compression strategy. Specifically, across AMC, AIME, GPQA-D, and BBEH benchmarks (Table 1), TRAAC (Qwen3-4B) achieves an average absolute improvement of 8.4% in accuracy while a relative reduction in reasoning length of 36.8% compared to the base model. Furthermore, TRAAC outperforms other baselines like AdaptThink by achieving a 26% gain on Qwen3-4B, underscoring its ability to adaptively allocate token budgets based on problem difficulty.

Conclusion

TRAAC is a post-training RL method that successfully mitigates under-adaptivity through an online, difficulty-adaptive, attention-based compression module.

Improvements for AI systems

Here are specific, actionable improvements for an AI system based on the TRAAC (Think Right with Adaptive, Attentive Compression) framework:


  1. The AI system will implement a dynamic reasoning budget allocation mechanism that adapts in real-time to the inherent difficulty of the prompt or query. Instead of a fixed maximum token limit, it will use task-difficulty estimation (derived from rollout pass rates) to dynamically adjust its thinking capacity during inference.

  2. The system will employ an attention-based pruning module that evaluates the importance score of every reasoning step based on its aggregated attention weight from the final delimiter token (""). Steps with low calculated importance scores will be pruned, ensuring only relevant material contributes to the final output.

  3. The training process will utilize a Group Reward Policy Optimization (GRPO) framework that optimizes for a multi-objective reward system combining:

  4. The primary objective of correctness (high weight).

  5. A format reward ensuring proper use of structural tokens ("", "").

  6. A difficulty-adaptive length reward that penalizes unnecessary verbosity on easy problems while allowing extended reasoning on hard ones, as determined by the problem's estimated difficulty bin.

  7. The improved AI system can perform complex mathematical and logical reasoning tasks with significantly higher accuracy (up to 8.4% absolute gain on benchmarks like AIME and AMC) while simultaneously achieving substantial efficiency gains (up to a 36.8% reduction in reasoning length).

  8. It will demonstrate superior performance across diverse, out-of-distribution datasets (like GPQA-D and BBEH), indicating strong generalization capabilities that transfer learned compression strategies from math domains to non-math reasoning tasks.

  9. The system will exhibit optimized behavior on Optimal Thinking benchmarks, specifically avoiding both the overthinking trap (wasting tokens on simple queries) and the underthinking pitfall (failing to explore complex solutions), resulting in higher F1 scores by balancing accuracy and efficiency across all difficulty levels.

  10. The system will be more robust against computational constraints, maintaining high performance even when scaled to larger context windows or longer reasoning trajectories, as evidenced by its scalability across training and test-time length expansions (e.g., up to 15k tokens).

Sources

Related papers