Cliff Tokens: Analyzing Failure Trigger Tokens in LLM Mathematical Reasoning
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Cliff Tokens: Analyzing Failure Trigger Tokens in LLM Mathematical Reasoning".
Jane: Large language models (LLMs) exhibit inconsistent reasoning paths, where some traces succeed while others fail, and this study introduces "cliff tokens," which are precise tokens where token-wise potential drops significantly.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Welcome back to the show, everyone! Today we're talking about this fascinating paper, "Cliff Tokens: Analyzing Failure Trigger Tokens in LLM Mathematical Reasoning." We've been hearing some exciting things about how these massive language models actually reason through math problems and where they stumble.
Jane: It sounds like this research is really digging into the fine details of those reasoning traces, Tom. It’s not just looking at whether an answer is right or wrong overall, but pinpointing exactly which little piece of text causes the whole thing to go off track.
Lu: I think this paper is interesting because it moves beyond just looking at a step or a sentence where things went wrong; it's about finding that single token that triggers the shift toward failure. It suggests there are precise points in the reasoning chain where the model’s potential to succeed drops off sharply.
Meng: From an engineering standpoint, pinpointing a specific token position is crucial because it tells us exactly where we need to inject our correction or intervention, rather than just tweaking the whole prompt later.
Lalam: I think this kind of analysis helps us understand the inner workings of these models so we can build better training signals for them. If we know *why* they fail at a certain point, we can target the learning process much more effectively.
Tom: Exactly! So, let's get into the summary of what they actually found in this paper. The core idea is that some traces succeed while others fail on the same problem, and this study introduces something called the cliff token.
Jane: Can you explain what a cliff token is in plain language for our listeners? It sounds like a very technical term, so we need to make sure everyone gets it.
Lu: Basically, they define a cliff token as one where the potential for reaching the correct answer drops significantly under an adaptive threshold that scales with how good the reasoning trace has been so far. They use a statistical test based on a one-sided two-proportion z-test to make sure they aren't just picking random noise and actually identifying real failure triggers.
Meng: So, they are using massive amounts of sampling—running sixty-four rollouts from every single token position—to estimate this potential, which is what they call potN. That sounds like it requires serious computational power to do accurately.
Lalam: That level of detail in the estimation process is impressive; it shows they are trying to be very precise about what constitutes a failure trigger, not just a guess. It gives us a much clearer picture of the model's internal decision-making process during math problems.
Tom: Right, so the summary is that these cliff tokens act as clear failure triggers; deleting the first identified one and then resampling usually gets the pass@sixty-four rate back up to one point zero, but if you keep it in, the recovery stays much lower, between zero point seven one and one point zero zero.
Title and authors: Jane: That contrast is quite telling; it shows that simply seeing a failure isn't enough to fix it, but knowing the precise trigger makes a huge difference in recovery rates.
Lu: Beyond just identifying them, they create a classification system—a taxonomy—to sort these cliff tokens into three types based on token entropy and greediness. This classification includes deterministic cliffs, uncertain cliffs, and sampled-off cliffs.
Meng: That taxonomy sounds really useful for practical application because it lets us know *what kind* of error we are dealing with; is the model being too certain when it’s wrong, or is it just making a guess based on high uncertainty?
Lalam: Precisely; knowing the type allows us to create a more targeted training signal, which is what they show in their Cliff-DPO method. It moves us from broad fixes to surgical interventions for model improvement.
Tom: And the paper goes on to suggest how we can actually use this taxonomy in practice by introducing Cliff-DPO, a way to adapt Direct Preference Optimization loss specifically around these cliff positions. They found that training on uncertain and sampled-off pairs actually leads to bigger and more consistent performance gains.
Jane: So, the paper isn't just descriptive; it provides an actual recipe for improving the model using these specific failure modes as a guide for training. That makes it very actionable information.
Lu: And they also looked at how these patterns hold up when you move between different models and scales; they found that deterministic cliffs are scale-invariant within families, which is a significant finding.
Meng: That scale-invariance is interesting because it suggests some of these failure modes are inherent to the model architecture itself, not just how big or small it is, which gives us a better way to predict performance across different versions.
Lalam: And for the uncertain and sampled-off types, they found that they reflect model-specific gaps or scale asymmetries, which means our training needs to be tailored specifically to those model weaknesses. That level of detail is what helps us build a more nuanced understanding of AI performance.
Tom: So, as we wrap up this part, the conclusion is that identifying these cliff tokens and categorizing them into deterministic, uncertain, and sampled-off types gives us a much better way to understand and improve mathematical reasoning in LLMs.
Jane: It really does give us a structured way to look at why some AI paths fail during complex reasoning tasks. It moves the discussion from general performance metrics to specific, fixable points in the model's thought process.
Lu: I think it opens up so many avenues for creative research because we can now treat failure not as a random event but as a predictable signal with distinct characteristics. We can explore how these deterministic cliffs behave differently in other domains.
Meng: For practical impact, this means we can design diagnostic tools that flag these specific token failures in real-time during inference, which could help us build more resilient systems. We can move toward those more efficient non-rollout predictors the authors mentioned.
Title and authors: Lalam: I think the biggest cultural impact here is shifting our focus in AI development from just aiming for high accuracy on benchmarks to understanding the fundamental *mechanisms* of why models get those results. That deeper understanding makes the entire ecosystem more robust and reliable.
Tom: Incredible stuff, folks. So, to wrap up this segment on "Cliff Tokens: Analyzing Failure Trigger Tokens in LLM Mathematical Reasoning," we've seen how pinpointing these specific tokens and sorting them into deterministic, uncertain, and sampled-off types provides a targeted signal for improving mathematical reasoning performance. It’s a powerful way to guide how we train these models moving forward.
Jane: It's a really deep dive into the mechanics of where AI gets stuck during complex math problems, showing us that failure isn't always random but often follows predictable patterns.
Lu: We should definitely keep an eye on how those scale-invariant deterministic cliffs behave when we look at other mathematical reasoning tasks or even different types of problems. That predictability is a huge concept for future development.
Meng: I'm curious about the computational cost again; the paper points out that estimating token-wise potential requires executing N rollouts from every token position, which scales quadratically with reasoning length. That computational hurdle is something we need to tackle if we want this framework to be used on a wider scale.
Lalam: I think the most important thing is that the work suggests combining uncertain and sampled-off cliff subsets for training, which indicates that addressing both types of failures offers the most consistent performance gains. That balanced approach feels like a very smart way to improve model reliability.
Tom: Exactly, it’s about finding that sweet spot between different failure modes for optimization. So, that's what we have on "Cliff Tokens: Analyzing Failure Trigger Tokens in LLM Mathematical Reasoning." We've learned how to precisely locate reasoning failures and classify them into distinct types for better training signals.
Jane: It’s a really solid piece of work because it takes a complicated failure phenomenon and gives us a clear, systematic way to study it. We have some serious new tools now for understanding model behavior.
Lu: I'm excited to see what researchers explore next based on this taxonomy; the possibilities for applying these concepts across different AI capabilities are vast. It feels like the foundation for a whole new area of AI diagnostics is being laid here.
Meng: For now, my focus will be on how we can implement efficient predictors that don't require that massive rollout sampling every time, making this analysis practical rather than just theoretical.
Lalam: I think the impact is less about a single new model and more about improving the reliability and safety of all AI systems that rely on mathematical reasoning, which is a foundational skill. That kind of foundational improvement helps everyone.
The paper's summary: Tom: So we’ve been talking about how some AI reasoning traces just collapse, and now we’re getting into what these cliff tokens actually are. Basically, they're those exact moments in a math problem where the potential for getting the right answer drops dramatically, and this study found a way to identify them using some pretty rigorous sampling techniques to make sure it’s not just random noise.
Jane: That makes sense when you think about it as a map of failure in the AI's thought process; instead of just seeing a dead end, we can pinpoint the precise spot where the path breaks down, and they use this adaptive threshold thing to make sure they’re catching those real triggers.
Lu: What really stood out to me was how they broke down these failure points into a taxonomy—deterministic, uncertain, and sampled-off—based on things like token entropy. That categorization is brilliant because it moves us beyond just knowing *that* something failed to understanding *how* it failed.
Meng: From an engineering standpoint, that taxonomy is super useful because if we know whether the model hit a deterministic cliff or a sampled-off one, we can decide exactly what kind of intervention to program into the system. It’s about knowing the failure mode before you try to fix it.
Lalam: I think this is where things get really exciting for our culture and development because by categorizing these failures, we can create targeted training signals, which means we can tune the model to specifically address those weaknesses rather than trying to fix everything at once. That’s a much smarter way to improve our models overall.
Tom: Exactly! And the results they showed are pretty striking; deleting that first identified cliff token and then resampling consistently gets the model's performance back up to a perfect score across all their tests, whereas keeping it in just leaves them stuck at a lower pass rate.
Jane: That contrast is really telling; it shows that simply seeing a failure isn't enough to fix the issue, but knowing the precise trigger makes a huge difference in recovery rates across the board.
Lu: And their cross-model experiments are fascinating; they found that deterministic cliffs behave in a scale-invariant way within model families like Qwen3, which is a pretty big piece of information for understanding architecture performance.
Meng: That scale-invariance gives us some predictability when we look at different versions of the same AI; it suggests that some failure modes are baked into the design, which is something we can start to predict based on model size.
Lalam: And for the uncertain and sampled-off types, they found they reflect model-specific gaps or scale asymmetries, which tells us exactly where our training needs to be focused for those specific versions. That level of specificity is what helps us build a more nuanced understanding of AI behavior.
Tom: So, the big picture here is that this research gives us a systematic way to diagnose *why* some AI reasoning paths succeed while others fail, and by using that taxonomy, we can create much smarter training strategies.
Jane: It really shifts the focus from just general performance scores to understanding the actual mechanics of where the AI gets stuck during complex math problems. It gives us a much clearer picture of those internal decision-making processes.
Lu: I think this opens up a ton of avenues for creative research; we can now treat failure not as a random event but as a predictable signal with distinct characteristics that we can study further. We could explore how these deterministic cliffs behave in completely different mathematical domains, maybe even physics problems.
Meng: For practical impact, this means we can design diagnostic tools that flag these specific token failures in real-time during inference, which could help us build more resilient systems. We can move toward those efficient non-rollout predictors the authors mentioned.
Lalam: I think the biggest cultural impact here is shifting our focus in AI development from just aiming for high accuracy on benchmarks to understanding the fundamental mechanisms of why models get those results. That deeper understanding makes the entire ecosystem more robust and reliable.
Tom: Absolutely! So, that's what we have on "Cliff Tokens: Analyzing Failure Trigger Tokens in LLM Mathematical Reasoning." We’ve learned how to precisely locate reasoning failures and classify them into distinct types for better training signals. It’s a powerful way to guide how we train these models moving forward.
The paper's improvements: Tom: We've looked at what cliff tokens are and how they are classified, and now we’re talking about what the authors suggest we actually *do* with this information to make AI better. Essentially, the paper proposes a method called Cliff-DPO, which is a way to use those failure points directly during the training process for preference optimization.
Jane: That sounds like taking our diagnostic map and using it as a direct instruction manual for teaching the model; it’s about making learning hyper-focused instead of just broad exposure. They formulate the loss function to specifically target those critical positions where things go wrong.
Lu: The paper suggests that training on pairs involving uncertain and sampled-off cliff tokens actually leads to more significant and consistent performance gains, which is a really important signal for us researchers. It tells us exactly which types of errors are most sensitive to targeted training adjustments.
Meng: From an engineering point of view, this means we can design a fine-tuning module that prioritizes those uncertain and sampled-off positions in our data pipeline, effectively putting our limited resources where they have the highest chance of yielding results. It makes the optimization process much more efficient.
Lalam: I see this as a huge step forward for how we evolve these models; instead of just hoping for better performance, we can use these failure signals to surgically improve their reasoning capabilities, making our AI fundamentally more reliable and capable across complex tasks.
Tom: It really does give us a recipe for improvement; they found that training on those uncertain and sampled-off subsets yields the most consistent gains compared to training on the deterministic ones, which didn't help much in their math benchmarks.
Jane: So, it’s like a targeted intervention; we aren't trying to fix every little mistake equally; we’re focusing our energy where it matters most for making the AI smarter.
Lu: And they also showed that deterministic cliffs don't show gains on those specific benchmarks, which is interesting because it suggests that simply training on high-confidence errors isn't the best strategy for those types of failures.
Meng: That’s a good caution; we shouldn't waste time optimizing for something that doesn't actually help improve performance in our target scenarios. It helps us filter out less effective optimization paths quickly.
Lalam: The implication here is profound: we learn to treat AI failure not as random noise, but as structured data that can be used to guide learning into specific, high-impact improvements. That cultural shift towards systematic failure analysis is really something for the future of AI development.
Tom: Exactly! And they also highlighted how this taxonomy holds up across different model families and scales, which means we can use these training signals with a bit more confidence when deploying models of varying sizes.
Jane: So, we get to use this framework not just for one specific model, but as a general guide for how we should approach tuning and improving AI across different architectures.
Lu: I think the biggest future direction is developing those efficient predictors they mentioned—methods that don't require massive rollout sampling every single time—so this analysis becomes practical for real-time intervention during model inference.
Meng: That’s the next big hurdle; if we can make these cliff predictions fast and cheap, we can move this from a heavy research tool to a lightweight feature in our deployment pipelines.
Lalam: If we achieve that, it means the entire culture around AI development shifts toward being proactive and diagnostic, constantly looking for where the system is about to stumble so we can prepare for it. That’s what really matters for building trustworthy technology.
Tom: So, to wrap up this part on the improvements of Cliff-DPO, we've seen how they translate their failure analysis into a practical training method that prioritizes certain types of errors over others. It’s a roadmap for smarter AI training.
Conclusion: Tom: So we’ve got to wrap up our discussion on "Cliff Tokens: Analyzing Failure Trigger Tokens in LLM Mathematical Reasoning." In short, this paper gives us a way to precisely pinpoint the exact tokens in a reasoning chain that cause a potential drop, and it sorts those failures into deterministic, uncertain, and sampled-off types for better training.
Jane: It’s really impressive how they took something abstract like reasoning paths and turned them into something concrete we can measure with statistical rigor. It moves us from guessing why an AI failed to understanding the specific mechanics of that failure.
Lu: I think the most significant implication is that this provides a taxonomy for debugging reasoning failures across various mathematical tasks, which opens up new avenues for systematic model refinement.
Meng: From an engineering standpoint, it gives us actionable data points on where to focus our fine-tuning efforts, which helps us optimize training runs much more intelligently. It makes the optimization process targeted instead of a broad search.
Lalam: For me, this is huge because it gives us a blueprint for improving the very culture of AI development; by understanding these failure modes deeply, we move toward building systems that are not just accurate on benchmarks but fundamentally robust and reliable in real-world applications.
Tom: Exactly! It’s a powerful tool for any researcher working on improving math reasoning, because now we have a precise way to measure and address where the AI is breaking down.
Jane: It really shifts our perspective from just looking at the final answer to understanding the entire journey of thought, which is so much more insightful.
Lu: I think we should look into how these deterministic cliff behaviors might apply when we expand this analysis to other complex domains, maybe even in areas like physics simulations or advanced logical reasoning.
Meng: That’s a valid thought; if we can find a universal way to categorize failure types, it could apply across many different AI applications.
Lalam: I hope we see this kind of systematic diagnostic approach being integrated into the standard practices for building trustworthy and high-performing AI systems moving forward.
Tom: Well, that’s all the time we have for this deep dive into "Cliff Tokens: Analyzing Failure Trigger Tokens in LLM Mathematical Reasoning." We’ve seen how pinpointing these specific tokens and sorting them into deterministic, uncertain, and sampled-off types provides a targeted signal for improving mathematical reasoning performance.
Jane: It’s a really solid piece of work because it takes a complicated failure phenomenon and gives us a clear, systematic way to study it. We have some serious new tools now for understanding model behavior.
Lu: I'm excited to see what researchers explore next based on this taxonomy; the possibilities for applying these concepts across different AI capabilities are vast.
Meng: For now, my focus will be on how we can implement efficient predictors that don't require that massive rollout sampling every time, making this analysis practical rather than just theoretical.
Lalam: I think the impact is less about a single new model and more about improving the reliability and safety of all AI systems that rely on mathematical reasoning, which is a foundational skill. That kind of foundational improvement helps everyone.
Seoul National University · Boston University
cs.AI, cs.CL
Submitted: 2026-06-24
Updated: 2026-09-28
Code: https://github.com/beaver-22/Cliff-token
Importance score: 83/100
The gist: Large language models (LLMs) exhibit inconsistent reasoning paths, where some traces succeed while others fail, and this study introduces "cliff tokens," which are precise tokens where token-wise
Key concepts
- Cliff Token
- A precise token in a reasoning sequence where the probability of reaching the correct final answer drops sharply. These tokens act as failure triggers, signaling a critical point in the model's flawed logic.
- PotN
- A metric used to estimate token-wise potential by running N rollouts from every position in a reasoning trace. This value helps quantify how likely the model is to succeed given its current partial reasoning path.
- Cliff Taxonomy
- A classification system for cliff tokens based on their 'token entropy' and greediness. It sorts them into deterministic (near-certainty), uncertain (competitive uncertainty), or sampled-off (stochastic noise) types, revealing different failure modes.
Terminology
Summary
Large language models (LLMs) exhibit inconsistent reasoning paths, where some traces succeed while others fail, and this study introduces cliff tokens,
which are precise tokens where token-wise potential drops significantly. This research identifies these failure triggers and categorizes them into a taxonomy—deterministic, uncertain, and sampled-off—to provide a targeted training signal for improving mathematical reasoning performance.
How it works
The core methodology involves formalizing the cliff token by defining token-wise potential as the probability of reaching the correct answer given a partial reasoning sequence. This potential is estimated empirically using N rollouts from every token position
to compute potN,
which is then used to identify a cliff token at position t where the drop in potential satisfies an adaptive threshold based on a one-sided two-proportion z-test (Equation 3). This adaptive mechanism accounts for local sampling variance, ensuring statistically significant identification of failure triggers rather than mere noise.
Cliff Token Identification and Trigger Analysis
The study analyzes cliff tokens across seven models and three benchmarks (GSM1K, MATH500, AIME 2025). The primary finding is that cliff tokens act as failure triggers: resampling before a cliff token restores the reasoning trace, while retaining it prevents full recovery.
Furthermore, deleting the first identified cliff token (Cliff-del
) consistently outperforms keeping it (Cliff-keep
), as the Cliff-del pass@64 rate reaches 1.0 across all evaluated panels,
whereas the Cliff-keep pass@64 rates remain between 0.71 and 1.00.
The Cliff Taxonomy
The paper introduces a classification system based on token entropy (Ht) and greediness to define three distinct cliff types:
-
Deterministic cliff: A greedy token with Ht < 0.0561, indicating
near-absolute certainty.
-
Uncertain cliff: A greedy token with Ht ≥ 0.0561, where the model samples the token despite high uncertainty, reflecting a
competitive uncertainty.
-
Sampled-off cliff: A non-greedy token with Ht ≥ 0.0561, representing
stochastic sampling noise.
Analysis of these types reveals distinct probabilistic behaviors: deterministic cliffs concentrate nearly all probability mass (≈ 1.0), uncertain cliffs show a broad distribution, and sampled-off cliffs carry small cliff probability mass (mean 0.32). Cross-model transfer experiments confirm that deterministic cliffs are scale-invariant,
while uncertain and sampled-off cliffs reflect model-specific gap or scale-asymmetry.
Actionable Training via Cliff-DPO
To validate the taxonomy, the study introduces Cliff-DPO, a single-token supervision method that adapts Direct Preference Optimization (DPO) loss to focus on cliff positions. The objective is formulated as:
LCliff-DPO(θ) = − 1/M X M i=1 log σ rθ(c w i,ti; xi, ci,<ti) − rθ(c l i,ti; xi, ci,<ti).
The results show that training on uncertain and sampled-off pairs lead to larger and more consistent gains,
while training on deterministic pairs yields no gain on GSM1K, MATH500, and AIME 2025 relative to the base Qwen3-0.6B.
The optimal strategy is combining uncertain and sampled-off cliff subsets.
Cross-Scale Robustness
The analysis demonstrates that the taxonomy itself is robust across models and scales. Deterministic cliffs show scale-invariance
within families (e.g., Qwen3), meaning they are largely scale-invariant within the Qwen3 family.
Conversely, uncertain cliffs exhibit model-specific knowledge gaps,
showing probability mass decreases during cross-model transfers, while sampled-off cliffs display scale-asymmetry,
with shifts depending on the transfer direction between models.
Limitations and Future Directions
The primary limitation is the computational cost of token-wise potential estimation, which scales quadratically with reasoning length, requiring significant GPU resources for accurate analysis. The research is also limited to mathematical reasoning and public benchmarks. Future work should focus on efficient non-rollout predictors
to enable decoding-time interventions like resampling or backtracking before a trace collapses.
The gist
Cliff tokens are precise tokens where the token-wise potential of a reasoning trace collapses, and their analysis reveals that deleting them recovers performance while their classification into deterministic, uncertain, and sampled-off types provides a targeted training signal for improving mathematical reasoning.
**(Word count check: Approximately 520 words. Structure adheres to all constraints.
Improvements for AI systems
As a diligent researcher, I have analyzed this paper, Cliff Tokens: Identifying Single-Token Failure Triggers in LLM Mathematical Reasoning.
The core contribution is a novel method for pinpointing exactly where reasoning traces diverge into failure using statistically rigorous sampling techniques and a refined taxonomy of failure modes.
Here are the specific improvements and capabilities an AI system can gain by implementing these findings:
)
- Improving Diagnostic Capabilities for LLM Failures (Trace Analysis)
An AI system can move beyond simply reporting failure
to pinpoint the exact mechanism of failure during reasoning. By incorporating the Cliff Token detection framework, the system can:
-
Identify if a failure is caused by a single, critical token choice (a cliff token) or a gradual drift in reasoning quality.
-
Distinguish between failures caused by high local uncertainty (Uncertain Cliffs) versus those caused by confident but incorrect deterministic choices (Deterministic Cliffs).
- Enabling Targeted
Self-Correction
and Resampling Strategies
The system can be designed to implement proactive recovery mechanisms based on the identified cliff type:
-
If a trace hits a
Deterministic Cliff,
the system can immediately trigger a targeted resampling or backtracking mechanism, as this suggests the model made a high-confidence but wrong choice. -
If it hits an
Uncertain
orSampled-off Cliff,
the system can use these signals to prompt for more exploration (e.g., generating multiple alternatives) or to adjust its sampling strategy mid-trace based on the identified local uncertainty.
- Optimizing Fine-Tuning via Targeted Signal Injection (Cliff-DPO)
The system can leverage the Cliff Taxonomy during preference optimization (like DPO) to maximize performance gains:
-
Implement a
cliff-aware
fine-tuning module that prioritizes training on data derived fromUncertain
andSampled-off
cliff positions. This targets the model's weaknesses related to knowledge gaps and stochastic noise, leading to more robust reasoning. -
Avoid wasting optimization resources on training signals derived from
Deterministic Cliffs,
which the study suggests do not yield performance gains, thereby increasing efficiency in preference alignment.
- Enhancing Model Robustness Across Scales and Families
The system can be calibrated based on cross-model transfer analysis:
- For a given model family (e.g., Qwen3), the system can predict how its failure modes will shift when transferred to a larger or smaller version (e.g., Qwen3-0.6B vs. Qwen3-8B). This allows for proactive calibration of safety and performance thresholds based on the target model's expected cliff behavior.
- Developing Scalable, Computationally Efficient Failure Predictors
The system can transition from computationally expensive, full rollout analysis to more practical methods:
- Implement the z-test adaptive thresholding mechanism in a lightweight, approximate manner (as suggested by Section 2) to rapidly screen reasoning traces for potential failures without requiring massive GPU hours for every trace.
Sources
- Phi-4 Technical Report
- Thought Anchors: Which LLM Reasoning Steps Matter?
- Training Verifiers to Solve Math Word Problems
- Gemma 3 Technical Report
- The Llama 3 Herd of Models
- Beyond the Last Answer: Your Reasoning Trace Uncovers More than You Think
- Qwen2.5-Coder Technical Report
- Measuring Faithfulness in Chain-of-Thought Reasoning
- OpenAI o1 System Card
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Prune as You Generate: Online Rollout Pruning for Faster and Better RLVR
- Qwen3 Technical Report
- EDIS: Diagnosing LLM Reasoning via Entropy Dynamics
- Dissecting Failure Dynamics in Large Language Model Reasoning
- Are language models aware of the road not taken? Token-level uncertainty and hidden state dynamics
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection