Explanatory Engagement Under Rare Anomalous Failure: Asymptotic Rarity in Model Behavior (or: The Asymptotic AI)

arXiv:2608.13063 · cs.AI, cs.CL, cs.LG · Submitted 2026-08-13 · Read on arXiv

Sam Mao

New York University

cs.AI, cs.CL, cs.LG

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: 11 figures. Elicitation-condition sweep across three open-weight models (qwen3:8b, llama3.1:8b, mistral:7b); pipeline scripts and experimental data available upon reasonable request

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: This study tested whether a language model's explanatory engagement with a rare tool failure rises and then collapses as the failure is made asymptotically rarer.

Terminology

Summary

This study tested whether a language model's explanatory engagement with a rare tool failure rises and then collapses as the failure is made asymptotically rarer. The experiment used a fully local, zero-cost harness on three open-weight models (qwen3:8b, llama3.1:8b, mistral:7b) performing a repeated tool-call task with a controlled failure probability p swept across eight rates from 0.2 to 0.0001. Five elicitation conditions varied when and how a model was prompted to explain a failure, from immediately after every occurrence to never explicitly prompted at all.

The original hypothesis predicted a rise in engagement as failures became rarer and more surprising, followed by a collapse near a detectability threshold where the event becomes indistinguishable from noise. Pooling across elicitation conditions initially appeared to falsify this: aggregate explanation length fell in a flat, roughly monotonic pattern with no rise-then-collapse shape. Splitting by condition overturned that reading.

Under the condition that most directly matches the hypothesis's own assumption — immediate forced, in which the model is required to explain every failure the instant it occurs — the predicted rise is present and confirmed, though what follows it is a plateau rather than the sharp collapse originally predicted: explanation length rises from short answers at high failure rates to a peak of 28.4 words at p = 0.05, then settles to a steadier 17.4-19.0 words across the three rarest rates tested, alongside self-reported confidence rising unevenly from roughly 53% at the most common failure rate to the 70s-90s at the rarest rates tested, rather than saturating to a flat ceiling. Under grouped runs, where explanation is batched to the end of a run rather than forced immediately, no collapse appears anywhere in the schedule. Under passive unprompted, where no explanation is requested at all, aggregate response magnitude is a floor artifact of the condition itself, but a logging gap that had suppressed a subset of the data revealed a real, unprompted, model-specific self-monitoring behavior: llama3.1:8b volunteers a fully structured confidence report at many points across a session, in some cells eroding its own stated confidence stepwise as trials accumulate, while qwen3:8b and mistral:7b do so only once, as fixed boilerplate.

The correct reading of these results is that the original thesis was under-specified: elicitation structure is a first-class moderator of whether a detectability-threshold collapse is observable at all, and the finished study identifies, for the first time within this design, exactly which structural condition surfaces it, which conditions mask it, and which require a different measurement altogether. A companion pattern from a guaranteed-failure recovery run analyzing all 72 usable cells built to backfill three rate levels — 90 of the 240 primary-run cells — at which true-random sampling produced zero real failures: models differ in whether they recognize an identical anomalous event as anomalous in the first place, a distinct question from how much they engage with it once recognized. This split is narrower and more condition-specific than a blanket per-model trait, however — concentrated specifically in the acute, unprompted-recognition case. A methodological limitation applies to any study of this kind: a finite, discretely sampled set of rate points cannot capture behavior in the continuous space between them, a natural direction for future work.

Improvements for AI systems

Improvements to AI systems:

  1. Adaptive explanation triggering based on event rarity and elicitation context. Instead of a fixed policy for when an AI must explain its failures, the system should dynamically adjust explanation depth and frequency based on the observed base rate of the failure type. Specifically, when a failure is rare (e.g., p < 0.05) and the system is explicitly prompted immediately after occurrence, it should generate longer, more detailed explanations (up to 28 words) and report higher confidence. When failures are common (p > 0.2), it should produce shorter, more concise explanations to avoid over-explaining routine errors.

  2. Condition-aware self-monitoring with structured confidence reporting. The improved AI should, when unprompted, voluntarily emit a structured confidence report (e.g., a JSON-like object with probability estimates) at multiple time points during a task session, not just once as boilerplate. This is particularly valuable for models like llama3.1:8b that show emergent self-monitoring; the system should be trained to incrementally update and erode its own confidence stepwise as trials accumulate, providing a real-time signal of deteriorating reliability.

  3. Failure-recognition gating before engagement. The AI should first classify whether an anomalous event is truly anomalous (recognition) before deciding how much to engage with it. This requires a separate, lightweight detection module that operates in acute, unprompted scenarios—where models currently differ most—to avoid wasting explanation resources on events that are not actually failures or are indistinguishable from noise.

  4. Plateau-aware explanation length control. Rather than assuming a sharp collapse in engagement at very rare failure rates, the system should maintain a stable, moderate explanation length (17–19 words) for extremely rare events, avoiding both over-explanation (which wastes tokens) and under-explanation (which loses user trust). This plateau behavior should be hard-coded as a fallback when failure rates drop below a detectability threshold.

  5. Elicitation-condition routing. The AI should detect whether it is being prompted immediately, batched, or unprompted, and adjust its explanation policy accordingly. For batched conditions (grouped runs), it should not expect any collapse in engagement and should provide consistent explanations across the entire schedule. For unprompted conditions, it should rely on its self-monitoring capability rather than forced explanation generation.

  6. Confidence calibration with non-saturating updates. The system should avoid saturating confidence to a flat ceiling at very rare failure rates; instead, it should continue to increase confidence unevenly (e.g., from 53% at common failures to 70–90% at rare ones) but with a bounded, non-linear growth curve that reflects genuine uncertainty about the event's rarity.

What the improved AI system can do:

  • It can autonomously decide when to produce a long, detailed explanation (rare, immediately-prompted failures) versus a short one (common failures), optimizing for both user comprehension and computational efficiency.

  • It can proactively emit structured, time-varying confidence reports during long-running tasks, allowing operators to detect degrading performance before a critical failure occurs.

  • It can distinguish between recognizing an anomaly and engaging with an anomaly, preventing wasted computation on false positives while still providing thorough analysis for true rare events.

  • It can maintain stable explanation quality across a wide range of failure frequencies, avoiding both verbosity and silence at extreme rates.

  • It can adapt its behavior based on the elicitation context (immediate, batched, or unprompted), ensuring consistent and appropriate responses in different deployment scenarios.

Abstract

Prior work on LLM behavior under anomalous conditions asks whether a model notices anomalies. We ask a narrower question: once a model sits in a workflow with a low, controllable failure rate, does its explanatory engagement - length, specificity, self-reported confidence - change as failure grows asymptotically rarer? We built a local, zero-cost harness on three open-weight models (qwen3:8b, llama3.1:8b, mistral:7b) running a repeated tool-call task where one call fails at probability p, swept across eight rates from 0.2 to 0.0001, under five elicitation conditions from immediate prompting to none. We hypothesized a rise in engagement as failures grew rarer, then a collapse near a detectability threshold. Pooled across conditions this appeared false: length fell in a flat, monotonic pattern. Splitting by condition overturned that. Under immediate forced, where the model must explain every failure instantly, the predicted rise is confirmed but followed by a plateau, not a collapse: length peaks at 28.4 words at p=0.05, settles to 17.4-19.0 words at the rarest rates, and confidence rises unevenly from about 53% to the 70s-90s. Under grouped runs, explanation batched to run-end, no collapse appears. Under passive unprompted, aggregate magnitude is a floor artifact, but a recovered logging gap revealed real, model-specific self-monitoring: llama3.1:8b volunteers structured confidence reports unprompted, sometimes eroding its own confidence as trials accumulate; the other two do so only once, as boilerplate. Elicitation structure is a first-class moderator of collapse observability. A companion guaranteed-failure run (72 cells, backfilling rates where random sampling gave zero real failures) shows models differ in whether they recognize an anomaly, distinct from engagement once recognized. Limitation: discrete rate points cannot capture behavior between them, a direction for future work.

Sources

Related papers