Signed Lexical Confidence for Risk-Calibrated Intent Routing

arXiv:2610.00262 · cs.CL, cs.AI · Submitted 2026-09-24 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Signed Lexical Confidence for Risk-Calibrated Intent Routing".

Jane: Selective intent routing allows an assistant to act on reliable predictions while deferring uncertain requests,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, let’s talk about the title and the authors of this paper, "Signed Lexical Confidence for Risk-Calibrated Intent Routing." It’s clear they aren't just making another confidence score; they are specifically engineering a mechanism to route requests selectively.

Jane: Exactly, Tom; it focuses on selective intent routing, meaning the assistant can choose when to act on a prediction and when to defer something uncertain. The authors are Yezhou Cheng from the University of Wisconsin–Madison, Zehua Yang as an Independent Researcher, and Bojun Lin from Pinterest Inc., which gives us a mix of academic rigor and practical application experience.

Lu: What strikes me about the author mix is how they’re bridging the gap between deep linguistic analysis—the lexical model part—and the necessary deployment considerations for real-world systems.

Meng: I'm thinking about those implications; if this gate can genuinely reduce error ranking, it means we might be able to handle more ambiguous user requests with a much higher degree of certainty than we currently allow.

Lalam: For our culture and interaction style, this means the AI could become much more discerning; instead of blindly following the highest confidence score, it could use this signed evidence to ensure its responses are always aligned with what’s most probable and appropriate.

The paper's summary: Tom: To summarize the core idea of "Signed Lexical Confidence for Risk-Calibrated Intent Routing," they propose a new method that merges the sentence classifier’s margin score with a sparse lexical model’s support for the predicted intent to create a single confidence score.

Jane: That combination is key, Tom; they define this process by first getting an intent prediction and its semantic margin, and then calculating a separate signed support score that shows whether words or characters actually back up that specific intent.

Lu: The way they define the signed support as "explicitly directional with respect to the routed label" by making lexical preference for a competing intent negative evidence is a really clever way to inject context into the scoring process.

Meng: That directional aspect is important because it means we aren't just looking for agreement; we are actively penalizing evidence pointing toward other possible intents, which should help filter out more misleading signals.

Lalam: It feels like this system moves beyond just measuring what the model *thinks* and starts measuring how much the language actually supports that specific belief, which is a much deeper level of understanding for our AI.

The paper's improvements: Tom: The main improvement they highlight is moving away from a simple "hard agreement veto" to this continuous score calculation using a regularized logistic model, which lets the system trade evidence across different regions instead of just shutting things down completely.

Jane: That continuous score, s(x), allows the system to estimate correctness while still letting it respect the base classifier’s predictions, which is a big step because it keeps that prediction intact while adding a layer of risk calibration.

Lu: The methodology involves standardizing inputs based on a disjoint gate-fitting partition and then applying this logistic model, which suggests they’ve found a way to blend these two types of evidence coherently into one manageable metric.

Meng: I like that the method separates the feature combination from the final threshold selection; it means we can tune how much weight we give to the semantic margin versus the lexical support independently before calibrating against risk.

Lalam: This refinement means when we deploy this, we aren't just setting one static rule; we’re creating a flexible confidence landscape that adapts based on how strong the evidence is for a specific request, which should lead to much more nuanced behavior.

Conclusion: Tom: So, wrapping up this discussion on "Signed Lexical Confidence for Risk-Calibrated Intent Routing," the paper successfully introduces a compact confidence gate that uses prediction-aligned signed lexical evidence to refine intent routing.

Jane: It’s an interpretable way to strengthen risk-calibrated routing by aligning sparse lexical evidence with the classifier’s predicted intent, which we saw improves error ranking across BANKING77, CLINC150, and HWU64.

Lu: The core contribution is that this method provides a continuous score that captures both magnitude and direction of the evidence without imposing a categorical veto when things are ambiguous.

Meng: From an engineering standpoint, the most compelling result is how it demonstrably improves ranking quality; we saw reductions in the area under the risk–coverage curve by about fifteen percent to twelve percent on those specific benchmarks compared to just using a semantic-only gate.

Lalam: This work shows that by incorporating this signed lexical support, our AI can achieve better reliability and higher accepted coverage while still operating under a specified error target, which is really exciting for how we build trustworthy systems.

Yezhou Cheng, Zehua Yang, Bojun Lin

University of Wisconsin–Madison

cs.CL, cs.AI

Submitted: 2026-09-24

Updated: 2026-09-24

Code: https://github.com/alexa/dialoglue

Importance score: 82/100

The gist: Selective intent routing allows an assistant to act on reliable predictions while deferring uncertain requests, and this work introduces a signed lexical gate that combines sentence classifier logit

Key concepts

Semantic Margin (ms(x))
This measures how strongly the model predicts a specific intent compared to all other possible intents. It quantifies the confidence in the predicted label by finding the difference between the logit of the predicted intent and its strongest competitor, providing a measure of semantic certainty.
Signed Lexical Support (mℓ(x))
This feature measures how well specific words or characters in an input sentence support a particular intent. It is 'signed' because if a word strongly supports a competing intent, it provides negative evidence against the routed label, making the score directional.
Two-Feature Correctness Gate (s(x))
This combines the semantic margin and signed lexical support into one confidence score using a regularized logistic model. This continuous score captures both how much evidence supports or opposes an intent, allowing for nuanced trade-offs between different pieces of evidence rather than a simple yes/no veto.
Independent Risk Calibration
This stage independently selects an operating threshold based on quantiles from a separate development partition. This ensures the system's error rate stays within acceptable bounds during deployment by explicitly controlling the risk of making incorrect selections.

Terminology

Summary

Selective intent routing allows an assistant to act on reliable predictions while deferring uncertain requests, and this work introduces a signed lexical gate that combines sentence classifier logit margin with sparse lexical model support to enhance risk-calibrated intent routing.

How it works

The proposed system integrates two distinct pieces of evidence to estimate correctness: the semantic margin and the signed lexical support. First, a frozen sentence encoder and trained linear head produce the predicted intent and its semantic margin, defined as ms(x) = zyˆ(x) − max k≠ˆy z k(x) (Equation 1). Second, a separate linear support-vector classifier measures lexical support for that same intent using word/character TF-IDF features. This results in the signed support score, defined as ml(x) = wyˆ(x) − max k≠ˆy w k(x) (Equation 2). Crucially, this makes the score explicitly directional with respect to the routed label, as lexical preference for a competing intent becomes negative evidence.

Two-feature correctness gate

These two features are then combined into a single confidence score using a regularized logistic model. On a disjoint gate-fitting partition, the inputs are standardized using that partition’s means and standard deviations. The resulting score, s(x), is calculated via Equation 3: s(x) = sigmoid β0 + X j∈ s,l βj mj (x) − µj σj . This gate allows the system to estimate correctness while preserving the base classifier’s predictions. Unlike a hard agreement veto, this continuous score captures both magnitude and direction, allowing it to trade evidence across agreement regions instead of imposing a categorical veto.

Independent risk calibration

After the two-feature gate yields a score, an independent binomial calibration stage selects an operating threshold (t). This stage is crucial for risk control. The coverage C(t) and selective error R(t) are defined based on the score s(x) relative to t. To select the optimal threshold, a second, independent development partition supplies 50 score quantiles, from quantile 0 to 0.98 in steps of 0.02. The operating threshold is then chosen by selecting the candidate with the largest Nt among those satisfying a bound derived from Eq. (5), which ensures that "Pcal[R(tˆ) > α] ≤ δ," thereby providing a calibration guarantee against population selective error.

Evaluation and results

The proposed signed lexical support feature consistently improves error ranking across all tested datasets (BANKING77, CLINC150, and HWU64). Specifically, the signed gate reduces the area under the risk–coverage curve (AURC) by 15.8%, 15.1%, and 11.8% relative to a learned semantic-only gate on these datasets. At a nominal 5% error target, this ranking advantage translates into greater accepted coverage: The signed gate raises mean coverage from 87.80% to 89.63% on BANKING77 and from 79.16% to 84.30% on HWU64. Furthermore, matched controls demonstrate that the proposed feature improves average error ranking over the tested unsigned lexicalconfidence feature, with incremental benefits varying by dataset and metric, such as adding 0.36 coverage points on BANKING77 and 2.37 points on HWU64 when combined with an agreement gate.

Calibration sensitivity and deployment

The study rigorously evaluates how the operational policy depends on the calibration procedure. Comparisons between simultaneous binomial calibration and Learn then Test (LTT) fixed-sequence testing show that the independently specified sequence starts at the development-set threshold for 20% coverage and increases coverage in two-point steps, ending with accept-all. At a strict 2% target, while the semantic gate defers every request in some runs, simultaneous binomial calibration returns a nonempty policy in all 30 dataset–seed runs at the available budgets, demonstrating that stronger confidence ranking helps assemble a sufficiently large set of reliable calibration examples. Stress tests involving orthographic perturbations and unsupported intents quantify the effect of deployment changes, showing that without recalibration, the signed-gate error can reach 14.11% under specific shifts.

Conclusion

This work introduces a compact confidence gate that uses prediction-aligned signed lexical evidence to improve selective routing for a fixed intent classifier. The resulting score provides an interpretable way to strengthen risk-calibrated intent routing by aligning sparse lexical evidence with the classifier’s predicted intent, which is shown to improve error ranking and increase accepted coverage across multiple benchmarks. The analysis also separates the contribution of confidence quality from the statistical procedure used to calibrate it, providing insights into deployment strategies.

Improvements for AI systems

Here are the specific improvements to AI systems derived from this research, detailing what the improved system can accomplish:

  1. The proposed system incorporates a signed lexical gate that combines a sentence classifier's logit margin with a sparse lexical model's support for the predicted intent. This gate provides a continuous confidence score that is explicitly directional: it assigns positive evidence to lexical agreement and negative evidence to competing intents, while preserving the base classifier’s prediction.

  2. The system utilizes an independent binomial calibration stage based on this signed score to select an operating threshold calibrated against a specified risk target (e.g., 5% or 2% error).

  3. This results in a selective intent routing mechanism where the assistant acts on requests deemed reliable by this enhanced confidence gate and defers uncertain requests, effectively trading accepted-request coverage against error among accepted predictions.

  4. The improved system can achieve a measurable improvement in ranking quality across multiple intent classification benchmarks (BANKING77, CLINC150, HWU64), as evidenced by a reduction in the Area Under the Risk–Coverage Curve (AURC) compared to learned semantic-only gates.

  5. At nominal 5% error targets, the system can increase accepted coverage significantly on specific datasets (e.g., increasing coverage by 1.83 and 5.14 percentage points on BANKING77 and HWU64, respectively) while maintaining a controlled error rate of approximately 2.65%–2.73%.

  6. The system can maintain a non-empty policy across all tested dataset–run combinations when operating under stricter risk targets (e.g., 2% error target), overcoming the limitations of hard agreement rules which might lead to deferring all requests in some scenarios.

  7. The resulting confidence feature is compact and interpretable, allowing developers to refine risk-calibrated intent routing without needing to retrain or relabel the base intent classifier.

In summary, the improved AI system can transition from a binary accept/defer decision based solely on semantic confidence to a sophisticated, continuously graded routing policy that leverages sparse lexical cues to make more nuanced and reliable decisions under strict risk constraints.

Sources

Related papers