Staged Linguistic Seeding: Grounded Query Expansion for Verified-Unit QA in AI Contact Centers
cs.CL, cs.IR
Submitted: 2026-09-01
Updated: 2026-09-07
Comments: 14 pages, 1 figure, 7 tables. Accepted to the Grounding Language Models (GroundLM) Workshop at EMNLP 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Customer-service QA in an AI contact center (AICC) runs under deployment constraints that benchmark QA misses: tight voice-hotline latency and a high cost for unsupported or wrong automatic answers.
Terminology
Abstract
Customer-service QA in an AI contact center (AICC) runs under deployment constraints that benchmark QA misses: tight voice-hotline latency and a high cost for unsupported or wrong automatic answers. We deploy a system that answers only from a closed set of verified QA units: it returns a retrieved unit verbatim, or routes to clarify, abstain, or handoff. The index is enriched offline by staged linguistic seeding (SLS): a human authors a per-unit world-grounded slot recipe, gpt-4.1-mini renders it into variants, and a light human gate filters them. One methodology is reused across both domains, so inference stays a single retrieval pass with no query-time generation. On held-out query variants from two industrial domains, SLS lifts hybrid R@1 to 0.881/0.930 (+0.27/+0.34), with gains across all five retrievers tested. At the same gpt-4.1-mini generation budget, SLS beats doc2query by +0.20/+0.32, while cross-provenance evaluation provides additional evidence of transfer across generated-query distributions. Verified-unit answering also removes free-form generation's unsupported-content surface (7-13% versus approximately 0%). We report this as an application study, including negative results.
Sources
- Query expansion techniques for information retrieval: A survey
- A System for Human-AI collaboration for Online Customer Support
- M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation
- FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance
- EnrichIndex: Using LLMs to Enrich Retrieval Indices Offline
- Turning Conversations into Workflows: A Framework to Extract and Evaluate Dialog Workflows for Service AI Agents
- Ragas: Automated Evaluation of Retrieval Augmented Generation
- Precise Zero-Shot Dense Retrieval without Relevance Labels
- DRAMA: Diverse Augmentation from Large Language Models to Smaller Dense Retrievers
- Enabling Large Language Models to Generate Text with Citations
- Evaluation of semantic relations impact in query expansion-based retrieval systems
- Teaching language models to support answers with verified quotes
- Query Expansion by Prompting Large Language Models
- Document Expansion by Query Prediction
- AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions
- Measuring Attribution in Natural Language Generation Models
- ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems
- Parametric Retrieval Augmented Generation
- Query2doc: Query Expansion with Large Language Models
- A Financial Service Chatbot based on Deep Bidirectional Transformers
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering