KumbhDoot: A Scale-Ready, LLM-Bounded Architecture for Mass-Gathering Public-Service Assistants

arXiv:2608.07520 · cs.CY, cs.AI · Submitted 2026-07-06 · Read on arXiv

Saurabh Sakalkar, Abhishek Singh, Ramesh Raskar

Cisco · Project Nanda · Massachusetts Institute of Technology

cs.CY, cs.AI

Submitted: 2026-07-06

Updated: 2026-08-11

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

The gist: KumbhDoot is designed as a public-service assistant for a massive, multilingual event where infrastructure cost is a concern to support 80M users, network connectivity may be poor and some questions

Terminology

Summary

KumbhDoot is designed as a public-service assistant for a massive, multilingual event where infrastructure cost is a concern to support 80M users, network connectivity may be poor and some questions are safety-critical. Instead of sending every query to an LLM, the system uses an escalation ladder: 1. Look for a semantically similar previously approved question and answer. 2. Classify intent by comparing the query with approximately 520 labelled examples. 3. Retrieve official information using lexical and embedding search. 4. Format a response directly from approved records. 5. Use an LLM only for low-confidence retrieval, multi-part synthesis, personalization, or live external information. Its central object is the InfoBin: an approved information unit containing content, source, language, location, validity dates, freshness, safety rules, and an embedding. The architecture has three broad workload bands: Band A: direct retrieval or cached answer; no generative model. Band B: retrieval plus coded orchestration across records; no generative model. Band C: bounded LLM synthesis over retrieved information. It also proposes an offline SQLite bundle containing selected InfoBins, FTS5 search, and compact embeddings, with Redis and larger embeddings on the server. Overall, the system shows that LLM use can be eliminated in nearly all UAT (user acceptance testing) scenarios.

The system operates on a foundational design principle that prioritizes semantic similarity over starting with an LLM. Generative models are invoked only in instances where similarity-based retrieval is insufficient to produce a correct answer. The system utilizes a semantic cache—an embedding-indexed store—as a single retrieval primitive, which handles intent routing, answer caching, offline lookups, and multi-agent retrieval. A custom three-tier agent architecture operates directly on this store, ensuring decision paths remain inspectable and avoiding the use of generic multi-agent frameworks that would trigger implicit per-step LLM calls. The paper presents the architecture, an analytical cost model for its per-query economics, and an honest account of where similarity is sufficient and where generative reasoning remains necessary. It argues that for bounded, high-stakes, low-connectivity public-service domains, a similarity-first and LLM-bounded design is not merely cheaper but architecturally more appropriate than an LLM-default one.

The paper argues that the appropriate response is to invert the usual order of operations. The foundational question of information retrieval, how to represent meaning as numbers and compare them, predates generative AI by decades and is computationally cheap, deterministic, and auditable. KumbhDoot makes that operation, cosine similarity over approved content, the default path for the large majority of queries, and treats the LLM as a bounded fallback reserved for genuinely generative or low-confidence tasks. The contribution is a systems contribution rather than an algorithmic one. The individual techniques used, dense sentence embeddings, hybrid lexical-semantic retrieval, cross-encoder reranking, and embedding nearest-neighbor intent routing, are all established. The contribution is their integration into a single similarity-first retrieval substrate, the discipline of bounding LLM use behind explicit escalation gates, and a custom, inspectable agent orchestration layer, applied to a constrained public-service domain where reliability and offline resilience matter more than open-ended capability. The paper also contributes an explicit account of the boundary between what similarity can and cannot do.

A conversational front-end backed by an LLM treats every query as an act of generation. For a public-service system at Kumbh scale, six risks follow directly from that assumption: Cost (a pilgrimage produces enormous query redundancy, and regenerating each answer with a model converts a fundamentally cacheable workload into a recurring per-query expense), Latency (emergency requests cannot wait for multi-step generation), Hallucination (when a model invents details about route access, medical facilities, or official schedules, the error is not cosmetic; grounding answers in pre-approved content rather than free generation is a safety control), Connectivity (dense crowds degrade mobile networks; offline behavior is a core requirement), Privacy (a design that ships every utterance to a cloud model maximizes exposure of exactly the data that should be minimized), and Operations (administrators need aggregate operational intelligence, not raw individual conversations or default per-user tracking). These risks motivate treating the assistant as trusted event infrastructure with a layered design, rather than as a conversational shell over a model. The central design principle is to use the cheapest, fastest, and safest path that can answer a query reliably, escalating to a model only when no cheaper path suffices.

The KumbhDoot runtime is a layered stack in which each layer answers as many queries as it can before escalating. From the bottom up: a curated content layer of official advisories, facility details, maps, schedules, and emergency contacts; an InfoBin layer that structures this content into versioned, searchable knowledge units; a local semantic cache that ships a compact partition of those units to the device; an intent router that uses rules, lightweight NLP, and semantic matching with confidence scoring; a layer of bounded domain agents for food, travel, safety, accommodation, health, emergency, ritual, profile, and crowd concerns; a response composer that formats answers from approved content and may call a small model or LLM only when required; and an observability layer tracking latency, cost, cache hits, model use, sources, safety flags, and fallback paths. The design is deliberately narrower than a general-purpose agent platform: it uses agentic structure but keeps each agent bounded, auditable, and grounded in approved content.

Curated official information typically appears in a basic app as Info Cards. For an agentic system, those cards must become structured, versioned, and locally usable objects, which are called InfoBins. An InfoBin bundles approved content with metadata (domain, language, source authority, zone, validity period, last-updated time), a semantic representation that allows natural-language matching, and a safety layer covering source labels, freshness, revocation, and escalation rules. The runtime's first question for any query is which approved InfoBin already contains the answer; only when no local match is found does the system escalate.

The central technical idea is that a single embedding-indexed store, which is termed the semantic cache, can serve five functions that conventional systems implement separately: a knowledge base of curated records, an intent classifier, a paraphrase-aware response cache, an offline data bundle, and per-agent memory. Each function is expressed as cosine similarity against a different partition of the same store. The architectural variety comes from which partition is compared against, not from different algorithms. The semantic cache is an embedding-indexed retrieval layer, not a single physical database, and the system does still use conventional storage backends. In practice the cache is realized differently on each tier. On the server, domain partitions are held in memory and cache-backed stores, including a Redis store for paraphrase-aware query-to-answer caching, using higher-dimensional embeddings that are updated with live content. On the device, a curated offline partition ships as a SQLite bundle of roughly 700 KB that combines a full-text (FTS5) index with precomputed lower-dimensional embeddings stored as binary blobs, enabling hybrid retrieval with no network call. Both tiers expose the same abstraction, approved records queried by cosine similarity and hybrid lexical-semantic search, even though they use different embedding dimensions and backends. A current limitation is that the two tiers do not share a single physical store and that live synchronization from server indexes to the device bundle is not yet implemented; the offline bundle is materialized at build time.

Most agentic systems classify intent by calling an LLM, which is expensive, slow, and opaque. KumbhDoot reframes intent classification as a similarity problem: whether a query is close to questions that humans have labelled with a known intent. The system curates approximately 520 labelled exemplar queries across 26 intent categories in English, Hindi, and Marathi, embeds them once at startup, and at query time embeds the user query and compares it against all exemplars in a single matrix multiplication. The nearest exemplar yields intent and domain labels together with a traceable numerical similarity score. This is k-nearest-neighbor classification over multilingual sentence embeddings. It consumes zero generative LLM tokens, though the embedding model itself still runs, and it produces an auditable score rather than an opaque model decision. Query complexity is determined not by the exemplar match but by a separate heuristic component that counts domain-keyword hits, conjunction words across languages, and query length. The router precedes this with an emergency check that detects urgent phrases by keyword and routes them onto a rule-based fast path, and with language normalization for transliteration and common speech-recognition errors.

Because dense crowds make connectivity unreliable, the device tier is designed to remain useful offline. Emergency contacts, key facilities, saved itinerary, basic maps, FAQs, phrasebooks, and previously synced InfoBins are available without a network. When connectivity is absent the system degrades gracefully: voice falls back to text, agentic answers fall back to approved static guidance, and live data is labelled with its last-updated status. Non-critical requests are queued and resumed on reconnect, while emergency updates receive sync priority. Answers are explicitly labelled as available offline, last updated, or pending refresh, so users can judge freshness. On the device, retrieval is hybrid: a lexical FTS5 search is fused with semantic similarity over the embedding blobs using rank fusion. When a quantized on-device embedding model is present, the client embeds live queries locally; when those assets are absent, it falls back to word-overlap matching against pre-baked query embeddings, preserving basic offline function without live embedding.

For queries that require composition across domains, KumbhDoot uses a custom three-tier agent architecture in which all tiers share the same semantic retrieval layer. An orchestrator decomposes a complex request into sub-needs with dependency edges and schedules them into parallel execution waves, so that independent sub-needs run concurrently while dependent ones wait for their prerequisites and receive prior results as context. Domain agents, one per domain such as health, travel, commerce, crowd, identity, governance, and curated Kumbh content, each operate over a partition of the cache. Record-level sub-agents enrich retrieved items into structured answers with actionability signals, card type, emergency flags, and name matching. Two design choices distinguish this architecture. First, the orchestration logic is hand-written and domain-specific rather than built on a general-purpose multi-agent framework. Because the agents are wired to the semantic cache first, an LLM is reached only under explicit routing conditions rather than as a default at every orchestration step, which keeps each agent's decision path inspectable. Second, adding a domain is primarily a data and registry operation, loading curated records into a new partition and extending a data-driven registry, provided the shared schema and routing rules already accommodate that domain; it is not, in the common case, a matter of writing new orchestration code.

The runtime tries methods in cost order and reaches the LLM only when the cheaper layers cannot answer. A query first hits the paraphrase-aware query-to-answer cache; a confident single-record hit is returned by a coded formatter with no LLM call. Failing that, cosine-based intent routing selects a domain at zero generative cost. For confident single-domain retrieval, multiple records are formatted into a grounded, cited answer, still without an LLM. The LLM is invoked only when an explicit escalation gate fires: a complex multi-part synthesis, retrieval confidence below threshold, or a need for live or external-source data. Even then, generation is constrained to grounded, retrieved records, and a grounding-validation step strips any citation that does not correspond to a record the retriever actually returned, removing invented identifiers before the answer is shown.

The similarity-first design produces a distinctive cost structure that can be reasoned about analytically. In a pipeline where every processing step is an LLM call, cost accrues at each of intent classification, complexity assessment, decomposition, per-domain reasoning, and synthesis. In KumbhDoot, deterministic components handle each of these steps, and the model is invoked only as a bounded fallback. Cosine similarity over a fixed exemplar set is a single matrix multiplication; paraphrase matching in the response cache is a dot product; hybrid retrieval combines vector math with a lexical index and a small local reranker. None of these consume generative tokens. The embedding step does consume compute, so common paths cost zero generative LLM tokens rather than zero compute. On the device tier, where vectors are pre-baked and shipped in the bundle, even the embedding step can be avoided for cached content. This structure yields a per-query cost profile that scales favorably with volume. The first time any question is asked, the cost is likely to be one embedding and a retrieval. Every subsequent semantically equivalent question, in any of the supported languages and in any phrasing, is likely to be served from the cache. Because pilgrimage query distributions are heavily concentrated on a relatively small set of recurring needs, the share of queries answerable without a model call grows as the cache warms, and the marginal cost of the common case approaches the cost of similarity computation alone. The LLM cost is incurred only on the long tail of genuinely novel, multi-step, or externally-sourced queries.

A central claim of the paper is that the boundary between similarity and generation should be stated explicitly rather than blurred. In KumbhDoot's bounded and curated domain, semantic similarity combined with coded formatters is sufficient for a substantial class of needs: intent classification over a fixed label set, paraphrase-tolerant retrieval, FAQ-style answers, and structured factual responses drawn from approved records. For these, a generative model adds latency, cost, and hallucination risk without adding correctness. Similarity is not, however, a substitute for reasoning or open-ended generation, and several hard limits are identified: queries about records absent from the corpus yield low-confidence retrieval and must deflect or escalate rather than fabricate; retrieval finds the closest text, which does not guarantee correct intent at decision boundaries; multi-constraint planning across domains requires orchestration and often genuine prose generation; coded formatters produce accurate structured output but not fluent narrative; transactional actions such as booking or identity verification need tool-using agents, not cosine similarity; and multilingual coverage does not eliminate degradation from code-mixing, dialect variation, and speech-recognition error. These limits map onto a three-band view of the workload. In Band A, similarity alone answers the query through the response cache or direct retrieval, with no model call: for example, asking where a ghat is or what the bathing dates are. In Band B, similarity plus coded orchestration handles multi-record, structured requests, again without a model: for example, listing hospitals with emergency care near a locality. In Band C, the system escalates to LLM-backed synthesis and dialogue: for example, an open-ended personalized plan for an elderly traveller arriving on a given day. The design goal is not to push everything into Band A but to ensure that queries are answered in the cheapest band that can answer them correctly, and that escalation to Band C is a deliberate, gated decision rather than a default. On hallucination, by driving agents from curated, embedding-indexed records and using coded formatters for the majority of responses, KumbhDoot reduces opportunities for unconstrained generation and thereby mitigates hallucination relative to architectures that invoke a model at every step. It does not eliminate hallucination: the system still calls a model for some queries, and retrieval-based systems can themselves surface incorrect records. The grounding-validation step that removes unsupported citations is a mitigation, not a guarantee.

The paper presents an architecture and an analytical cost argument; it does not yet present a controlled empirical evaluation. The central quantitative claims, the fraction of production queries answerable without a generative model call, the realized per-query cost and latency by band, and the speedup relative to an LLM-default pipeline, require a controlled benchmark over a representative query distribution. A tiered benchmark is planned spanning simple, multi-record, and complex multi-domain queries, each paired with semantically equivalent paraphrases to test cache behavior, with full per-step tracing of tokens, latency, and cost. The baseline must be characterized honestly: an LLM-for-everything pipeline that forces a model through every step is a worst-case rather than a representative comparison, and a fair evaluation should also include a conventional RAG baseline that uses a model only for synthesis. Emergency systems need to have more escape hatches. A keyword fast path is useful, but it cannot be the principal emergency safety mechanism. There is a chance of missing indirect statements, transcription errors and code-mixed expressions, and it may create false positives. The production design should additionally include: a permanent one-tap emergency control, locally stored emergency numbers and instructions, explicit confirmation of location and emergency type, direct telephone/SMS paths, human escalation, and a safe fallback when classification confidence is uncertain. Cell broadcast is especially relevant for official one-to-many emergency alerts because it is designed to operate with limited impact from network congestion. The accuracy of cosine-based intent routing across English, Hindi, Marathi, and code-mixed input is asserted but not yet measured. Embedding quality degrades under dialect variation and speech-recognition error, and the confidence thresholds that govern answer, clarify, deflect, and escalate decisions need to be tuned and reported against a labelled multilingual test set. The minimum InfoBin bundle that yields useful offline behavior, and the cache hit rate achievable on-device, are open empirical questions. The favorable scaling argument depends on a query distribution concentrated enough for the cache to warm quickly, which should be validated on real traffic rather than assumed. Live synchronization from server indexes to the device bundle is not implemented; the offline partition is materialized at build time. Closing this gap, including signed updates and revocation for stale or corrected guidance, is necessary before the offline tier can be trusted for fast-changing advisories. Cached guidance can become stale during road closures, crowd changes, or emergencies. The system design includes last-updated labels, tiered refresh cadences, and revocation lists, but the operational discipline of keeping safety-critical InfoBins current, and the audit standards appropriate to a public-sector agentic application, require field validation. The intended privacy posture, bucketed rather than exact attributes, local-first storage, coarse and consented location, aggregate-by-default dashboards, and audit logs, is a design commitment that needs to be verified in deployment, including the question of how much personalization is useful before its privacy cost outweighs its benefit.

KumbhDoot addresses the specific challenges of mass-gathering environments, where large-scale, multilingual, safety-critical public services must operate despite limited connectivity and high cost sensitivity. By rejecting the standard LLM-default architectural pattern, the system prioritizes reliability and efficiency through an inverted operational order: semantic similarity over a curated knowledge base serves as the foundational principle, while large language models are restricted to a bounded, secondary role. The system is engineered to handle mass-gathering scenarios characterized by millions of concurrent users, intermittent network connectivity, code-mixed and multilingual input, and safety-critical information needs where errors carry physical consequences. Rather than generating every response via an LLM, the system centers on semantic similarity. It utilizes a single retrieval primitive—cosine similarity over an embedding-indexed semantic cache—to perform five key functions: intent routing, paraphrase caching, offline data lookup, and multi-agent retrieval. The implementation employs a hand-coded, three-tier agent architecture that operates directly on the semantic store. This bypasses the opacity of general-purpose multi-agent frameworks, ensuring that decision paths remain fully inspectable and that LLM calls are gated rather than automatic. The authors provide an analytical cost argument to support the design, while remaining transparent about the current implementation's limitations and the necessity of future empirical work to establish quantitative claims and validate research gaps. The architectural thesis is that for bounded, high-stakes, and low-connectivity domains, a similarity-first and LLM-bounded approach is both economically superior and architecturally appropriate. The value of generative AI in such contexts is maximized by reserving it exclusively for tasks that cannot be solved through reliable, deterministic retrieval.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in AI systems, along with the resulting capabilities:


What I change: Instead of routing every query to an LLM, I add a multi-tier retrieval pipeline: (a) semantic cache lookup for exact/paraphrase matches, (b) intent classification via k-NN over 520 labelled exemplars (26 intents, 3 languages), (c) hybrid lexical-semantic retrieval over approved InfoBins, (d) coded formatters for structured answers, and (e) LLM only for low-confidence retrieval, multi-part synthesis, personalization, or live data.

What the improved system can do: Answer 80–90% of common queries (FAQs, schedules, locations, emergency contacts) with zero generative LLM calls, reducing per-query cost by 10–100x and latency from seconds to milliseconds, while eliminating hallucination risk on those paths.

What I change: I replace separate intent classifiers, response caches, and retrieval indexes with one embedding-indexed store. The same cosine-similarity operation serves five functions: knowledge base lookup, intent routing, paraphrase-aware answer caching, offline data bundle, and per-agent memory.

What I change: I replace general-purpose agent frameworks (which trigger implicit per-step LLM calls) with a custom orchestrator, domain agents (health, travel, safety, etc.), and record-level sub-agents—all operating directly on the semantic cache. Orchestration logic is explicit, inspectable, and domain-specific.

What I change: I embed 520 curated exemplar queries across 26 intents once at startup. At query time, I embed the user query and compare against all exemplars in a single matrix multiplication. The nearest exemplar yields intent, domain, and a numerical confidence score.

What I change: After any LLM synthesis, I run a validation step that checks every citation in the generated answer against the actual retrieved records. Any citation not corresponding to a record the retriever returned is stripped before the answer is shown.

What I change: I ship a pre-baked offline bundle (SQLite with FTS5 + compact embeddings) to the device. When connectivity is absent, the system falls back to: hybrid lexical-semantic retrieval, pre-baked query embeddings, cached InfoBins, and static approved guidance. Voice falls back to text; live data is labelled with last-updated status.

What I change: Beyond keyword-based emergency detection, I add: a permanent one-tap emergency control, locally stored emergency numbers/instructions, explicit confirmation of location and emergency type, direct telephone/SMS paths, human escalation, and a safe fallback when classification confidence is low.

What I change: I classify each query into Band A (direct retrieval/cache—no model), Band B (retrieval + coded orchestration—no model), or Band C (bounded LLM synthesis). I route queries to the cheapest band that can answer correctly.

What I change: I define and enforce thresholds for retrieval confidence that govern four outcomes: answer directly, ask clarifying question, deflect (say I don't have that information), or escalate to LLM. I also add a nothing in, nothing out rule: if no record matches, the system never fabricates.

What I change: I structure the system so that adding a new domain (e.g., "lost & found or weather") is primarily a data operation: load curated InfoBins into a new partition, extend the data-driven registry, and reuse the existing schema and routing rules.

The improved AI system can:

  • Serve millions of concurrent users at near-zero marginal cost for common queries

  • Operate offline with a 700 KB bundle, maintaining core functionality

  • Answer in English, Hindi, and Marathi (including code-mixed input) with auditable intent classification

  • Eliminate hallucination on all Band A/B paths and strip unsupported citations on Band C

  • Handle emergencies deterministically with multiple escape hatches and human escalation

  • Scale cost linearly with novel queries only, not with total volume

  • Provide full auditability of every decision path (cache hit, intent score, retrieval confidence, LLM call reason)

  • Extend to new domains via data loading, not code changes

Sources

Related papers