On-Device Commercial Intent Retrieval Under Size, Latency, and Privacy Constraints: A 3 MiB Retrieval System with Typed Egress Boundaries

arXiv:2610.00170 · cs.IR, cs.CL · Submitted 2026-09-16 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "On-Device Commercial Intent Retrieval Under Size, Latency, and Privacy Constraints".

Tom: As a meticulous researcher, I have thoroughly analyzed both provided summaries of the paper "On-Device Commercial Intent Retrieval Under Size, Latency, and Privacy Constraints:

Jane: First, who's behind it and why it matters.

Title and authors: Tom: Speaking of structure, let's check out who wrote this piece, because it’s important to know the team behind the engineering study on "On-Device Commercial Intent Retrieval Under Size, Latency, and Privacy Constraints: A three MiB Retrieval System with Typed Egress Boundaries."

Jane: The paper is authored by Hyojung Han. It’s interesting how focused this work is; it’s not about inventing a brand new retrieval algorithm from scratch, but rather designing the system around very specific physical limits that are becoming increasingly important for user experience.

Lu: Han's focus on constraint satisfaction is what makes this paper stand out in my view. It shows a deep understanding of how to make a system work within hard engineering boundaries, which is something we need more of in complex AI deployments.

Meng: I’m interested in the fact that they framed their contribution not as a new algorithm, but as system design under constraint; that tells me the real value here is in the engineering discipline rather than just theoretical model performance.

Lalam: I think Han’s work underscores how crucial it is to bake privacy and size limits directly into the design gates from the very beginning, instead of trying to bolt them on later. That upfront constraint management really sets a strong precedent for future AI development.

The paper's summary: Tom: Now that we know who’s behind it, let's look at what they actually achieved in this paper, "On-Device Commercial Intent Retrieval Under Size, Latency, and Privacy Constraints: A three MiB Retrieval System with Typed Egress Boundaries."

Jane: Essentially, the authors built a retrieval system for behavioral advertising that is designed to run directly on the device. They set three major rules upfront: the total download size must be under three MiB, inference needs to happen in Tier-zero under twenty milliseconds at the p95 level, and absolutely no raw text or content embeddings are allowed to leave the device.

Lu: What’s really impressive is how they achieve this by using a static embedding table that they distilled from a Korean sentence transformer and then quantized it down to just four bits, meaning it doesn't need any inference runtime for the core search process at all.

Meng: The measured payload came in at two million nine hundred forty-two thousand six hundred fifty-two bytes, which is about ninety-three point five percent of that three MiB limit across their tests on Android devices—that’s a tight margin we have to respect when planning for mobile apps.

Lalam: And the privacy aspect is very explicit; they enforced this using a federated layer where the personal head data simply doesn't have a path to serialization, and they confirmed that seven canaries were undetected during testing, meaning they successfully kept identifiers and embeddings contained.

The paper's improvements: Tom: So, what are the actual improvements this system suggests over previous approaches? We need to look at how they made it better than what was already out there for on-device intent retrieval.

Jane: The main improvement here isn't necessarily a different model, but the system design itself, which is focused entirely on satisfying those three hard constraints simultaneously. They introduced a three-tier structure: Tier zero for always-on inference with no runtime needs, Tier one for optional inference using existing OS models, and a federated layer to manage the privacy boundaries.

Lu: I think the shift from relying on dynamic model inference to this static lookup table is the key design improvement; it drastically cuts down on computational load during query time by removing those complex transformer forward passes altogether.

Meng: From an engineering standpoint, that reduction in runtime complexity is huge because it directly impacts latency, and they managed to keep the Tier-zero latency p95 under five point zero eight zero milliseconds on some devices—that’s very close to what we need for a smooth user experience.

Lalam: The improvement in how they handled the privacy boundary by using type-checking on the upload schema means that even if there was a slight oversight, the protocol itself prevents sensitive data from leaving, which is a much stronger safeguard than just relying on policy documentation alone.

Conclusion: Tom: We’ve seen how they built this system and what their core design improvements are; let’s wrap up by looking at the final conclusions of "On-Device Commercial Intent Retrieval Under Size, Latency, and Privacy Constraints: A three MiB Retrieval System with Typed Egress Boundaries."

Jane: The authors conclude that while the design successfully enforces all those specific size, latency, and privacy constraints and achieves a measurable accuracy of seventy-five point zero percent on product surface top-five for real Korean commerce text, their success is highly conditional on the operational regime they are in.

Lu: They emphasize that the cost of this design isn't uniform; for instance, they found that when a listing named the kind of thing sold, there was a penalty roughly ten points higher in accuracy loss compared to when only a brand and model were named.

Meng: I think that uneven cost analysis is important because it tells us we can’t just optimize for one metric; we have to balance what matters most for the specific use case. The core thesis here seems to be that the real value is in successfully executing the intended function without ever moving raw behavioral data off the device.

Lalam: I really think their final measure of success—the on-device confidence signal derived from the ranker’s score gap—is a very mature way to define winning when accuracy isn't everything. It shows that even under these limitations, we can still derive meaningful intent signals locally.

Tom: So, as we wrap up this discussion on "On-Device Commercial Intent Retrieval Under Size, Latency, and Privacy Constraints: A three MiB Retrieval System with Typed Egress Boundaries," the big picture is that the competitive advantage is shifting toward executing functions locally rather than relying on a massive central server.

Jane: It’s a fascinating piece of engineering because it shows how you can achieve strong performance metrics while strictly adhering to severe real-world limitations like payload size and latency budgets.

Lu: This work opens up a lot of avenues for developing more efficient, privacy-aware AI agents that operate directly on user hardware.

Meng: From an implementation view, the lesson is clearly about making every byte count and designing the system with strict, falsifiable constraints in mind from the start.

Lalam: It confirms that robust privacy isn't just about policies; it’s about embedding those boundaries into the very structure of how data flows within an AI interaction.

Hyojung Han

cs.IR, cs.CL

Submitted: 2026-09-16

Updated: 2026-09-16

Comments: 36 pages, 40 tables, 1 figure. Evidence bundle (measurement ledgers and verification gates, no datasets): https://github.com/hyojunguy/ondevice-intent-evidence

Code: https://github.com/hyojunguy/ondevice-intent-evidence

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 83/100

The gist: As a meticulous researcher, I have thoroughly analyzed both provided summaries of the paper "On-Device Commercial Intent Retrieval Under Size, Latency, and Privacy Constraints: A 3 MiB Retrieval

Key concepts

On-Device Commercial Intent Retrieval
This is a system designed to find specific commercial goals or intents within a user's local data, such as identifying what kind of product someone intends to buy. The key challenge is doing this search entirely on the user's phone without sending any personal information or large files outside the device.
Typed Egress Boundaries
This refers to strict rules that control what data can leave the device. In this system, it means ensuring that no sensitive personal data—like raw text or content embeddings—can be serialized or transmitted externally. Privacy is enforced by protocol design rather than just documentation.
Tier-0 Inference
This describes the fastest possible mode of operation for the system, which runs continuously without needing any extra software dependencies at runtime. This tier is crucial for meeting the strict latency requirements, ensuring that commercial intent checks happen almost instantly.
Egress Control Canaries
These are internal checks or 'canaries' built into the system to detect if any unauthorized data is trying to leave the device. The paper tested these and found zero detections, meaning the privacy boundaries were successfully maintained during testing.

Terminology

Summary

As a meticulous researcher, I have thoroughly analyzed both provided summaries of the paper On-Device Commercial Intent Retrieval Under Size, Latency, and Privacy Constraints: A 3 MiB Retrieval System with Typed Egress Boundaries. The following synthesis integrates the technical details from both sources to construct a comprehensive and detailed description of the work.


This research presents a novel system designed for on-device commercial intent retrieval, specifically targeting behavioral advertising applications, while strictly adhering to severe operational constraints related to payload size (under 3 MiB), inference latency (Tier-0 inference under 20 ms at p95), and stringent privacy requirements (no raw text, content embeddings, or stable identifiers leaving the device).

The core innovation lies in transforming the three primary constraints—size, latency, and privacy—into falsifiable propositions that are adjudicated by a custom code-based system. The architecture employs a three-tier structure:

  1. Tier 0: Always-on inference with no runtime dependencies.

  2. Tier 1: Optional inference utilizing an on-device OS model for enhanced capability.

  3. Federated Layer: A mechanism ensuring that personal head data cannot be serialized or transmitted externally, enforcing a privacy boundary by type, preventing any personal slice data from having a path into the upload type.

Crucially, the system enforces egress control through seven canaries, which were detected as zero detections during testing. Furthermore, the final communication leg utilizes a closed OpenRTB 2.6 request protocol where an identifier has no field to reside within it, reinforcing privacy boundaries through protocol enforcement rather than documentation alone.

The system constructs a retrieval path over a 6,020-leaf commercial taxonomy. This path is built upon a static embedding table that was distilled from a Korean sentence transformer and quantized to 4 bits. The search strategy combines two elements:

  1. Dense Cosine Score: Used for initial ranking.

  2. Lexical-Overlap Term: Incorporated as a hybrid comparison factor against teacher models.

Performance evaluation yielded several key results:

  • Payload Size: The measured deployment payload was 2,942,652 bytes (93.5% of the 3 MiB limit), with Android being the largest measured platform among three tested devices.

  • Latency: Tier-0 p95 latency ranged between 3.670 ms and 5.080 ms across various devices, including two iPhones running different SoC generations and a budget Android tablet (Snapdragon 695). Critically, every measured device was slower than every server CPU, with server CPU latencies ranging from 2.00 to 3.37 ms for the same code.

  • Accuracy: The system achieved a product surface top-5 accuracy of 75.0% when tested against real Korean commerce text, significantly outperforming the 18.4% permutation baseline. However, performance exhibited a sharp split: 83.5% accuracy was observed when the query contained some leaf name as a substring, whereas it dropped to 45.2% when no such substring was present.

The cost associated with adhering to the size constraint is not uniform across all listing types:

  • The penalty (in terms of accuracy loss) was roughly ten points higher when a listing named the kind of thing sold.

  • The penalty was roughly twenty points higher when a listing named only a brand and model.

Task adaptation demonstrated nuanced results: fine-tuning the teacher model could yield substantial gains (up to +10.5 percentage points) in certain regimes, but this benefit was not detectable in others, suggesting that the system's success is dictated by capacity versus training constraints rather than simple retraining. The performance on real Korean commerce text was further influenced by the presence of a manufacturer model code within the query string.

The paper concludes that while the design successfully enforces all specified constraints and achieves measurable accuracy under these severe limitations, its success is conditional upon specific operational regimes. The cost of this design is inherently non-uniform, depending critically on whether the listing provides category context or only brand/model information.

The ultimate measure of success for this system is not solely accuracy, but rather an on-device confidence signal derived from the ranker’s own score gap between the first and fifth mid-category. The central thesis redefines the competitive moat: **the advantage lies not in achieving superior accuracy compared to a server, but in successfully executing the intended function even though the critical event (the raw behavioral data) never leaves the device.

Improvements for AI systems

Based on the provided scientific paper, here are specific, high-impact improvements for AI systems:


  1. The core improvement is a shift from traditional model-based inference to a highly constrained, static retrieval architecture optimized for commercial intent discovery directly on the device.

  2. The improved AI system can perform:

Describe the new system as follows: A low-latency, privacy-preserving on-device retrieval engine that operates entirely within a strict 3 MiB payload budget and requires zero runtime inference for its primary function. It achieves this by using a static, 4-bit quantized embedding table derived from knowledge distillation of a large Korean sentence transformer.

  1. The system can handle:

Summarize the specific capabilities:

  • It can process real-time colloquial queries (e.g., 패딩 사고 싶은데) and route them to the correct commercial category (leaf) with high accuracy (up to 83.5% when an exact leaf name is present).

  • It can perform zero raw text egress, ensuring strict privacy by design.

  • It can utilize a two-tier structure: Tier 0 provides always-on intent vector generation via static lookup, and Tier 1 optionally leverages existing OS models (like Gemini Nano) for deeper interpretation when available.

  1. The system can optimize:

Detail the technical efficiencies achieved:

  • It minimizes size by dropping category and catalog vectors from 8-bit to 4-bit per row, achieving a 425 KB reduction in the total payload compared to an 8-bit version while maintaining high quality (e.g., L2 top-5 product surface accuracy remains stable).

  • It eliminates the need for complex transformer forward passes during query time by relying on static embedding lookups and inner product matching, achieving sub-millisecond latency (p95 of 3.670 ms on reference iPhones).

  1. The system can adapt:

Detail the learning and personalization capabilities:

  • It implements a Federated Learning (FedPer) structure where only the personal head parameters are trained locally, ensuring raw user data never leaves the device. This allows for personalized intent matching based on local user patterns without compromising privacy invariants enforced by types.
  1. The system can enhance:

Detail the accuracy gains achievable through task adaptation:

  • It demonstrates that task-specific contrastive training (fine-tuning a model on the user's own utterance/leaf pairs) can yield significant accuracy boosts (+10.5 pp when an anchor is present) compared to a generic teacher, allowing for high performance in niche areas.
  1. The system can manage:

Detail the robust privacy and security mechanisms:

  • It enforces a privacy boundary by types, ensuring that stable identifiers, raw text, or content embeddings cannot be serialized into any upload request (OpenRTB 2.6), effectively blocking leakage at the protocol level rather than relying on policy alone.
  1. The system can monitor:

Detail the diagnostic and validation framework:

  • It features a rigorous budget gate mechanism that adjudicates against three constraints (Size, Latency, Privacy) in real-time, providing measurable feedback on the cost of design choices (e.g., quantifying the trade-off between a 3 MiB limit and accuracy).

  • It includes an explicit accounting layer for campaign-wide differential privacy budget adherence, dynamically adjusting learning based on device population size to ensure privacy guarantees hold or reverting to a safer, less private mode when conditions are not met.

Abstract

We study commercial intent inference that runs entirely on the user's device, under three constraints frozen before the work began: the downloaded payload under 3 MiB, Tier-0 inference under 20 ms at p95, and no raw text, content embedding, or stable identifier leaving the device. Under them we build a retrieval path over a 6,020-leaf commercial taxonomy: a static embedding table distilled from a Korean sentence transformer, quantized to 4 bits, no inference runtime. Our main result is where that constraint costs accuracy. On real Korean commerce text labelled by others (22,900 AI-Hub shopping reviews), mid-category top-5 on real product names is 75.0% against an 18.4% permutation baseline, but splits on one observable: a query containing some leaf name as a substring scores 83.5%, one containing none 45.2%. A generic 196.6x larger teacher seemed to localize the gap (+20.1 pp without an anchor, +0.1 with). That null was two effects cancelling: the same teacher fine-tuned on the student's own contrastive pairs reaches 0.8586 and beats the pure-encoder student by +10.6 pp with an anchor and +20.9 pp without. The cost is not uniform, but it is not free anywhere; where the anchor is absent, task adaptation buys the teacher nothing, so what the constrained encoder lacks there is capacity. The expensive regime is detectable on-device from the ranker's own score margin: declining the least confident fifth lifts the rest to 0.8296. A second axis we first reported, a manufacturer model code, does not survive source-category fixed effects (-4.0 pp, p=0.51); the anchor does (+13.0 pp). Payload is 2,942,652 bytes, all three library links measured. Tier-0 p95 is 4.431 and 3.670 ms on two iPhones (A14, A16) and 5.080 ms on a budget Android tablet (Snapdragon 695), all slower than three server CPUs on the same code. Taxonomy supervision is mostly synthetic Korean utterances.

Sources

Related papers