LocalSUG: City-Preference-Enhanced LLM for Query Suggestion in Local-Life Services

arXiv:2603.04946 · cs.CL · Submitted 2026-03-05 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "LocalSUG: City-Preference-Enhanced LLM for Query Suggestion in Local-Life Services".

Jane: The paper was written by Jinwen Chen, Shiwen Zhang, Shuai Gong, Zheng Zhang, Yachao Zhao et al. from Meituan.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Summary: Tom: Okay, we've established that "LocalSUG: City-Preference-Enhanced LLM for Query Suggestion in Local-Life Services" is about making suggestions more contextually aware. Now, let's dig into the summary they provide—it’s really impressive how they frame this problem of insufficient locality in existing models.

Jane: The core idea from the summary seems to be that general LLMs are great at theory but struggle with practical, ground-level reality; they lack that 'sense of place.'

Tom: Exactly. They're showing that simply knowing what an item is isn't enough; you have to know *where* it is and *why* it’s relevant to the user’s immediate context.

Meng: What struck me reading the summary was how they formalize the concept of 'preference.' It implies that the model isn't just matching keywords, but it's prioritizing results based on implied user needs related to their location.

Lu: This moves us into the realm of cognitive modeling, doesn't it? We’re moving away from pure information retrieval and towards predictive assistance that anticipates real-world needs based on context.

Jane: It’s like having a helpful friend who knows your neighborhood really well, rather than just searching through the entire internet for you. They filter out the noise because they know what's actually useful *right now*.

Tom: And this enhanced ability to understand that 'preference' based on location is critical because if the suggestions are irrelevant, the user just gets frustrated and leaves.

Lalam: For culture, this means that local commerce doesn't have to compete with globally optimized, generic e-commerce giants; the AI can champion the unique, small-scale enterprise by making it highly visible and discoverable through superior suggestion technology.

Meng: I'm thinking about integration—if we could hook this into municipal services or public transit apps, it would be revolutionary for urban planning and civic engagement.

Lu: But Meng, while the potential is huge, we have to consider the data pipeline complexity; maintaining real-time accuracy across diverse city infrastructures will require enormous computational overhead.

Jane: So the summary shows that they solved a conceptual problem—the lack of place-based awareness—which is a massive leap for general AI systems.

Tom: It really sets the stage for how these advanced models can actually improve daily life, moving beyond just generating text to genuinely assisting navigation and discovery. But how did they achieve this enhancement? Let's look at what improvements they suggest in the next section.

Paper discussion segment 2: Tom: So, to recap our last segment, we’re looking at LocalSUG as a system designed to give users better search suggestions in local services like food delivery by leveraging LLMs for City-Preference-Enhanced query suggestion.

Jane: It really boils down to the idea that the AI needs to know what you're looking for *and* where you are, right? A generic model can't tell you which pizza place is open nearby, but LocalSUG makes that specific local knowledge a part of the search process.

Lu: That’s just scratching the surface of what this could be. I see this as a huge step toward a truly "intelligent" city where AI isn's not just retrieving data but anticipating needs based on hyper-local dynamic contexts, transforming how we interact with urban infrastructure entirely.

Meng: But Lu, while that vision is exciting, it makes me wonder about the practical overhead. How do we manage all these shifting local preferences—merchants opening and closing daily—without the system crashing under a massive data synchronization load?

Lalam: Meng raises a valid point about scale, but I think the cultural impact is huge because this allows local businesses to compete on merit. Instead of drowning in generic global noise, your neighborhood gets its own intelligent champion that highlights unique, small-scale offerings based on real community needs.

Tom: That’s exactly what I mean; it' isn't just about finding the nearest store, it’s about finding the *best* local store for a specific intent. It bridges that gap between what we want and where we are.

Jane: And by injecting those city-specific candidates directly into the prompt, they avoid making mistakes based on old data, ensuring the suggestions are fresh and relevant to what's actually available today.

Lu: I think this opens up a whole new paradigm in recommendation engineering, moving us away from simple historical popularity towards a dynamic, real-time understanding of consumer desire.

Meng: Speaking of implementation, if we're using Qwen3-zero point 6B as the backbone, we need to ensure the entire pipeline is optimized for high concurrency without sacrificing that latency requirement.

Lalam: Imagine the cultural shift when this technology could be applied to public services too; a suggestion engine that knows which bus route is locally optimal or which community resource is most relevant right now, transforming civic participation.

Tom: It's fascinating how solving this local knowledge problem can have such wide-reaching implications for both the tech industry and our daily lives.

Jane: So, if LocalSUG enables this level of localized intelligence, we need to look at how they manage the optimization process next.

Paper discussion segment 3: Tom: It’s fascinating how they found that forty-one percent of those generated queries don't even exist in the candidate list! That proves this LLM is doing way more than just reading off a suggestion card.

Jane: Right, it means the system isn't limited by what someone else has typed before, which is huge for user experience. Instead, it’s synthesizing new ideas based on multiple inputs—like history plus current location.

Meng: From an engineering standpoint, that compositional ability sounds highly effective until you run into data sparsity; how do you train a model to reliably invent useful concepts when the training data itself is patchy or noisy?

Lu: Well, instead of just looking at historical queries, we could force the model to consider structured knowledge graphs alongside the text. Imagine linking "MixC" not just as a place name, but as an intersection of cuisines and events.

Lalam: Thinking about that compositional intelligence makes me think about how much human connection is lost when technology only offers suggestions; this system is helping restore a sense of personalized discovery, which is fundamental to community life.

Jane: Exactly! It’s not just getting someone to *find* something, it's making them feel like the system actually *understands* their lifestyle.

Tom: And that understanding goes beyond just knowing where they were last week; it's predicting what they might need next week based on their pattern of life.

Meng: If we could integrate real-time environmental data—like localized weather or sudden community events—that compositional capability would move from useful to essential for daily life services.

Lu: Maybe the next iteration needs to incorporate temporal decay, making sure that suggestions fade gracefully if they haven't been relevant for a long time, keeping the suggestion space fresh.

Lalam: That freshness mirrors the changing nature of human interest; a good suggestion should feel like serendipity, not just optimized prediction. So, if we can make these suggestions feel genuinely useful and compositional, what happens when we apply this level of localized intelligence to something bigger than just finding a noodle shop?

Conclusion: Tom: So, to wrap things up for our listeners, we’ve seen that LocalSUG is a huge step forward in making query suggestion in local services more intelligent by using LLMs with city-specific context.

Jane: It’s wonderful to think about how this technology has successfully combined the power of generative AI with the practical need to make sure suggestions are locally relevant and up-to-date, without getting bogged down in outdated information.

Lu: I think the biggest shift here is that it finally moves us past simply looking at what was popular; it allows for a dynamic, real-time understanding of how local preferences are evolving.

Meng: From an operational standpoint, we' can be confident that managing this framework is robust and efficient enough to handle high traffic volumes in a production environment without sacrificing quality.

Lalam: By empowering local businesses with this kind of visibility, it ensures that the unique flavor and needs of the community can thrive in the digital space.

Tom: That sense of local empowerment is something we really hope to see reflected in how people use these services more broadly.

Jane: And knowing that this entire system—the training, the reward balancing, and the accelerated inference—is designed for a high-concurrency environment gives us peace of mind about its scalability.

Lu: I’m particularly excited about how this provides a stable platform for future research into more complex recommendation strategies.

Meng: It's definitely an efficient way to handle the constraints; we can actually deploy this at scale, which is a massive win for real-world implementation.

Lalam: It feels like the right moment to see these localized AI advances integrated into our everyday lives.

Tom: We’re really looking forward to seeing how LocalSUG: City-Preference-Enhanced LLM for Query Suggestion in Local-Life Services is put into practice.

Jane: We’ll be back next week with another exciting paper from the world of AI, so make sure you stay tuned!

Jinwen Chen, Shiwen Zhang, Shuai Gong, Zheng Zhang, Yachao Zhao, Lingxiang Wang, Haibo Zhou, Wei Lin, Hainan Zhang*

Meituan

cs.CL

Submitted: 2026-03-05

Updated: 2026-08-25

Importance score: 87/100

The gist: LocalSUG is an end-to-end generative query suggestion framework designed specifically for local-life service platforms, such as food delivery and hotel booking services.

Key concepts

LocalSUG
A system designed for better search suggestions in local services, such as food delivery. It leverages LLMs to incorporate city-specific knowledge into the query process, ensuring that results are highly relevant and up-to-date.
City-Preference-Enhanced LLM
An AI model that moves beyond general internet searches by understanding the user's immediate location and implied needs. It prioritizes suggestions based on local context, giving the system a 'sense of place.'
Compositional Ability
The system's capacity to synthesize new ideas for search queries, rather than just retrieving existing ones. It combines multiple inputs, such as user history and current location, to create novel and useful suggestions.

Terminology

Summary

LocalSUG is an end-to-end generative query suggestion framework designed specifically for local-life service platforms, such as food delivery and hotel booking services. It addresses the limitations of traditional multi-stage cascading architectures, which often struggle to explore long-tail, compositional, or emerging intents, by leveraging the semantic generalization of Large Language Models (LLMs) while overcoming critical deployment hurdles like city-specific relevance and real-time latency requirements.

The Challenges of LLM Deployment

Directly deploying LLMs for query suggestion in local-life platforms introduces three primary obstacles. First, there is insufficient city-preference awareness, as user intents vary by region due to merchant availability and local habits; without specific references, models may produce locally invalid suggestions. Second, the system faces exposure bias in preference optimization, where training on sequence-level historical logs creates a mismatch with real-world list-wise beam search decoding. Finally, strict online latency constraints remain a barrier because high-capacity LLMs incur substantial computational overhead that is often incompatible with large-scale production environments.

How it works

The LocalSUG framework utilizes Qwen3-0.6B as a backbone and employs a three-phase architecture to manage context, training, and deployment. The process begins with candidate mining and multi-source context construction, where city preferences are treated as dynamic external references rather than fusing them into model parameters. This allows the model to adapt to changing city preferences without stale memorization. The input is constructed by concatenating:

  • The user prefix p.

  • Candidate queries C cand fetched from a city-preference-enhanced co-occurrence cache.

  • Trending hot words W hot from internal platform services.

  • User behavior history H beh and user profiles U prof.

Optimization via Beam-Search-Driven GRPO

To bridge the training-inference gap, the authors introduce a Beam-Search-Driven GRPO algorithm that aligns training with inference-time decoding. This method optimizes sequence-level objectives through a multi-objective reward mechanism that balances relevance and business metrics. The unified reward function R(y i, Y G) incorporates several key components:

  1. Gap Shaping: Penalizing the tail candidates to force high-quality results into the top- K beam.

  2. Rank & Hit: Assigning a hit bonus and rank reward if the ground truth exists in the group, while penalizing non-target candidates ranked higher than it.

  3. Format & Miss Penalty: Penalizing format errors like repetition or invalid tokens, and penalizing instances where the target is missing from the top- K.

Inference Acceleration and Results

To enable large-scale deployment in high-concurrency industrial environments, LocalSUG implements Quality-Aware Accelerated Beam Search (QA-BS) and vocabulary pruning. QA-BS utilizes quality verification to prune active beams if cumulative log-probability falls below a threshold and employs adaptive termination to prevent high-latency tail processing. Combined with pruning the LM head to the top 30,000 most frequent tokens, these techniques achieve a 3.11× decoding speedup. Extensive online A/B testing demonstrates that LocalSUG reduces the few/no-result rate by 3.98%, boosts PV CTR by 0.35%, and increases Unique Item Exposure by 7.14%.

Improvements for AI systems

Improvement 1: Dynamic City-Preference Injection via Term Co-occurrence

  • Capability: The system can adapt instantaneously to evolving local environments—such as new merchant openings, store closures, or regional consumption shifts—without requiring model retraining. By injecting daily-updated, city-specific co-occurrence candidates into the prompt as external references rather than fusing them into model parameters, the AI ensures that suggestions remain locally valid and prevent stale or irrelevant recommendations.

Improvement 2: Beam-Search-Driven GRPO with Multi-Objective Reward Shaping

  • Capability: The system eliminates the exposure bias mismatch between sequence-level training and list-wise inference. By utilizing Group Relative Policy Optimization (GRPO) that samples outputs via beam search and applies a multi-objective reward function (incorporating Gap Shaping, Rank/Hit bonuses, and Format/Miss penalties), the AI produces more coherent, diverse, and correctly ranked suggestion lists that are directly optimized for business metrics like CTR and conversion.

Improvement 3: Quality-Aware Accelerated Beam Search (QA-BS) and Vocabulary Pruning

  • Capability: The system enables the deployment of large language models in high-concurrency, real-time production environments with strict latency constraints. By implementing adaptive branch pruning (based on cumulative log-probability thresholds), early termination when candidate capacity is saturated, and pruning the LM head to only high-frequency tokens, the AI achieves significant decoding speedups (up to 3.11×) while maintaining near-original generation quality and diversity.

Sources

Related papers