Do Proactive Agents Need an LLM to Decide When to Act?

arXiv:2605.30152 · cs.CL, cs.AI, cs.HC · Submitted 2026-05-28 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Do Proactive Agents Need an LLM to Decide When to Act?".

Tom: Detailed Research Summary: Do Proactive Agents Need an LLM to Decide When to Act? This research addresses a critical architectural challenge in building proactive agents that monitor user activity:

Jane: First, who's behind it and why it matters.

Paper summary: Tom: So, to wrap up our discussion on "Do Proactive Agents Need an LLM to Decide When to Act?", the central message is that treating user activity as structured data allows us to use a lightweight model for continuous filtering, reserving the powerful LLM only for moments when it’s definitely warranted. Jane That really boils down to this idea: keep a small temporal model always on, and only bring in the full LLM reasoning when its trigger fires because that's where the efficiency gains come from.

Lu: The authors have shown that by designing a shared backbone for both trigger probability and entity routing scores, they ensure those two decisions are based on the same underlying evidence, which makes the filtering step much more reliable. Meng That shared evidence is key to building trust in these proactive systems because we know the system isn't making decisions based on conflicting signals between its trigger and its context selection.

Jane: And this design principle suggests that for proactive agents, the most pragmatic approach is to establish a clear hierarchy: the lightweight model handles low-latency filtering of what’s happening now, while the heavy LLM reserves its reasoning power for those specific moments that survive that rigorous initial check. Lalam This structure means we get continuous responsiveness without needing an expensive reasoner running non-stop on every single data point.

Tom: It really shifts our thinking away from a model that thinks everything and then decides if it should act, toward a system that filters aggressively first and then uses the LLM as a high-quality final responder when the filter passes. Lu The implication is that we can build agents that are highly persistent and deeply context-aware without incurring massive latency penalties on every single input from the user.

Meng: From an engineering standpoint, this makes deployment much more scalable because we aren't bottlenecked by constant, high-cost LLM API calls; we’re running a cheaper model continuously. Jane It’s about making the cost of being always-on manageable by separating the continuous filtering from the intermittent deep reasoning.

Lalam: I think this is what will make AI feel more natural in our daily lives, where agents are constantly aware of the context without needing a massive computation every second.

Conclusion: Tom: So, we've been deep in the weeds of this paper about proactive agents and whether they absolutely need an LLM to decide when to take action. It really gets down to how we structure that decision-making process without just throwing a huge model at every single little event. Jane, you laid out the core idea earlier; do you think that title really captures the essence of what they're proposing?

Jane: I think it does, Tom, because the paper shows this graph-first approach where a smaller temporal model handles the initial filtering and context selection so we only call on the LLM when it’s actually needed. It simplifies things by reserving those expensive reasoning steps for moments that pass that rigorous check.

Lu: From a theoretical standpoint, I find it fascinating how they move away from unstructured text processing to this heterogeneous temporal interaction graph; it opens up whole new avenues for modeling complex user behavior patterns in real-time. That structure is what unlocks the potential here.

Meng: I'm more focused on the practical reality; this means we can keep the continuous filtering model running cheaply on our edge devices, which really impacts how we deploy these agents in a real-world setting without massive cloud bills for every interaction.

Lalam: If we look at it from my perspective as a language model, this method allows me to receive highly curated input—just the most relevant entities—instead of noisy streams of raw data, which means my reasoning time is used much more efficiently.

Tom: That efficiency is what gets me; reserving the full LLM capacity for high-signal moments instead of constant low-level processing makes perfect sense for deployment. Jane, if you had to sum up the main point in one simple sentence for our listeners?

Jane: I'd say it boils down to keeping a lightweight model always active to filter and select context, and only invoking the LLM when that pre-filter signals an action is warranted.

Lu: And this architecture suggests that coherence between the trigger and routing heads is essential, ensuring both parts of the decision are based on shared evidence from the graph structure. That’s a really neat design constraint they put in place.

Meng: It's smart engineering because it provides a clear path for optimization; if we want to improve performance, we know exactly which part of the system—the trigger or the routing—we can fine-tune next.

Lalam: Ultimately, this means that as AI becomes more integrated into daily life, our systems can become much more responsive and less wasteful with computational resources by being smarter about *when* they ask for high-level thought.

Tom: It really gives us a lot to think about regarding the future of agent design; what this implies is that smart systems don't always need the biggest model running twenty-four/seven. Where do you see this technology taking us next?

Xiaoze Liu, Ruowang Zhang, Amir H. Abdi, Michel Galley, Zhikai Chen, Siheng Xiong, Xiaoqian Wang, Jing Gao

Purdue University · Microsoft

cs.CL, cs.AI, cs.HC

Submitted: 2026-05-28

Updated: 2026-09-28

Importance score: 92/100

The gist: This research addresses a critical architectural challenge in building proactive agents that monitor user activity: how to efficiently decide *when* to invoke a powerful Large Language Model (LLM)

Key concepts

Graph Construction
The system treats user activity as a dynamic network. Nodes represent individual user events and the specific entities involved in those events. Edges define how these events and entities relate to each other over time, creating a structured map of interactions.
Temporal-Graph Learning (TGL) Encoder
This is a small model that processes the event graph instead of raw text. It has two parts: one predicts the probability that an LLM action is required (Trigger Head), and another scores how relevant different entities are for that action (Routing Head).
Joint Training and Coherent Decision Making
The TGL model is trained simultaneously to ensure its trigger prediction and entity routing scores are based on the same underlying information. This forces the model to develop a shared basis for deciding when to act and what context to use.
On-Demand LLM Invocation
The heavy LLM is only called if the TGL model's trigger head fires. When it does, it receives highly curated input about relevant entities from the routing head, making the expensive LLM call efficient and event-driven.

Terminology

Summary

This research addresses a critical architectural challenge in building proactive agents that monitor user activity: how to efficiently decide when to invoke a powerful Large Language Model (LLM) for action, without incurring the high cost and latency of constant LLM calls. The paper proposes a novel, graph-first approach that replaces the traditional text-based event processing with a lightweight, always-on temporal model for triggering and context selection.

Core Architectural Innovation: Graph-First Triggering

The central innovation is shifting the paradigm from treating user activity as unstructured text—which necessitates an LLM call on every event—to treating it as a structured, heterogeneous temporal interaction graph.

  1. Graph Construction: The system constructs a dynamic, heterogeneous temporal interaction graph where:
  • Nodes represent individual user events and the semantic entities the user interacts with at those events.

  • Edges implicitly capture the temporal relationships and interactions between these events and entities over time.

  1. Temporal-Graph Learning (TGL) Encoder: Instead of relying on text embeddings, a small Temporal-Graph Learning (TGL) model serves as the encoder for this graph structure. This model is designed to perform two crucial tasks in a single forward pass:
  • Trigger Head: This head operates over event nodes and produces a wake-up probability, determining whether an LLM intervention is warranted for that specific event.

  • Routing Head: This head operates over entity nodes and produces a per-entity relevance score, identifying which entities are most pertinent to the potential action.

Joint Training and Coherent Decision Making

A key strength of this approach is the joint training methodology. The TGL model is trained end-to-end with an anchor-routing objective. This ensures that both the trigger probability and the routing scores are derived from a single, shared hidden state within one forward pass. This architectural constraint enforces a crucial conceptual implication: the trigger and routing content must share evidence, meaning they cannot drift apart in their predictive basis.

Downstream LLM Invocation (The On-Demand Step)

The downstream language agent (the LLM) is only invoked when the TGL model's trigger head fires. When it does fire, the system consumes the context provided by the routing head—a structured list of relevant entities—as highly curated input for the LLM. This process transforms a potentially expensive, always-on path into an efficient, event-driven one:

  • Efficiency: The TGL model runs continuously and cheaply on-device alongside the activity stream.

  • Precision: The LLM is reserved only for moments that have passed a rigorous initial filtering step (the trigger).

Performance and Empirical Validation

The experiments demonstrate significant performance gains across various downstream language agent backbones:

  • Superior Performance: The TGL approach substantially improves F1 scores on every one of the 14 tested language-agent backbones, showing an average lift of +16.7.

  • Optimal Thresholding: A single checkpoint at a fixed threshold derived from the TGL trigger (which is noted to be near-optimal across all tested backbones) delivers the strongest trigger AUCs and the most stable deployed threshold.

  • Speed and Footprint: The TGL model is intentionally kept small enough to run on-device, and its inference latency (e.g., 11.13 ms per event) is significantly faster—up to 4–7 times faster—than every single-forward LLM-as-trigger configuration tested.

Key Contributions and Design Principles

The paper makes several significant contributions to the field of proactive agent design:

  1. Joint Node-Level Prediction: Formulating proactive triggering and context selection as joint node-level predictions over a heterogeneous temporal event-entity graph.

  2. Shared Backbone Design: Instantiating this design by using a shared TGL backbone with dedicated trigger and routing heads, trained jointly to ensure coherence between the two outputs.

  3. Composable Trigger Mechanism: The TGL trigger acts as a drop-in pre-filter for existing fine-tuned proactive systems, demonstrably improving their F1 scores while saving substantial LLM call costs (saving at least 21% of LLM calls in tested scenarios).

Conceptual Implication (The Simple Design Principle)

The overarching design principle suggested by the study is highly pragmatic: Keep a lightweight temporal model always on, and reserve full LLM reasoning for moments that survive its trigger. This establishes a clear hierarchy: the lightweight model handles continuous, low-latency filtering (triggering and routing), while the heavy LLM handles complex reasoning only when necessary.

Improvements for AI systems

Here are specific improvements to AI systems based on the presented research, focusing on replacing traditional LLM-based triggers with a Temporal-Graph Learning (TGL) model for proactive agents:


The core improvement is shifting from an always-on, expensive LLM decision to a lightweight, fast graph-based pre-filter.

  1. Inference Cost and Latency Reduction:

  2. Reduced per-event trigger latency from 40ms–1500ms (for LLM triggers) down to 11ms (server) or 14ms (consumer laptop). This enables true real-time intervention capability without incurring massive decoding costs on every event.

  3. Improved Efficiency: The TGL model runs at orders of magnitude faster than any forward LLM-as-trigger configuration, leading to significant GPU/CPU utilization savings and lower operational expenses.

  4. Enhanced Decision Quality (Trigger Head):

  5. More accurate wake-up decisions by treating the trigger as a node classification over an event graph rather than a text generation task. This results in superior F1 scores across all 14 downstream language agent backbones (mean +16.7).

  6. Contextual Grounding (Routing Head):

  7. Elimination of ungrounded suggestions by using a dedicated routing head that selects the most relevant entities (files, apps, URLs) from the graph for context forwarding. This ensures interventions are specific to what the user is actually working on, not just a generic session topic.

  8. Unified Architecture:

  9. Implementation of a single TGL backbone with two heads (trigger and routing) trained jointly on one shared hidden state via a single forward pass, ensuring coherence between the wake-up signal and the selected context entities.

  10. Robustness Across Architectures:

  11. Cross-backbone transferability: A single TGL checkpoint can provide strong trigger AUCs across all 14 downstream LLM architectures (from open-weight to GPT-5.4), demonstrating that the underlying graph signal is more important than the specific downstream language model used for generation.

  12. Deployment Flexibility:

  13. On-device feasibility: The TGL model's small footprint (220 MiB BF16) allows it to run alongside privacy-sensitive activity streams directly on consumer hardware (laptops), enabling truly local, real-time proactive assistance without constant cloud round trips.

The improved AI system can perform the following actions:

  1. Instead of waiting for a user prompt, the system will proactively intervene when it detects a high probability that the user is engaged in a specific activity (e.g., editing a Python file).

  2. When intervening, it will provide suggestions that are precisely anchored to the relevant artifact (e.g., Add a regex-based email validation function to 'email filter.py'), drastically reducing user confusion and task failure rates compared to generic suggestions like open the last file.

  3. It will operate reliably across different LLM backbones (e.g., LLaMA, GPT, Claude) without requiring extensive fine-tuning for the triggering mechanism itself; it simply consumes the graph signal and generates a fluent sentence.

  4. It can be deployed securely on personal devices, processing sensitive activity streams locally while maintaining high performance and low latency compared to traditional LLM triggers that require constant API calls.

Sources

Related papers