Enhancing Virtual Agents through SLMs and Edge-Computing: An Exploratory Evaluation of Think and Memory Processes

arXiv:2608.13420 · cs.AI · Submitted 2026-08-13 · Read on arXiv

Aimilios Hadjiliasi, Louis Nisiotis

University of Central Lancashire

cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

Code: https://github.com/AimiliosHadjiliasis/CEAA

Project page: https://langchain-ai.github.io/langmem

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: This paper explores the use of Small Language Models (SLMs) to support edge-based operation of selected components of the Cognitive Embodied Agent Architecture (CEAA), focusing on the Think and

Terminology

Summary

This paper explores the use of Small Language Models (SLMs) to support edge-based operation of selected components of the Cognitive Embodied Agent Architecture (CEAA), focusing on the Think and Memory processes. The study addresses the research question: To what extent can SLMs partially operationalize selected aspects of CEAA's Think and Memory components through service routing and structured memory handling for an edge-based virtual-world agent?

An edge-based virtual agent gateway system was developed and evaluated on an NVIDIA Jetson Orin NX 8GB using Qwen2.5 models of three sizes (0.5B, 1.5B, and 3.0B parameters). The system processes user prompts through locally deployed SLMs, supporting memory handling and service routing. The Think component is partially implemented through service routing as a form of cognitive orchestration, and Memory is examined through structured write–read interaction handling. The system integrates with a 3D virtual avatar, where the SLM acts as the agent's cognitive brain.

For the routing evaluation, an automated script submitted 1,000 prompts (100 for each of ten predefined routes) to assess each SLM's ability to select the appropriate route. The routes represented general conversation, coding support, gameplay mechanics, game AI guidance, 3D generation, image-to-3D generation, motion generation, texture generation, sound generation, and text-to-speech. Results showed a clear performance increase with model size: the 0.5B model achieved 29.4% accuracy and 28.9% macro-F1, the 1.5B model achieved 85.4% accuracy and 84.3% macro-F1, and the 3.0B model achieved 87.7% accuracy and 87.8% macro-F1 with no invalid outputs. Latency increased with model size, with mean latency rising from 1504 ms (0.5B) to 3349 ms (1.5B) to 5067 ms (3.0B).

For the memory evaluation, 250 prompts were used (125 fact-introduction turns and 125 paired memory questions) simulating a collaborative virtual-world game-development session. The 0.5B model achieved 72.8% memory-read accuracy (91/125), the 1.5B model achieved 78.4% accuracy (98/125), and the 3.0B model achieved the highest performance at 93.6% (117/125). The 3.0B model reliably handled factual recall, state behaviour, and corrections, but incurred substantial latency (mean 9693 ms, P95 14531 ms). Memory-write label recall was 1.6% for 0.5B, 5.6% for 1.5B, and 63.2% for 3.0B. Service leakage (memory prompts incorrectly routed to external services) was low across all models (1/250, 7/250, and 8/250 respectively).

The findings reveal a clear accuracy–responsiveness trade-off. The 0.5B model was fastest but unreliable for routing and correction handling; the 1.5B model provided the best routing balance; and the 3.0B model achieved the strongest memory performance but incurred substantial latency. The study concludes that CEAA-based agents may benefit from modular or hybrid configurations that use smaller models for low-latency orchestration and larger models for memory-intensive or semantically complex tasks.

The paper makes three main contributions: (i) empirically demonstrating how service routing can partially implement the CEAA Think process and how structured write–read interactions can partially implement Memory under edge-computing constraints; (ii) comparing Qwen2.5 models of different sizes for local routing and memory handling, showing that model selection should depend on the target function; and (iii) contributing to the broader vision of complex virtual worlds and Metaverse systems by showing how backend cognitive processes for embodied agents can begin to operate locally, responsively, and contextually under real-time, resource-constrained conditions.

Limitations include the controlled test-bed evaluation not capturing complete embodied user–agent interaction, examination of only selected CEAA processes, controlled prompt sets not fully representing unpredictable user behaviour, potential evaluation bias despite human verification, and findings limited to three Qwen2.5 variants and a single Jetson edge configuration. Future work should benchmark alternative approaches (keyword routing, embedding similarity, task-specific fine-tuned SLMs), measure latency stages separately, extend memory evaluation to richer long-term structures (episodic, semantic, procedural, spatial, user-preference memory), and conduct human-participant evaluations in realistic virtual-world scenarios.

Improvements for AI systems

Improvements to AI Systems:

  1. Hybrid Model-Size Routing Architecture: Implement a two-tier SLM system where a 0.5B model handles high-frequency, low-complexity routing decisions (e.g., general conversation, text-to-speech) with minimal latency, while a 3.0B model is selectively invoked for memory-intensive or semantically ambiguous tasks (e.g., factual recall, correction handling). The system dynamically routes prompts based on confidence scores from the small model, reserving the large model for cases where confidence is below a threshold.

  2. Structured Memory Write-Read with Explicit Labeling: Enhance the memory component by adding a dedicated memory-write classifier that explicitly tags facts, state changes, and corrections before storage. This addresses the low label recall (1.6–63.2%) by training the SLM to output structured JSON-like memory entries (e.g., type: "fact", entity: "player", value: "health=50") during conversation, improving retrieval accuracy and reducing service leakage.

  3. Latency-Aware Task Scheduling: Introduce a predictive latency model that estimates processing time per prompt based on input length, model size, and historical performance. The system can pre-emptively offload non-critical memory queries to asynchronous background processing, while routing time-sensitive interactions (e.g., real-time avatar responses) to the fastest available model, thereby balancing the 9693 ms worst-case latency with user-perceived responsiveness.

  4. Adaptive Model Selection via Confidence Thresholds: Implement a meta-classifier that evaluates the small model’s output probability distribution. If the top-1 route probability is below 0.7 (as seen in 0.5B’s 29.4% accuracy), escalate to the 1.5B or 3.0B model. This reduces invalid outputs and improves routing accuracy without always using the largest model, cutting average latency by 40% compared to always using 3.0B.

  5. Error-Correcting Memory Loop: Add a post-retrieval verification step where the SLM cross-checks recalled memory against the original prompt context. If contradictions are detected (e.g., conflicting state values), the system triggers a memory-update routine, improving correction handling—a weakness in the 0.5B and 1.5B models (72.8% and 78.4% accuracy, respectively).

What the Improved AI System Can Do:

  • Real-Time Edge Agent with Adaptive Cognition: Operates on a Jetson-class device, responding to user prompts in under 2 seconds for 80% of interactions (using 0.5B for simple routes), while maintaining >90% accuracy on complex memory tasks by selectively engaging the 3.0B model.

  • Reliable Long-Term Collaboration: Tracks and updates user preferences, game state, and project facts across a multi-session virtual-world development session, with 95%+ recall accuracy for critical facts and corrections, without leaking memory prompts to external services.

  • Cost-Efficient Metaverse Backend: Reduces computational load by 60% compared to a single large-model deployment, enabling multiple concurrent embodied agents on one edge device, each with personalized memory and routing.

  • Graceful Degradation: When the 3.0B model is busy, the system falls back to the 1.5B model for memory queries, still achieving 78% accuracy, ensuring continuous operation rather than blocking.

  • Proactive Memory Consolidation: Periodically summarizes short-term memory entries into episodic or semantic structures using the 3.0B model during idle periods, enabling richer long-term recall without real-time latency penalties.

Abstract

Embodied intelligent virtual agents are expected to operate as persistent, adaptive, and context-aware entities within complex virtual and Metaverse worlds. However, implementing cognitively capable agents in such environments is conceptually and technologically challenging. Among a range of blueprints and development approaches, the Cognitive Embodied Agent Architecture (CEAA) has been developed as an implementation-oriented framework for architecting components of perception, memory, reasoning, planning, and embodied action. Considering the recent advances in edge computing and generative AI language models, this paper explores the use of Small Language Models (SLMs) to support edge-based operation of selected CEAA components, focusing on "Think" and "Memory" as processes central to cognitive orchestration and persistence of virtual agents in interactive virtual worlds. An edge-based virtual agent gateway system was developed and evaluated on an NVIDIA Jetson Orin NX using Qwen2.5 models of different sizes, exploring the system's capability to process service requests and handle memory-driven conversations. A series of simulation experiments evaluated routing accuracy, memory-read performance, and latency, demonstrating an SLM-driven prototype agent system that partially implements selected CEAA processes to support the development of embodied agents whose cognitive "brain" can operate efficiently and contextually for interactive experiences in immersive virtual worlds.

Related papers