Total Recall at What Cost? Benchmarking the Serving Cost of Agentic Memory Systems

arXiv:2608.11879 · cs.CL, cs.IR · Submitted 2026-08-12 · Read on arXiv

Natchanon Pollertlam, Witchayut Kornsuwannawit

Bricks Technology

cs.CL, cs.IR

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: 11 pages, 2 figures, 8 tables

Code: https://github.com/openai/tiktoken

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

Terminology

Summary

Summary

This paper benchmarks the serving cost of agentic memory systems for long-running conversational agents. The authors compare three memory systems—Mem0, Hindsight, and Mastra Observational Memory—against two reference strategies: a fixed-size rolling window (floor) and full-transcript resubmission (ceiling). All systems are evaluated across two backbone models (gpt-oss-20b and Gemma 4 26B A4B) at two reasoning-effort settings (low, medium), on synthetic conversations of up to 400 turns. Each cost measurement is paired with answer accuracy on 665 LoCoMo questions.

Key findings:

  1. Cost cannot be predicted from conversation length and message size alone. The authors fit a separable cost model: log(C+1) = a + p log(L+1) + q log(t+1), where p is the message-size exponent and q is the depth exponent. The two baselines fit well: full-history has p≈q≈1 (cost proportional to cumulative content), and rolling-window has p≈0.9, q≈0.1 (nearly flat in depth). However, the memory systems have high held-out error (LOOCV-MAPE): 18–22% for Mem0, 46–48% for Hindsight, and 41–69% for Mastra OM. The paper states: "A low in-sample R2 on its own does not show that a system has separated cost from conversation size. Mastra OM has the highest in-sample R2 of the three memory systems, yet it has the worst held-out error. The held-out test shows that per-turn cost depends on internal memory state."

  2. Break-even analysis shows high sensitivity to system and backbone. The break-even length (when a memory system becomes cheaper than full-transcript resubmission) spans from turn 0 for the cheapest systems to never within 400 turns for the most expensive. Specifically: Mastra OM breaks even at turn 0, Mem0 at turn 82, and Hindsight only at turn 356 under gpt-oss-20b at 200 tokens per turn. Across all measured 400-turn cells, break-even spans Mastra OM 0–86, Mem0 0–342, and Hindsight 60–never. By turn 400, the full transcript costs up to 12.7× a memory system that has broken even.

  3. No system wins on both cost and accuracy. Accuracy spans 21–54% across systems and settings. Mem0's accuracy varies the most (0.214–0.516). The lowest cost-per-correct-answer is Mastra OM on gpt-oss-20b low (0.028 at the 100-turn reference cell, 0.278 at the 400-turn cell), while Mem0 leads on Gemma 4 26B A4B (0.037–0.038 at 100 turns, 0.325–0.339 at 400 turns). Hindsight is the most expensive at the reference cell (≈0.24 per 100-turn conversation under gpt-oss-20b).

  4. Backbone choice drives cost as much as the memory system does. The paper notes: Mem0's reference-cell cost falls from 0.059– 0.065 under gpt-oss-20b to 0.019 under Gemma 4 26B A4B. However, Mastra OM behaves in the opposite way—its gpt-oss-20b low cell is the cheapest Mastra OM setting. The paper concludes: backbone and memory are a joint decision.

  5. Raising reasoning effort does not always improve accuracy. "Mem0's accuracy on gpt-oss-20b drops from 0.322 to 0.214 at the higher reasoning level. This is likely because reasoning tokens use up the max tokens budget and leave less room for the answer. Mastra OM moves in opposite directions on the two backbones (+7.3 pp on Gemma 4 26B A4B, −5.3 pp on gpt-oss-20b)."

Limitations noted: The cost benchmark uses synthetic dialogues; accuracy is reported on LoCoMo only; the cost model is descriptive, not mechanistic; Hindsight's ingest backbone was not benchmark-controlled (it ran under a fixed configuration across all settings); provider-routing variance can affect costs by single-digit percent; and a serving-stack error truncated gpt-oss-20b full-history measurements at turn 374 of the 400-turn cell.

Conclusion: The main cost finding is that the model predicts well for window-based strategies but not for memory systems... This shows that their cost is driven by internal memory state, not by conversation length or message size. The choice of memory system depends on the expected conversation length, not on the memory system alone.

Improvements for AI systems

Improvements to AI Systems:

  1. Adaptive memory-selection controller: Build a router that predicts conversation depth and message size at runtime, then selects between Mem0, Hindsight, Mastra OM, or a rolling window based on the fitted cost model’s break-even thresholds (e.g., switch to Mastra OM for short chats, Mem0 for mid-length, Hindsight only for very long). This reduces cost by up to 12.7× at turn 400 versus full-history, while avoiding the “never breaks even” trap for expensive systems.

  2. Memory-state-aware cost predictor: Replace length-based cost estimation with a learned regressor that takes internal memory state features (e.g., number of stored entities, retrieval hit rate, compaction events) as inputs, since the paper shows length/message-size alone yields 18–69% held-out error. This enables accurate per-turn cost forecasting for budgeting and auto-scaling in production agents.

  3. Backbone-memory co-optimizer: Given that backbone choice flips cost rankings (Mem0 is 3× cheaper on Gemma 4 26B A4B, while Mastra OM is cheapest on gpt-oss-20b low), implement a joint selection algorithm that evaluates both dimensions together—e.g., a lookup table or small ML model that recommends (backbone, memory system, reasoning effort) for a given expected conversation length and accuracy target, rather than treating them independently.

  4. Reasoning-effort budget allocator: Detect when raising reasoning effort degrades accuracy (as seen with Mem0 on gpt-oss-20b, where accuracy dropped from 0.322 to 0.214 due to token budget exhaustion). Add a safeguard that monitors max tokens consumption and dynamically lowers reasoning effort or increases token budget when the answer generation is likely to be truncated, preserving accuracy without cost blowup.

  5. Cost-per-correct-answer optimizer: Use the paper’s cost-per-correct-answer metric (e.g., Mastra OM at 0.028 on gpt-oss-20b low vs. Mem0 at 0.037–0.038 on Gemma) to drive an online reinforcement-learning loop that adjusts memory system, backbone, and reasoning effort per conversation cluster, maximizing accuracy per dollar spent rather than raw accuracy or raw cost.

  6. Break-even-aware conversation-length estimator: For new conversations, estimate the expected total turns (from user behavior patterns or intent classification) and use the break-even curves (turn 0–356 across systems) to pre-commit to the cheapest memory system that will remain cost-effective for the full conversation, avoiding mid-conversation migrations that incur overhead.

What the improved AI system can do: It can run long-horizon conversational agents (e.g., customer support, personal assistants, coding copilots) at 3–12× lower serving cost while maintaining or improving answer accuracy, by dynamically choosing memory architecture, backbone model, and reasoning effort based on predicted conversation length, internal memory state, and token-budget constraints—rather than using a fixed configuration.

Abstract

Long-running conversational agents increasingly rely on a memory system to avoid resending the whole conversation each turn, yet how much that costs to serve has received little systematic benchmarking. We compare three memory systems (Mem0, Hindsight, and Mastra Observational Memory) against two reference strategies -- a fixed-size rolling window and resubmitting the full transcript -- across two backbones and conversations of up to 400 turns, pairing every cost measurement with answer accuracy on 665 LoCoMo questions. First, a memory system's serving cost cannot be predicted from conversation length and message size alone: a regression that tracks the two reference strategies closely misses the memory systems by 18-69%, their cost driven instead by internal memory behavior. Second, a break-even analysis shows that whether -- and when -- a memory system becomes cheaper to serve than the full transcript is highly sensitive to the system and the backbone, from the first tens of turns for the cheapest to never within 400 turns for the most expensive. Third, no system wins on both axes: accuracy spans 21-54%, and the backbone choice drives cost as much as the memory system does.

Sources

Related papers