When Your Agent Opens the Chat App: Agent-Controlled Search over Raw Chat Logs Rivals Structured Memory

arXiv:2608.12888 · cs.CL · Submitted 2026-08-16 · Read on arXiv

Ruizhe Li, Licheng Zhang, Benfeng Xu, Mingxuan Du, Zheren Fu, Weidong Chen

University of Science and Technology of China · MetaStone Technology

cs.CL

Submitted: 2026-08-16

Updated: 2026-08-18

License: http://creativecommons.org/licenses/by/4.0/

The gist: The paper introduces ReFind, an agent-controlled search interface for conversational memory that deliberately builds no semantic structure over raw chat logs.

Terminology

Summary

The paper introduces ReFind, an agent-controlled search interface for conversational memory that deliberately builds no semantic structure over raw chat logs. The central research question is: how much of their reported benefit comes from the structure itself, rather than from having competent retrieval over the history at all? ReFind "leaves the conversation archive unmodified, indexes it lexically at turn granularity, and combines a generic iterative keyword-search loop with four chat-native controls grounded in empirical refinding work: session-aware rank fusion, local context expansion, temporal narrowing, and skipping already-inspected sessions." A separate reasoning stage answers from the collected evidence.

The system architecture has two decoupled stages: Stage 1 (retrieval) uses a ReAct-style agent loop that lets the LLM autonomously decide search keywords and parameters, conducting multiple rounds of search over the history and saves valuable fragments as notes. Stage 2 (reasoning) groups the collected notes by session, sorts them chronologically, and presents them with the question for answer generation. The search engine uses BM25 at turn granularity with k1=1.2, b=0.75, and applies Reciprocal Rank Fusion (RRF) with k=60 to combine turn-level and session-level rankings. The four chat-native controls are: context window expansion (returns ±2 neighboring turns around each hit, truncated at session boundaries), temporal filtering (filters turns by timestamp when the agent supplies a date range), session deduplication (excludes sessions returned in prior rounds), and RRF two-level reranking.

The main results show ReFind attains the highest mean accuracy (58.2) of any system compared, above the strongest graph- and tree-based memory systems (HippoRAG 2, 53.2), all under a GPT-4o-mini backbone matched to every reused baseline. Across six benchmarks (single-/multi-hop QA, LongMemEval, EventQA, and single-/multi-hop FactConsolidation) totaling roughly 2,800 questions, ReFind is best on five of six benchmarks: SH-QA 83.0, MH-QA 69.0, LME 51.3, FC-SH 62.7, and FC-MH 8.8. On EventQA, ReFind and BM25-RAG are closely matched (74.1 and 74.6). On the LongMemEval-S/M subsets with GPT-5-mini, ReFind reaches 93.2 ± 3.3 and 89.3 ± 6.0 across five runs, above every backbone-matched structured system including STITCH (86.0/80.0) and GAM (70.0/60.0).

The ablation study decomposes the gains. A matched Generic Agentic BM25 control (retaining the multi-round controller but removing all four chat-native controls) reaches 78.7 ± 4.6 (S) and 82.2 ± 3.8 (M), below the full method by 14.5 and 7.1 points. Component ablations show: context window removal is the largest S drop (−9.2 points), session deduplication matters more on M (−9.3 points) than S (−1.2 points), RRF reranking removal costs −3.9/−4.9 points, and temporal filtering has the smallest S effect (−1.9 points). A one-search control (keeping the chat-native retrieval behavior for the first query but removing reaction to returned evidence) trails the full method by 8.5 (S) and 20.4 (M) points. Backend comparisons show BM25 has the highest mean on both subsets, with Dense and Hybrid variants not exceeding BM25's mean.

The paper concludes that preserving raw records and exposing their conversational structure to an adaptive agent outperforms first transforming them into a memory representation. The key insight is that much of the benefit credited to elaborate memory structures is recoverable by giving an agent controllable search over the unmodified record, with no LLM-based index construction at all. The system ties model computation to expressed information needs rather than to every archive update, averaging 2.5–2.6 searches and 5.0 LLM calls per query. The paper suggests a modular design principle: begin with faithful storage and a controllable search interface, then add derived structures only for workloads that demand a separate latency or abstraction layer.

Improvements for AI systems

Improvements to AI Systems:

  1. Add a raw-first retrieval mode – Instead of forcing all conversational history through semantic/graph indexing, AI systems can keep the original chat logs untouched and use a lexical BM25 index at turn granularity. This avoids information loss from abstraction and reduces compute overhead from index construction.

  2. Implement an agentic multi-round search loop – The AI system can autonomously decide search keywords and parameters, iteratively refine queries based on returned evidence, and save relevant fragments as notes. This ties model computation to the user's expressed information need, not to every archive update, reducing unnecessary processing.

  3. Integrate four chat-native retrieval controls:

  • Session-aware rank fusion: Combine turn-level and session-level BM25 rankings using Reciprocal Rank Fusion (k=60) to improve relevance across conversational boundaries.

  • Local context expansion: Return ±2 neighboring turns around each hit (truncated at session boundaries) to preserve conversational context without building explicit structure.

  • Temporal narrowing: Allow the agent to filter turns by timestamp when the user specifies a date range, improving precision for time-sensitive queries.

  • Session deduplication: Exclude already-inspected sessions in subsequent search rounds to avoid redundant evidence and speed up convergence.

  1. Decouple retrieval from reasoning – Use a two-stage pipeline: Stage 1 (retrieval) collects evidence via the agentic loop; Stage 2 (reasoning) groups notes by session, sorts chronologically, and generates answers from that evidence. This separation improves accuracy and interpretability.

  2. Adopt a generic agentic BM25 baseline – For any new memory system, first evaluate a simple agentic BM25 without chat-native controls to isolate the true benefit of added structure. This prevents over-attributing gains to complex architectures.

  3. Prioritize BM25 over dense/hybrid backends – When building conversational memory retrieval, use BM25 (k1=1.2, b=0.75) as the default backend, as it outperformed dense and hybrid variants in the paper. Only add dense retrieval if a specific workload demands it.

What the Improved AI System Can Do:

  • Answer multi-hop questions over long conversational histories with higher accuracy (e.g., 58.2 mean vs. 53.2 for best structured system) without building any semantic index.

  • Handle single- and multi-hop QA, long-term memory evaluation, and fact-consolidation tasks across 2,800 questions, achieving state-of-the-art results on 5 of 6 benchmarks.

  • Maintain high performance even with a small LLM backbone (GPT-4o-mini or GPT-5-mini), reaching 93.2% on LongMemEval-S and 89.3% on LongMemEval-M.

  • Reduce computational cost by averaging only 2.5–2.6 searches and 5.0 LLM calls per query, tying computation to user needs rather than archive size.

  • Provide a modular design: start with faithful storage and controllable search, then add derived structures only for workloads requiring a separate latency or abstraction layer.

Abstract

Agent-memory systems increasingly buy retrieval quality with structure, transforming raw conversation histories into summaries, embeddings, trees, or knowledge graphs before any question is asked. We ask how much of that benefit comes from the structure itself, rather than from competent retrieval over the raw history. We present ReFind, an agent-controlled search interface that builds no semantic structure at all: it leaves the conversation archive unmodified, indexes it lexically at turn granularity, and combines a generic iterative keyword-search loop with four chat-native controls grounded in empirical refinding work: session-aware rank fusion, local context expansion, temporal narrowing, and skipping already-inspected sessions. A separate reasoning stage answers from the collected evidence. Across a broad suite of conversational-memory tasks (single- and multi-hop QA, event ordering, and fact consolidation), roughly 2,800 questions on precise-retrieval and fact-tracking capabilities evaluated under the incremental multi-turn setting of MemoryAgentBench, ReFind attains the highest mean accuracy (58.2) of any system compared, above the strongest graph- and tree-based memory systems (HippoRAG 2, 53.2), all under a GPT-4o-mini backbone matched to every reused baseline. Controlled comparisons to single-shot BM25, a matched generic-agentic BM25 control, component removals, and agentic dense/hybrid variants separately support the roles of agent control, chat-native controls, and lexical retrieval. On LongMemEval-S/M, the same interface reaches 93.2 +/- 3.3 and 89.3 +/- 6.0 with GPT-5-mini. The results indicate that for precise, evidence-grounded questions over chat archives, much of the benefit credited to elaborate memory structures is recoverable by giving an agent controllable search over the unmodified record, with no LLM-based index construction at all.

Sources

Related papers