Harness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents
Abstract
Empirical evaluation of diverse memory substrates for long-horizon LLM agents reveals regime-dependent trade-offs, motivating adaptive substrate routing for reliable agent memory.
Memory is becoming core infrastructure for long-horizon LLM agents, yet existing evaluations offer limited guidance on which memory substrate, namely the underlying medium in which memory is represented and stored, should be used under different operating regimes. We present a controlled harness evaluation of memory substrates for memory-augmented agents, covering dense and sparse indices, text records, structural stores, hierarchical stores, refinement-based memories, parametric updates, and activation-compatible context mechanisms. Across three backbone models and four benchmark suites spanning user-centric question answering and agent-centric decision-making, we instrument 26 performance and efficiency metrics under a unified harness. Our results show that no single substrate consistently dominates: broad retrieval benefits long-context factual QA, while excessive retrieval can harm sequential decision-making by shifting attention away from action-critical context. Scalability introduces a further routing axis, as substrates that perform well at moderate history lengths can become costly or brittle at longer horizons. These findings motivate substrate routing as a necessary component of adaptive agent memory systems and provide empirical guidance for designing efficient, reliable, and regime-aware long-term memory for LLM agents. Code will be made available upon acceptance.
Community
We present a controlled evaluation of memory substrates for memory-augmented LLM agents, covering dense and sparse retrieval, text and structural stores, hierarchical and refinement-based memories, parametric updates, and activation-compatible context mechanisms.
Across three backbone models and four benchmark suites, we evaluate 26 performance and efficiency metrics under a unified harness. Our results show that no single memory substrate consistently dominates: different substrates excel under different tasks, context lengths, and operating regimes. These findings highlight substrate routing as an important direction for building adaptive and efficient long-term memory systems for LLM agents.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Beyond Memory Leaderboards: Evaluating Scientific Memory as Budgeted Context Restoration (2026)
- SelfMem: Self-Optimizing Memory for AI Agents (2026)
- Oracle Agent Memory as an Enterprise Memory Substrate for Long-Horizon AI Agents (2026)
- LycheeMemory V2: Efficient Long-Term Memory for LLM Agents via Semantic Segment-Level Consolidation (2026)
- What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents (2026)
- Memory as a Controlled Process: Learned Adaptive Memory Management for LLM Agents (2026)
- TrajWiki: Source-Grounded Memory Trajectories for Long-Horizon Dialogue Agents (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.15008 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper