Agent / long-term memory
Zep
Zep: A Temporal Knowledge Graph Architecture for Agent Memory
Heavily superseded — a standard baseline that newer methods routinely beat
1 papers critique it · 9 beat it on benchmarks
What papers say
Verbatim critique sentences, each from a paper that cites Zep as a baseline.
LLMs cannot use memory tools effectively and using such tools increases redundancy in planning and overall tool use.
Beaten on benchmarks
Head-to-head results where a newer method reports beating Zep. Values are copied from the source paper's tables — verify against the cited paper.
HingeMem beats Zep
63.9 vs 16.3
SwiftMem beats Zep
0.569 vs 0.200
BLEU-1 · [Temporal Reasoning]
SwiftMem: Fast Agentic Memory via Query-aware IndexingChronos beats Zep
91.73 vs 57.90
MS · [GPT-4o model, multi-session aggregation]
Chronos: Temporal-Aware Conversational Agents with Structured Event Retrieval for Long-Term MemoryMem0 beats Zep
0.708 vs 1.292
Latency total p50 (seconds) · [overall]
Mem0: Building Production-Ready AI Agents with Scalable Long-Term MemorySmartSearch beats Zep
88.4 vs 63.8
Overall · [full benchmark]
SmartSearch: How Ranking Beats Structure for Conversational Memory RetrievalMnemosyne beats Zep
54.55 vs 42.80
Overall (%) · [LoCoMo benchmark]
Mnemosyne: An Unsupervised, Human-Inspired Long-Term Memory Architecture for Edge-Based LLMsGRAVITY beats Zep
60.9 vs 47.8
LLM-judge accuracy · [LongMemEval Macro]
GRAVITY: Architecture-Agnostic Structured Anchoring for Long-Horizon Conversational MemoryDeltaMem beats Zep
75.13 vs 65.99
gpt-5+NoMem beats Zep
0.206 vs 0.214
gemini-2.5-pro+NoMem beats Zep
0.144 vs 0.140
Correctness · [Gemini-2.5-pro]
MEMTRACK: Evaluating Long-Term Memory and State Tracking in Multi-Platform Dynamic Agent Environments
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.
- Jun 9, 2026
- May 30, 2026
- MemGuardMemGuard: Preventing Memory Contamination in Long-Term Memory-Augmented Large Language ModelsMay 27, 2026
- DeferMemDeferMem: Query-Time Evidence Distillation via Reinforcement Learning for Long-Term Memory QAMay 21, 2026
- May 20, 2026
- May 3, 2026
- Apr 23, 2026
- Apr 2, 2026
- ChronosChronos: Temporal-Aware Conversational Agents with Structured Event Retrieval for Long-Term MemoryMar 17, 2026
- Mar 15, 2026
- Jan 13, 2026
- Agentic Memory (AgeMem)Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model AgentsJan 5, 2026