Agent / long-term memory
Mem0
Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
Heavily superseded — a standard baseline that newer methods routinely beat
3 papers critique it · 15 beat it on benchmarks
What papers say
Verbatim critique sentences, each from a paper that cites Mem0 as a baseline.
While graph databases provide structured organization for memory systems, their reliance on predefined schemas and relationships fundamentally limits their adaptability.
“LLMs cannot use memory tools effectively and using such tools increases redundancy in planning and overall tool use.”
“they operate in an ``open loop'' without feedback on whether the constructed memories benefit downstream tasks”
Beaten on benchmarks
Head-to-head results where a newer method reports beating Mem0. Values are copied from the source paper's tables — verify against the cited paper.
MemGuard beats Mem0
71.49 vs 24.50
Engrama beats Mem0
0.5997 vs 0.2848
SwiftMem beats Mem0
11 vs 784
Search latency · [Overall]
SwiftMem: Fast Agentic Memory via Query-aware IndexingMemBuilder beats Mem0
85.75 vs 47.00
accuracy · [LongMemEval benchmark]
MemBuilder: Reinforcing LLMs for Long-Term Memory Construction via Attributed Dense RewardsHingeMem beats Mem0
63.9 vs 35.5
GAM beats Mem0
42.55 vs 31.73
LoCoMo Multi Hop F1 · [Qwen2.5-14b]
General Agentic Memory Via Deep ResearchSmartSearch beats Mem0
88.4 vs 66.4
Overall · [full benchmark]
SmartSearch: How Ranking Beats Structure for Conversational Memory RetrievalDeltaMem beats Mem0
75.13 vs 57.86
AgeMem beats Mem0
54.31 vs 41.95
Average · [Qwen3-4B-Instruct]
Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model AgentsGRAVITY beats Mem0
64.0 vs 51.7
LLM-judge accuracy · [LongMemEval Macro]
GRAVITY: Architecture-Agnostic Structured Anchoring for Long-Horizon Conversational MemoryGAM beats Mem0
12.55 vs 10.27
Avg F1 · [Qwen 2.5-7B]
GAM: Hierarchical Graph-based Agentic Memory for LLM Agentsgemini-2.5-pro+NoMem beats Mem0
0.144 vs 0.118
Correctness · [Gemini-2.5-pro]
MEMTRACK: Evaluating Long-Term Memory and State Tracking in Multi-Platform Dynamic Agent Environments
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.
- Jun 9, 2026
- May 30, 2026
- MemGuardMemGuard: Preventing Memory Contamination in Long-Term Memory-Augmented Large Language ModelsMay 27, 2026
- DeferMemDeferMem: Query-Time Evidence Distillation via Reinforcement Learning for Long-Term Memory QAMay 21, 2026
- May 20, 2026
- May 3, 2026
- Apr 23, 2026
- Apr 2, 2026
- ChronosChronos: Temporal-Aware Conversational Agents with Structured Event Retrieval for Long-Term MemoryMar 17, 2026
- Mar 15, 2026
- Jan 13, 2026
- Agentic Memory (AgeMem)Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model AgentsJan 5, 2026