Agent / long-term memory
LightMem
LightMem: Lightweight and Efficient Memory-Augmented Generation
Superseded — cited as a baseline and beaten by newer methods
1 papers critique it · 4 beat it on benchmarks
What papers say
Verbatim critique sentences, each from a paper that cites LightMem as a baseline.
Taking LightMem on LoCoMo as an example: open-domain (75.9% accuracy) and single-hop (70.7% accuracy) questions are handled well, but multi-hop reasoning drops to 60.6% and temporal reasoning to just 45.8%.
Beaten on benchmarks
Head-to-head results where a newer method reports beating LightMem. Values are copied from the source paper's tables — verify against the cited paper.
DeferMem beats LightMem
0.00 vs 28.25
Token cost · [LongMemEval-S]
DeferMem: Query-Time Evidence Distillation via Reinforcement Learning for Long-Term Memory QADeltaMem beats LightMem
66.43 vs 58.38
QA · [Question Answering]
DeltaMem: Towards Agentic Memory Management via Reinforcement LearningGRAVITY beats LightMem
75.8 vs 70.1
LLM-judge accuracy · [LoCoMo]
GRAVITY: Architecture-Agnostic Structured Anchoring for Long-Horizon Conversational MemoryGAM beats LightMem
73.59 vs 69.51
LongBenchv2 F1 · [GPT-4o-mini]
General Agentic Memory Via Deep Research
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.
- Jun 9, 2026
- May 30, 2026
- MemGuardMemGuard: Preventing Memory Contamination in Long-Term Memory-Augmented Large Language ModelsMay 27, 2026
- DeferMemDeferMem: Query-Time Evidence Distillation via Reinforcement Learning for Long-Term Memory QAMay 21, 2026
- May 20, 2026
- May 3, 2026
- Apr 23, 2026
- Apr 2, 2026
- ChronosChronos: Temporal-Aware Conversational Agents with Structured Event Retrieval for Long-Term MemoryMar 17, 2026
- Mar 15, 2026
- Jan 13, 2026
- Agentic Memory (AgeMem)Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model AgentsJan 5, 2026