Agent / long-term memory
HippoRAG 2
From RAG to Memory: Non-Parametric Continual Learning for Large Language Models
Superseded — cited as a baseline and beaten by newer methods
1 papers critique it · 5 beat it on benchmarks
What papers say
Verbatim critique sentences, each from a paper that cites HippoRAG 2 as a baseline.
However, these methods typically define cross-session relationships as simple clusters without modeling the nature of relationships or temporal evolution.
Beaten on benchmarks
Head-to-head results where a newer method reports beating HippoRAG 2. Values are copied from the source paper's tables — verify against the cited paper.
HingeMem beats HippoRAG 2
63.9 vs 39.1
PREMem beats HippoRAG 2
67.50 vs 45.95
LLM-as-a-judge score · [Qwen2.5-72B on LongMemEval]
Pre-Storage Reasoning for Episodic Memory: Shifting Inference Burden to Memory for Personalized DialogueREMem-I beats HippoRAG 2
93.1 vs 66.9
EM · [Test of Time]
REMem: Reasoning with Episodic Memory in Language AgentGSW beats HippoRAG 2
0.894 vs 0.787
Recall · [Overall Recall]
Beyond Fact Retrieval: Episodic Memory for RAG with Generative Semantic WorkspacesMemGAS beats HippoRAG 2
60.20 vs 57.60
GPT4o-as-Judge · [LongMemEval-s]
From Single to Multi-Granularity: Toward Long-Term Memory Association and Selection of Conversational Agents
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.
- HingeMemHingeMem: Boundary Guided Long-Term Memory with Query Adaptive Retrieval for Scalable DialoguesApr 8, 2026
- Jan 13, 2026
- Nov 25, 2025
- Generative Semantic Workspace (GSW)Beyond Fact Retrieval: Episodic Memory for RAG with Generative Semantic WorkspacesNov 10, 2025
- Oct 7, 2025
- PREMemPre-Storage Reasoning for Episodic Memory: Shifting Inference Burden to Memory for Personalized DialogueSep 13, 2025