Agent / long-term memory
Reflexion
Reflexion: Language Agents with Verbal Reinforcement Learning
Superseded baseline#11 of 63 most-superseded · first seen Mar 20, 2023
Superseded — cited as a baseline and beaten by newer methods
3 papers critique it · 2 beat it on benchmarks
What papers say
Verbatim critique sentences, each from a paper that cites Reflexion as a baseline.
Matrix incorporates a unique iterative self-refinement mechanism that allows agents to systematically improve their understanding of document structures and extraction patterns.
“A Reflexion agent that accumulates thousands of verbal self-critiques is still running the same frozen model at every session; its filing cabinet grows while its capacity does not.”
“Reflexion approximates reinforcement learning by storing self-critiques from synthetic environments, but it relies on binary correctness signals from the environment rather than pre-labeled data and restricts memory retrieval to identical tasks, limiting its ability to generalize.”
Beaten on benchmarks
Head-to-head results where a newer method reports beating Reflexion. Values are copied from the source paper's tables — verify against the cited paper.
EMPO^2 beats Reflexion
75.9 vs 17.1
Average · [ScienceWorld]
Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy OptimizationAriGraph beats Reflexion
0.79 vs 0.27
normalized score · [Cleaning]
AriGraph: Learning Knowledge Graph World Models with Episodic Memory for LLM Agents
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.
- Feb 26, 2026