AIJun 29

ManimAgent: Self-Evolving Multimodal Agents for Visual Education

arXiv:2606.3029616.9Has Code
Predicted impact top 25% in AI · last 90 daysOriginality Highly original
AI Analysis

This work addresses the problem of wasted reflection experience in LLM-based agents for code generation, offering a memory mechanism that improves efficiency over no-memory and RAG baselines.

ManimAgent introduces a self-evolving multimodal agent that retains cross-task reflection experience via a dual-channel Episodic Memory Bank, achieving improved Pass@1 and reduced reflection rounds on a Manim code-generation task without weight updates or human seeds.

Multi-round reflection lets agents built on large language models recover from failures within a single task, but each task remains an isolated episode: lessons learned across many reflection rounds on one task are discarded before the next begins. We study this gap on a code-generation task: from a scientific paper section, the agent writes Python in the open-source Manim library to render a mathematical animation. We present ManimAgent, a self-evolving multimodal agent that carries reflection experience across tasks through a dual-channel Episodic Memory Bank grown entirely from its own task stream, with no weight updates and no human seeds. After each animation converges, a vision-language model scores the rendered keyframes; the resulting signals populate a positive channel M+ that stores success rationales as soft Reference Examples, and a negative channel M- that stores validated failure patterns as hard Known Pitfalls. On a fixed-probe evaluation against no-memory, matched-budget retrieval-augmented generation, and shuffled-memory baselines, blind human Pass@1 rises and reflection rounds fall as memory size grows. We will release the code, frozen memory snapshots, and the task stream.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes