AICLJul 21

AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents

arXiv:2607.1875412.3h-index: 2Has Code
Predicted impact top 17% in AI · last 90 daysOriginality Incremental advance
AI Analysis

For developers and researchers building LLM agents, AgentDebugX provides a structured debugging workflow with improved root-cause attribution and automated recovery, addressing a critical bottleneck in agent reliability.

AgentDebugX is an open-source toolkit for debugging LLM agents that organizes debugging as a closed loop of Detect, Attribute, Recover, and Rerun. Its core component, DeepDebug, achieves 28.8% exact agent-and-step attribution accuracy on the Who and When benchmark (vs. 21.7% for the strongest baseline) and repairs 13 of 73 failed tasks on GAIA in a single rerun, improving overall accuracy from 55.8% to 63.6%.

LLM agent failures are difficult to debug because the step where an error surfaces is often not the one that caused it. Existing observability tools replay execution traces but provide little support for identifying the root cause or translating diagnosis into recovery. We present AgentDebugX, an open-source debugging framework that organizes debugging as a closed loop of Detect, Attribute, Recover, and Rerun. At its core, DeepDebug performs multi-turn root-cause diagnosis through global trajectory understanding, structure-guided investigation, and cross-examination. On the Who and When benchmark, DeepDebug achieves the best strict attribution accuracy among the evaluated methods on both tested open-weight backbones, reaching 28.8 percent exact agent-and-step accuracy on qwen3.5-9b versus 21.7 percent for the strongest single-pass baseline. On GAIA, DeepDebug repairs 13 of 73 failed tasks in a single rerun, compared with 4 to 6 for three decoupled self-correction baselines, improving overall accuracy from 55.8 percent to 63.6 percent. AgentDebugX exposes this workflow through a Python library, CLI, web console, and installable agentic skill, and provides an opt-in Error Hub for sharing scrubbed failure-diagnosis-repair bundles and reusing them as debugging memory.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes