Not All Skills Help: Measuring and Repairing Agent Knowledge
For developers of LLM agents, this work addresses the bottleneck of skill curation by providing a method to match skills to tasks at inference time, significantly improving performance without modifying model weights.
LLM agents accumulate skills from experience, but current methods rely solely on LLM judgment for curation, causing causal heterogeneity where skills help on some tasks and hurt on others. The proposed ASSAY framework separates generation from curation via per-skill causal attribution, achieving 69.3% task-goal completion on AppWorld's hardest split (47.4% relative improvement) and 8.7% relative improvement on tau-bench retail, setting new state-of-the-art results without weight updates.
LLM agents can improve without weight updates by accumulating natural-language skills from experience, but current systems entrust every decision about which skills to keep and how to apply them to LLM judgment alone. We argue that this conflates two distinct roles: generating a skill from experience is a creative act that judgment handles well, while deciding whether that skill actually helps requires empirical evidence across many tasks. Measuring per-skill causal contributions via randomized masking, we find that skill libraries exhibit pervasive causal heterogeneity: individual skills routinely help on some task types while hurting on others, yet their opposing effects cancel in aggregate, making them invisible to global curation methods. We propose ASSAY, a framework that separates generation from curation: it computes a per-skill causal attribution on a small development set, restructures the library offline, and suppresses skills with negative predicted effect for each test task. Across seven base models spanning four providers and two benchmarks (AppWorld and tau-bench), ASSAY consistently improves over prior skill-curation approaches. On AppWorld's hardest split, DeepSeek-V3 achieves 69.3% task-goal completion (47.4% relative improvement), a new state of the art among all published methods including weight-tuned approaches. On tau-bench retail, GPT-4.1 improves by 8.7% relative, advancing past o4-mini, o1, and GPT-4.5 on the public leaderboard without any weight modification. Ablation traces the dominant gain to per-task masking, confirming that the bottleneck is matching skills to tasks at inference time, not removing bad skills globally. Code is available at https://github.com/aiming-lab/assay.