Rethinking Memory as Continuously Evolving ConnectivityJizhan Fang, Buqiang Xu, Zhixian Wang et al.
For LLM agents operating in dynamic environments, FluxMem addresses the brittleness of static memory by enabling adaptive connectivity evolution, leading to consistent SOTA results across diverse benchmarks.
22.5ROMar 19
Aegis: Automated Error Generation and Attribution for Multi-Agent SystemsFanqi Kong, Ruijie Zhang, Huaxiao Yin et al.
This work addresses the reliability and interpretability of multi-agent systems for developers and researchers, though it is incremental as it builds on existing error attribution methods with automated data generation.
KernelSkill: A Multi-Agent Framework for GPU Kernel OptimizationQitong Sun, Jun Han, Tianlin Li et al.
This addresses GPU kernel optimization for AI systems, offering a more interpretable and efficient approach compared to prior LLM-based methods.
18.8CVMar 12
VQQA: An Agentic Approach for Video Evaluation and Quality ImprovementYiwen Song, Tomas Pfister, Yale Song
This addresses the problem of inefficient or inaccessible video quality evaluation for users of video generation models, offering a novel black-box optimization method.
A Multi-Agent Perception-Action Alliance for Efficient Long Video ReasoningYichang Xu, Gaowen Liu, Ramana Rao Kompella et al. · gatech
This addresses the challenge of scaling video reasoning to real-world long videos for applications like video question-answering, though it appears incremental as an optimization of existing multi-agent and VLM approaches.
FedGUI: Benchmarking Federated GUI Agents across Heterogeneous Platforms, Devices, and Operating SystemsWenhao Wang, Haoting Shi, Mengying Yuan et al.
For researchers developing privacy-preserving GUI agents, this benchmark enables systematic study of cross-platform heterogeneity, filling a gap in existing federated learning benchmarks.
SkillGen: Verified Inference-Time Agent Skill SynthesisYuchen Ma, Yue Huang, Han Bao et al.
For LLM agent developers, SkillGen automates the creation of high-quality, reusable skills that improve agent performance without retraining, addressing the bottleneck of manual skill authoring.
19.1AIMar 20
A Subgoal-driven Framework for Improving Long-Horizon LLM AgentsTaiyi Wang, Sian Gooding, Florian Hartmann et al.
This addresses the challenge of autonomous control in dynamic digital environments for AI developers, offering a novel approach to enhance agent robustness, though it is incremental in combining planning and RL techniques.
25.7SEApr 9
Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness EngineeringChenyu Zhou, Huacan Chai, Wenteng Chen et al.
This provides a unified review framework for researchers and practitioners building LLM agents, though it is primarily conceptual rather than presenting new experimental results.
19.6CRMar 16Code
ClawWorm: Self-Propagating Attacks Across LLM Agent EcosystemsYihao Zhang, Zeming Wei, Xiaokun Luan et al.
This addresses critical security risks for users of interconnected multi-agent systems, exposing vulnerabilities that could lead to autonomous attacks without attacker intervention.
20.4ROMar 12
BrainMem: Brain-Inspired Evolving Memory for Embodied Agent Task PlanningXiaoyu Ma, Lianyu Hu, Wenbing Tang et al.
This addresses the challenge of persistent memory for embodied agents in complex 3D environments, offering a scalable solution for generalizable embodied intelligence.
29.1AIMay 25
ScientistOne: Towards Human-Level Autonomous Research via Chain-of-EvidenceRui Meng, Bhavana Dalvi Mishra, Jiefeng Chen et al.
For researchers and developers of autonomous AI systems, this work addresses the critical problem of output verifiability, providing a framework and system that eliminates common failure modes like hallucinated references and unreproducible results.
Chat2Workflow: A Benchmark for Generating Executable Visual Workflows with Natural LanguageYi Zhong, Buqiang Xu, Yijun Wang et al.
For developers and practitioners in industrial automation, this benchmark addresses the costly and error-prone manual construction of executable visual workflows by evaluating LLMs' ability to automate this process.
19.7SEApr 4Code
SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution ScenariosMinh V. T. Thai, Tue Le, Dung Nguyen Manh et al.
This addresses the need for better benchmarks in AI-assisted software engineering to reflect real-world development challenges, though it is incremental as it builds on existing benchmarking efforts.
TrinityGuard: A Unified Framework for Safeguarding Multi-Agent SystemsKai Wang, Biaojie Zeng, Zeming Wei et al.
This addresses safety risks for developers and users of multi-agent systems, though it appears incremental as it builds on existing standards like OWASP.
29.4MAJun 1
Multi-Agent Computer UseJing Yu Koh, Ruslan Salakhutdinov, Daniel Fried
For researchers and practitioners building computer use agents, this work addresses the limitations of single-agent systems for complex, long-horizon tasks by introducing a multi-agent coordination framework.
40.2AISep 3
Bioinfoysis Technical ReportQingyang Shao, Xin Zhang, Zhouyang Yuan et al.
This work addresses the challenge of maintaining data, computations, and evidence connections in long-horizon bioinformatics tasks for researchers and practitioners, offering a substantial improvement over existing LLM agent systems.
17.5CLMar 29
AgentSwing: Adaptive Parallel Context Management Routing for Long-Horizon Web AgentsZhaopeng Feng, Liangcai Su, Zhen Zhang et al.
Addresses the critical bottleneck of finite context capacity in LLM-based autonomous agents for long-horizon tasks, offering a principled framework for future context management strategies.
Synergy: A Next-Generation General-Purpose Agent for Open Agentic WebXiaohang Nie, Zihan Guo, Kezhuo Yang et al.
This addresses the problem of agent interoperability and social integration for developers and users in decentralized digital ecosystems, representing a novel architectural approach rather than an incremental improvement.
17.7SEMar 11
Resolving Java Code Repository Issues with iSWE AgentJatin Ganhotra, Sami Serhan, Antonio Abu Nassar et al.
This addresses the need for better automated issue resolution in enterprise software development, where Java is widely used, but it is incremental as it builds on existing agent-based methods with a focus on a specific language.