Zhiyuan Zeng, Yichi Zhang, Yong Shan et al.
This addresses the problem of LLMs' shallow reasoning in software engineering for developers and AI researchers, representing a new paradigm rather than incremental work.
Software development, testing, maintenance
Zhiyuan Zeng, Yichi Zhang, Yong Shan et al.
This addresses the problem of LLMs' shallow reasoning in software engineering for developers and AI researchers, representing a new paradigm rather than incremental work.
Jian Yang, Wei Zhang, Jiajun Wu et al.
This addresses performance gaps in industrial code intelligence for domains like chip design and embedded systems, though it appears incremental as it builds on existing foundation model approaches.
Gangda Deng, Zhaoling Chen, Zhongming Yu et al.
This addresses the need for benchmarks that assess AI agents in dynamic, real-world software environments, which is incremental as it builds on existing evaluation methods.
Cursor Research, Aaron Chan, Ahmed Shalaby et al. · berkeley, microsoft-research
This addresses the need for efficient coding models in software engineering, though it appears incremental as it builds on previous Composer models.
Aili Chen, Chi Zhang, Junteng Liu et al.
This addresses the challenge of robust generalization for tool-using LLMs in agentic tasks, offering a scalable solution to improve performance on out-of-distribution benchmarks, though it is incremental as it builds on existing synthesis methods.
Joongwon Kim, Wannan Yang, Kelvin Niu et al.
For developers of coding agents, this work addresses the bottleneck of scaling test-time compute for long-horizon tasks by focusing on representation and reuse of prior experience.
Jiale Zhao, Guoxin Chen, Fanzhe Meng et al.
This addresses the data bottleneck for training LLM-based software engineering agents, though it is incremental in automating data construction rather than a fundamental breakthrough.
Lintang Sutawika, Aditya Bharat Soni, Bharath Sriraam R R et al. · cmu
This addresses the need for efficient code search in software development, offering a simpler, agent-based approach that is incremental over prior methods using specialized tools.
Dayuan Fu, Shenyu Wu, Yunze Wu et al.
This provides a scalable, open-source solution for academic researchers to train software engineering agents, addressing a barrier in the field.
Xinping Lei, Xinyu Che, Junqi Xiong et al.
Provides a comprehensive evaluation framework for web coding capabilities of LLMs, addressing gaps in existing benchmarks.
John Yang, Kilian Lieret, Jeffrey Ma et al.
For AI software engineering, it reveals that current models fail at holistic codebase construction and produce monolithic implementations unlike human code.
Jialin Yang, Dongfu Jiang, Lipeng He et al. · amazon-science, utoronto
This work addresses the need for better evaluation of LLMs in software development workflows, where generating structured outputs is critical, though it is incremental as it builds on prior benchmarking efforts.
Kaixuan Wang, Tianxing Chen, Jiawei Liu et al.
This addresses the problem of limited simulation data for robotic manipulation researchers, though it is incremental as it builds on existing simulation-based learning paradigms.
Chenyu Zhou, Huacan Chai, Wenteng Chen et al.
This provides a unified review framework for researchers and practitioners building LLM agents, though it is primarily conceptual rather than presenting new experimental results.
Simon Yu, Derek Chong, Ananjan Nandi et al.
Provides an efficient infrastructure for programming meta-agents, enabling runtime intervention, counterfactual optimization, and tree-RL training.
Jian Yang, Wei Zhang, Shawn Guo et al.
This work addresses the need for more dynamic and efficient code generation models for developers and researchers, though it appears incremental with architectural enhancements.
Bihui Yu, Xinglong Xu, Junjie Jiang et al.
This work addresses the problem of automating visual typesetting optimization for scientific document preparation, a critical but overlooked stage in document automation.
Yihao Zhang, Zeming Wei, Xiaokun Luan et al.
This addresses critical security risks for users of interconnected multi-agent systems, exposing vulnerabilities that could lead to autonomous attacks without attacker intervention.
Qijun Han, Haoqin Tu, Zijun Wang et al.
For developers of autonomous GUI agents, this work provides a practical modular solution to common failure modes, though it is an incremental engineering contribution combining existing ideas.
Srijan Bansal, Jiao Fangkai, Yilun Zhou et al.
This addresses a critical gap for autonomous software engineering by systematically evaluating LLMs' debugging capabilities, revealing foundational limitations in fault reasoning that hinder agentic coding tools.