Aichen Cai, Anmeng Zhang, Anyu Li et al.
This work addresses efficiency challenges for users of mid-scale LLMs, though it appears incremental with novel components like FiberPO and architectural optimizations.
AI systems, knowledge representation, planning
Aichen Cai, Anmeng Zhang, Anyu Li et al.
This work addresses efficiency challenges for users of mid-scale LLMs, though it appears incremental with novel components like FiberPO and architectural optimizations.
Valentin Gabeur, Shangbang Long, Songyou Peng et al.
This work suggests a potential paradigm shift in computer vision by positioning generative pretraining as a foundational approach for building generalist vision models that unify generation and understanding tasks.
Mateusz Dziemian, Maxwell Lin, Xiaohan Fu et al. · eth-zurich
This addresses a critical security threat for users of AI agents in high-stakes settings, revealing fundamental weaknesses in current models.
Pranjal Aggarwal, Marjan Ghazvininejad, Seungone Kim et al. · meta-ai
This work addresses the challenge of automated assessment for mathematical reasoning in STEM fields, offering incremental improvements through enhanced training methods.
Nicolas Carion, Laura Gustafson, Yuan-Ting Hu et al.
For researchers and practitioners in computer vision, SAM 3 provides a more accurate and unified model for concept-driven segmentation and tracking, with a new benchmark and dataset.
Xuhui Zhou, Weiwei Sun, Qianou Ma et al. · cmu
This work addresses the critical issue of inaccurate user simulation in NLP agent evaluation, which can mislead development, and is incremental in providing empirical validation and a new metric.
Alexandre Lacoste, Nicolas Gontier, Oleh Shliazhko et al. · ibm-research
This addresses a critical productivity issue for AI researchers by standardizing benchmark integration to prevent further fragmentation as new benchmarks emerge.
Ruisi Wang, Zhongang Cai, Fanyi Pu et al.
This provides a systematic understanding of reasoning emergence in video generation models, potentially guiding future research to exploit these dynamics for AI intelligence.
Xinhao Deng, Yixiang Zhang, Jiaqing Wu et al.
This addresses security risks for users and developers of autonomous LLM agents, but it is incremental as it builds on existing threat analysis frameworks.
Jian Yang, Wei Zhang, Jiajun Wu et al.
This addresses performance gaps in industrial code intelligence for domains like chip design and embedded systems, though it appears incremental as it builds on existing foundation model approaches.
Shubham Parashar, Shurui Gui, Xiner Li et al.
This addresses the challenge of inefficient reasoning improvement in small LLMs for mathematical and coding tasks, representing an incremental advancement in RL-based training methods.
Yuwen Du, Rui Ye, Shuo Tang et al.
This work democratizes frontier search agent research for the broader AI community by providing open-source data and models, addressing a bottleneck previously dominated by industrial giants.
Yichen Zhang, Da Peng, Zonghao Guo et al.
This addresses the challenge of mismatched decoding regimes and representations in multimodal AI, offering an efficient solution for unified tasks, though it appears incremental in building on existing UMM approaches.
Xin An, Jingyi Cai, Xiangyang Chen et al.
This work addresses multimodal parsing challenges for AI systems handling documents, images, and audio-visual data, representing a novel method for a known bottleneck.
Yuanhong Zheng, Ruichuan An, Xiaopeng Lin et al.
This addresses the limitation of current personalization methods to static/offline data for future AI assistants, though it is incremental as it builds on existing vision-language models.
Wenxuan Zhang, Lemeng Wu, Changsheng Zhao et al.
This work addresses the problem of efficient policy optimization for diffusion-based language models, offering incremental improvements in training and generation efficiency for AI researchers and practitioners.
Jinguang Tong, Jinbo Wu, Kaisiyuan Wang et al.
This work addresses a frontier in expressive digital human creation for applications like animation and virtual reality, representing a novel method for a known bottleneck rather than an incremental improvement.
Gangda Deng, Zhaoling Chen, Zhongming Yu et al.
This addresses the need for benchmarks that assess AI agents in dynamic, real-world software environments, which is incremental as it builds on existing evaluation methods.
Fang Wu, Haokai Zhao, Da Xing et al.
This work addresses the underexplored problem of noise valuation for diffusion model training, offering a method to improve training efficiency and generation quality.
Hao Zhang, Mingjie Liu, Shaokun Zhang et al.
This work addresses the problem of infrastructure inefficiency for researchers and developers training multi-turn LLM agents, though it is incremental as it focuses on improving existing rollout orchestration methods.