Back to Explore
cs.AIComputer Science

Artificial Intelligence

AI systems, knowledge representation, planning

35.5CVApr 22
Image Generators are Generalist Vision Learners

Valentin Gabeur, Shangbang Long, Songyou Peng et al.

This work suggests a potential paradigm shift in computer vision by positioning generative pretraining as a foundational approach for building generalist vision models that unify generation and understanding tasks.

51.1CVMar 28Code11k
SAM 3: Segment Anything with Concepts

Nicolas Carion, Laura Gustafson, Yuan-Ting Hu et al.

For researchers and practitioners in computer vision, SAM 3 provides a more accurate and unified model for concept-driven segmentation and tracking, with a new benchmark and dataset.

29.0AIMar 11Code
Mind the Sim2Real Gap in User Simulation for Agentic Tasks

Xuhui Zhou, Weiwei Sun, Qianou Ma et al. · cmu

This work addresses the critical issue of inaccurate user simulation in NLP agent evaluation, which can mislead development, and is incremental in providing empirical validation and a new metric.

33.4AIMar 16
CUBE: A Standard for Unifying Agent Benchmarks

Alexandre Lacoste, Nicolas Gontier, Oleh Shliazhko et al. · ibm-research

This addresses a critical productivity issue for AI researchers by standardizing benchmark integration to prevent further fragmentation as new benchmarks emerge.

28.5CVMar 17
Demystifing Video Reasoning

Ruisi Wang, Zhongang Cai, Fanyi Pu et al.

This provides a systematic understanding of reasoning emergence in video generation models, potentially guiding future research to exploit these dynamics for AI intelligence.

32.3SEMar 17Code
InCoder-32B: Code Foundation Model for Industrial Scenarios

Jian Yang, Wei Zhang, Jiajun Wu et al.

This addresses performance gaps in industrial code intelligence for domains like chip design and embedded systems, though it appears incremental as it builds on existing foundation model approaches.

28.8AIMar 10Code1.4k
Logics-Parsing-Omni Technical Report

Xin An, Jingyi Cai, Xiangyang Chen et al.

This work addresses multimodal parsing challenges for AI systems handling documents, images, and audio-visual data, representing a novel method for a known bottleneck.

25.4CVMar 20Code56
PEARL: Personalized Streaming Video Understanding Model

Yuanhong Zheng, Ruichuan An, Xiaopeng Lin et al.

This addresses the limitation of current personalization methods to static/offline data for future AI assistants, though it is incremental as it builds on existing vision-language models.