Back to Explore
cs.CVComputer Science

Computer Vision

Image recognition, object detection, visual understanding

35.5CVApr 22
Image Generators are Generalist Vision Learners

Valentin Gabeur, Shangbang Long, Songyou Peng et al.

This work suggests a potential paradigm shift in computer vision by positioning generative pretraining as a foundational approach for building generalist vision models that unify generation and understanding tasks.

31.1CVMar 16Code3k
Kimodo: Scaling Controllable Human Motion Generation

Davis Rempe, Mathis Petrovich, Ye Yuan et al.

This addresses the need for scalable, high-quality human motion data for applications in robotics, simulation, and entertainment, representing a significant advancement over previous limited datasets.

32.7CVMar 13Code307
Multimodal OCR: Parse Anything from Documents

Handong Zheng, Yumeng Li, Kaile Zhang et al.

This addresses the limitation of conventional OCR systems that ignore graphical elements, enabling more comprehensive document parsing for applications like document reconstruction and multimodal pretraining.

51.1CVMar 28Code11k
SAM 3: Segment Anything with Concepts

Nicolas Carion, Laura Gustafson, Yuan-Ting Hu et al.

For researchers and practitioners in computer vision, SAM 3 provides a more accurate and unified model for concept-driven segmentation and tracking, with a new benchmark and dataset.

28.5CVMar 17
Demystifing Video Reasoning

Ruisi Wang, Zhongang Cai, Fanyi Pu et al.

This provides a systematic understanding of reasoning emergence in video generation models, potentially guiding future research to exploit these dynamics for AI intelligence.

27.1CVMar 16
Grounding World Simulation Models in a Real-World Metropolis

Junyoung Seo, Hyunwook Choi, Minkyung Kwon et al.

This work addresses the challenge of creating realistic, dynamic simulations of actual urban environments for applications in urban planning, autonomous systems, or virtual reality, representing a novel method for a known bottleneck rather than a foundational breakthrough.

25.4CVMar 20Code56
PEARL: Personalized Streaming Video Understanding Model

Yuanhong Zheng, Ruichuan An, Xiaopeng Lin et al.

This addresses the limitation of current personalization methods to static/offline data for future AI assistants, though it is incremental as it builds on existing vision-language models.

29.2CVApr 14
Lyra 2.0: Explorable Generative 3D Worlds

Tianchang Shen, Sherwin Bahmani, Kai He et al. · nvidia, utoronto

This work tackles the problem of generating large-scale, consistent 3D environments from video for applications in virtual reality and simulation, offering a significant improvement over existing methods that degrade over long trajectories.