Back to Explore
cs.MMComputer Science

Multimedia

Multimedia systems, content analysis

46.9CVJun 1Code11k
Cosmos 3: Omnimodal World Models for Physical AI

Aditi, Niket Agarwal, Arslan Ali et al.

This work provides a scalable, general-purpose backbone for embodied agents by unifying multiple modalities into a single framework, which is a significant step for Physical AI research.

19.0CLMay 27Code1k
Rethinking Memory as Continuously Evolving Connectivity

Jizhan Fang, Buqiang Xu, Zhixian Wang et al.

For LLM agents operating in dynamic environments, FluxMem addresses the brittleness of static memory by enabling adaptive connectivity evolution, leading to consistent SOTA results across diverse benchmarks.

33.3CVApr 9Code97
A Survey on 3D Gaussian Splatting

Guikun Chen, Wenguan Wang

It addresses the need for a comprehensive overview of this emerging method for researchers in computer graphics and vision, but it is incremental as it surveys existing developments rather than introducing new findings.

35.3CVApr 22Code
Building a Precise Video Language with Human-AI Oversight

Zhiqiu Lin, Chancharik Mitra, Siyuan Cen et al.

For researchers and practitioners in video understanding and generation, this work provides a scalable method to produce high-quality, structured captions that improve both VLM performance and text-to-video generation control.

20.4SDJun 3
Audio Interaction Model

Zhifei Xie, Zihang Liu, Ze An et al.

This work addresses the need for a single model that can handle multiple streaming audio tasks (e.g., voice chatting, ASR) in real time, unifying capabilities that were previously separate.

25.1CVApr 9
LPM 1.0: Video-based Character Performance Model

Ailing Zeng, Casper Yang, Chauncey Ge et al.

This addresses the problem of creating lifelike virtual characters for conversational agents, live streaming, and games, representing a novel method rather than an incremental improvement.

14.6CVMar 20
EgoForge: Goal-Directed Egocentric World Simulator

Yifan Shen, Jiateng Liu, Xinzhuo Li et al.

This work addresses the problem of simulating dynamic egocentric environments for applications like smart-glasses, though it is incremental as it builds on existing generative world models.

25.8AIApr 27
Co-Director: Agentic Generative Video Storytelling

Yale Song, Yiwen Song, Nick Losier et al.

For AI video generation, Co-Director addresses semantic drift in agentic pipelines, offering a principled optimization approach that generalizes to cinematic narratives.