Yang Yang, Tianyi Zhang, Wei Huang et al.
This addresses the challenge of maintaining fidelity and coherence in streaming video diffusion for interactive applications, representing an incremental improvement over existing methods.
Image processing, video analysis
Yang Yang, Tianyi Zhang, Wei Huang et al.
This addresses the challenge of maintaining fidelity and coherence in streaming video diffusion for interactive applications, representing an incremental improvement over existing methods.
Firat Ozdemir, Yun Cheng, Salman Mohebi et al. · eth-zurich
Provides a unified open foundation model for diverse Earth system data, enabling broader adoption in climate science and downstream tasks.
Yebin Yang, Di Wen, Lei Qi et al.
This work addresses a domain-specific problem for researchers and practitioners in 3D animation and human-computer interaction, offering incremental advances in multi-person motion editing.
Mengfei Duan, Hao Shi, Fei Teng et al.
This addresses the need for comprehensive and safe perception in open-world exploration for embodied agents, representing a novel method rather than an incremental improvement.
Muyang He, Hanzhong Guo, Junxiong Lin et al.
For researchers and practitioners in video generation and world modeling, this survey provides a structured taxonomy and identifies efficiency as a key bottleneck, but it is a review paper with no new experimental results.
Xiangyan Qu, Zhenlong Yuan, Jing Tang et al. · tsinghua
This work addresses efficiency challenges in image editing for AI practitioners, offering an incremental improvement over existing test-time scaling approaches.
Tianyi Zeng, Jincheng Gao, Tianyi Wang et al.
This addresses a core problem in video editing for users of diffusion models, though it is incremental as it builds on existing DiT-based methods.
Moein Heidari, Ali Mehrabian, Mohammad Amin Roohi et al.
This work addresses the challenge of noisy and inconsistent echocardiography analysis for clinical applications, representing an incremental improvement over existing methods.
Lifeng Chen, Tianqi You, Hao Liu et al.
This work addresses the workload of radiologists by improving efficiency in medical report generation, though it is incremental as it builds on existing diffusion methods.
Guoqiang Zhao, Zhe Yang, Sheng Wu et al.
This addresses perception challenges for quadruped robots in complex environments, though it is incremental as it adapts existing occupancy prediction methods to a new robotic platform.
Hang Wang, Chao Shen, Chenhao Lin et al.
For digital forensics and media authenticity, this work addresses the growing challenge of AI-generated video detection by exploiting a previously overlooked cross-modal temporal fingerprint.
Runze Wang, Yuxuan Song, Youcheng Cai et al.
This work addresses the scalability problem for real-time 3D reconstruction in streaming settings, offering a plug-and-play solution to improve memory efficiency and performance.
Xinjie Zhang, Peng Zhang, Shicheng Zheng et al.
This work addresses the high cost of training and deploying large visual generators by co-designing a tokenizer, backbone, and system for efficient high-resolution image generation and editing.
Lubin Gan, Jing Zhang, Heng Zhang et al.
This work addresses critical issues in medical imaging for histopathology, offering a domain-specific improvement that is incremental but effective.
Jingyun Liang, Min Wei, Shikai Li et al.
For video generation tasks requiring precise 3D human motion control, this method offers a novel token-based pipeline that improves 3D awareness and reduces artifacts.
Yueqian Lin, Jingyang Zhang, Qinsi Wang et al.
This addresses the problem of temporal integration and cross-modal associations in computational systems for researchers in multimodal AI, though it appears incremental as it builds on known hippocampal mechanisms.
Qian Qi, Jiangyun Tang, Jim Lee et al.
This reveals a critical failure mode for watermarking systems used in copyright protection and content provenance, highlighting an incremental but important vulnerability in existing methods.
Yuanfan Zheng, Kunyu Peng, Xu Zheng et al.
This work solves the problem of comprehensive 360° scene understanding for real-world applications like autonomous driving or robotics, but it appears incremental as it builds on existing domain adaptation methods for panoramic segmentation.
Wojciech Zielonka, Tobias Kirschstein, Timo Bolkart et al.
It serves as a comprehensive reference for researchers and practitioners in computer graphics and vision, but is a survey, not a novel contribution.
Xiaoxu Peng, Dong Zhou, Jianwen Zhang et al.
It addresses security risks for autonomous driving systems, offering a scalable solution, though it appears incremental as it builds on existing detection and tuning methods.