Let All be Whitened: Multi-teacher Distillation for Efficient Visual RetrievalZhe Ma, Jianfeng Dong, Shouling Ji et al.
Visual retrieval aims to search for the most relevant visual items, e.g., images and videos, from a candidate gallery with a given query item. Accuracy and efficiency are two competing objectives in retrieval tasks. Instead of crafting a new method pursuing further improvement on accuracy, in this paper we propose a multi-teacher distillation framework Whiten-MTD, which is able to transfer knowledge from off-the-shelf pre-trained retrieval models to a lightweight student model for efficient visual retrieval. Furthermore, we discover that the similarities obtained by different retrieval models are diversified and incommensurable, which makes it challenging to jointly distill knowledge from multiple models. Therefore, we propose to whiten the output of teacher models before fusion, which enables effective multi-teacher distillation for retrieval models. Whiten-MTD is conceptually simple and practically effective. Extensive experiments on two landmark image retrieval datasets and one video retrieval dataset demonstrate the effectiveness of our proposed method, and its good balance of retrieval performance and efficiency. Our source code is released at https://github.com/Maryeon/whiten_mtd.
6.2CVFeb 17, 2025
VoLUT: Efficient Volumetric streaming enhanced by LUT-based super-resolutionChendong Wang, Anlan Zhang, Yifan Yang et al.
3D volumetric video provides immersive experience and is gaining traction in digital media. Despite its rising popularity, the streaming of volumetric video content poses significant challenges due to the high data bandwidth requirement. A natural approach to mitigate the bandwidth issue is to reduce the volumetric video's data rate by downsampling the content prior to transmission. The video can then be upsampled at the receiver's end using a super-resolution (SR) algorithm to reconstruct the high-resolution details. While super-resolution techniques have been extensively explored and advanced for 2D video content, there is limited work on SR algorithms tailored for volumetric videos. To address this gap and the growing need for efficient volumetric video streaming, we have developed VoLUT with a new SR algorithm specifically designed for volumetric content. Our algorithm uniquely harnesses the power of lookup tables (LUTs) to facilitate the efficient and accurate upscaling of low-resolution volumetric data. The use of LUTs enables our algorithm to quickly reference precomputed high-resolution values, thereby significantly reducing the computational complexity and time required for upscaling. We further apply adaptive video bit rate algorithm (ABR) to dynamically determine the downsampling rate according to the network condition and stream the selected video rate to the receiver. Compared to related work, VoLUT is the first to enable high-quality 3D SR on commodity mobile devices at line-rate. Our evaluation shows VoLUT can reduce bandwidth usage by 70% , boost QoE by 36.7% for volumetric video streaming and achieve 3D SR speed-up with no quality compromise.
5.1MMNov 30, 2018
A Robust Algorithm for Tile-based 360-degree Video Streaming with Uncertain FoV EstimationArnob Ghosh, Vaneet Aggarwal, Feng Qian
We propose a robust scheme for streaming 360-degree immersive videos to maximize the quality of experience (QoE). Our streaming approach introduces a holistic analytical framework built upon the formal method of stochastic optimization. We propose a robust algorithm which provides a streaming rate such that the video quality degrades below that rate with very low probability even in presence of uncertain head movement, and bandwidth. It assumes the knowledge of the viewing probability of different portions (tiles) of a panoramic scene. Such probabilities can be easily derived from crowdsourced measurements performed by 360 video content providers. We then propose efficient methods to solve the problem at runtime while achieving a bounded optimality gap (in terms of the QoE). We implemented our proposed approaches using emulation. Using real users' head movement traces and real cellular bandwidth traces, we show that our algorithms significantly outperform the baseline algorithms by at least in $30\%$ in the QoE metric. Our algorithm gives a streaming rate which is $50\%$ higher compared to the baseline algorithms when the prediction error is high.
9.7NIApr 30, 2018
LBP: Robust Rate Adaptation Algorithm for SVC Video StreamingAnis Elgabli, Vaneet Aggarwal, Shuai Hao et al.
Video streaming today accounts for up to 55\% of mobile traffic. In this paper, we explore streaming videos encoded using Scalable Video Coding scheme (SVC) over highly variable bandwidth conditions such as cellular networks. SVC's unique encoding scheme allows the quality of a video chunk to change incrementally, making it more flexible and adaptive to challenging network conditions compared to other encoding schemes. Our contribution is threefold. First, we formulate the quality decisions of video chunks constrained by the available bandwidth, the playback buffer, and the chunk deadlines as an optimization problem. The objective is to optimize a novel QoE metric that models a combination of the three objectives of minimizing the stall/skip duration of the video, maximizing the playback quality of every chunk, and minimizing the number of quality switches. Second, we develop Layered Bin Packing (LBP) Adaptation Algorithm, a novel algorithm that solves the proposed optimization problem. Moreover, we show that LBP achieves the optimal solution of the proposed optimization problem with linear complexity in the number of video chunks. Third, we propose an online algorithm (online LBP) where several challenges are addressed including handling bandwidth prediction errors, and short prediction duration. Extensive simulations with real bandwidth traces of public datasets reveal the robustness of our scheme and demonstrate its significant performance improvement as compared to the state-of-the-art SVC streaming algorithms. The proposed algorithm is also implemented on a TCP/IP emulation test bed with real LTE bandwidth traces, and the emulation confirms the simulation results and validates that the algorithm can be implemented and deployed on today's mobile devices.
16.7CYDec 1, 2017
DeepWear: Adaptive Local Offloading for On-Wearable Deep LearningMengwei Xu, Feng Qian, Mengze Zhu et al.
Due to their on-body and ubiquitous nature, wearables can generate a wide range of unique sensor data creating countless opportunities for deep learning tasks. We propose DeepWear, a deep learning (DL) framework for wearable devices to improve the performance and reduce the energy footprint. DeepWear strategically offloads DL tasks from a wearable device to its paired handheld device through local network. Compared to the remote-cloud-based offloading, DeepWear requires no Internet connectivity, consumes less energy, and is robust to privacy breach. DeepWear provides various novel techniques such as context-aware offloading, strategic model partition, and pipelining support to efficiently utilize the processing capacity from nearby paired handhelds. Deployed as a user-space library, DeepWear offers developer-friendly APIs that are as simple as those in traditional DL libraries such as TensorFlow. We have implemented DeepWear on the Android OS and evaluated it on COTS smartphones and smartwatches with real DL models. DeepWear brings up to 5.08X and 23.0X execution speedup, as well as 53.5% and 85.5% energy saving compared to wearable-only and handheld-only strategies, respectively.
5.9MMApr 26, 2017
A Rate Adaptation Algorithm for Tile-based 360-degree Video StreamingArnob Ghosh, Vaneet Aggarwal, Feng Qian
In the 360-degree immersive video, a user only views a part of the entire raw video frame based on her viewing direction. However, today's 360-degree video players always fetch the entire panoramic view regardless of users' head movement, leading to significant bandwidth waste that can be potentially avoided. In this paper, we propose a novel adaptive streaming scheme for 360-degree videos. The basic idea is to fetch the invisible portion of a video at the lowest quality based on users' head movement prediction and to adaptively decide the video playback quality for the visible portion based on bandwidth prediction. Doing both in a robust manner requires overcome a series of challenges, such as jointly considering the spatial and temporal domains, tolerating prediction errors, and achieving low complexity. To overcome these challenges, we first define quality of experience (QoE) metrics for adaptive 360-degree video streaming. We then formulate an optimization problem and solve it at a low complexity. The algorithm strategically leverages both future bandwidth and the distribution of users' head positions to determine the quality level of each tile (i.e., a sub-area of a raw frame). We further provide theoretical proof showing that our algorithm achieves optimality under practical assumptions. Numerical results show that our proposed algorithms significantly boost the user QoE by at least 20\% compared to baseline algorithms.
3.1CVApr 9, 2013
Image Classification by Feature Dimension Reduction and Graph based RankingYao Nan, Qian Feng, Sun Zuolei
Dimensionality reduction (DR) of image features plays an important role in image retrieval and classification tasks. Recently, two types of methods have been proposed to improve the both the accuracy and efficiency for the dimensionality reduction problem. One uses Non-negative matrix factorization (NMF) to describe the image distribution on the space of base matrix. Another one for dimension reduction trains a subspace projection matrix to project original data space into some low-dimensional subspaces which have deep architecture, so that the low-dimensional codes would be learned. At the same time, the graph based similarity learning algorithm which tries to exploit contextual information for improving the effectiveness of image rankings is also proposed for image class and retrieval problem. In this paper, after above two methods mentioned are utilized to reduce the high-dimensional features of images respectively, we learn the graph based similarity for the image classification problem. This paper compares the proposed approach with other approaches on an image database.