13.0CLAug 11, 2025
Efficient Speculative Decoding for Llama at Scale: Challenges and SolutionsBangsheng Tang, Carl Chengyan Fu, Fei Kou et al.
Speculative decoding is a standard method for accelerating the inference speed of large language models. However, scaling it for production environments poses several engineering challenges, including efficiently implementing different operations (e.g., tree attention and multi-round speculative decoding) on GPU. In this paper, we detail the training and inference optimization techniques that we have implemented to enable EAGLE-based speculative decoding at a production scale for Llama models. With these changes, we achieve a new state-of-the-art inference latency for Llama models. For example, Llama4 Maverick decodes at a speed of about 4 ms per token (with a batch size of one) on 8 NVIDIA H100 GPUs, which is 10% faster than the previously best known method. Furthermore, for EAGLE-based speculative decoding, our optimizations enable us to achieve a speed-up for large batch sizes between 1.4x and 2.0x at production scale.
2.3AIMar 3, 2020
Hierarchical Context Enhanced Multi-Domain Dialogue System for Multi-domain Task CompletionJingyuan Yang, Guang Liu, Yuzhao Mao et al.
Task 1 of the DSTC8-track1 challenge aims to develop an end-to-end multi-domain dialogue system to accomplish complex users' goals under tourist information desk settings. This paper describes our submitted solution, Hierarchical Context Enhanced Dialogue System (HCEDS), for this task. The main motivation of our system is to comprehensively explore the potential of hierarchical context for sufficiently understanding complex dialogues. More specifically, we apply BERT to capture token-level information and employ the attention mechanism to capture sentence-level information. The results listed in the leaderboard show that our system achieves first place in automatic evaluation and the second place in human evaluation.
1.2NIOct 10, 2017
Link Quality Aware Channel Allocation for Multichannel Body Sensor NetworksWeifeng Gao, Zhiwei Zhao, Geyong Min et al.
Body Sensor Network (BSN) is a typical Internet-of-Things (IoT) application for personalized health care. It consists of economically powered, wireless and implanted medical monitoring sensor nodes, which are designed to continually collect the medical information of the target patients. Multichannel is often used in BSNs to reduce the spectrum competition of the tremendous sensor nodes and the problem of channel assignment has attracted much research attention. The health sensing data in BSNs is often required to be delivered to a sink node (or server) before a certain deadline for real time monitoring or health emergency alarm. Therefore, deadline is of significant importance for multichannel allocation and scheduling. The existing works, though designed to meet the deadline, often overlook the impact of the unreliable wireless links. As a result, the health sensing data can still be overdue because of the scheduled lossy links. Besides, potential collisions in the schedules also incur considerable delay in delivering the sensing data. In this paper, we propose a novel deadline- driven Link quality Aware Channel Assignment scheme (LACA), where link quality, deadlines and collisions are jointly considered. LACA prioritizes links with urgent deadlines and heavy collisions. Besides, LACA allows the exploition of the spare slots for retransmissions on lossy links, which can further reduce the retransmission delay. Extensive simulation experiments show that compared to the existing approaches, LACA can better utilize the wireless spectrum and achieve higher packet delivery ratio before the deadline.