4.6LGAug 25, 2024
Variational autoencoder-based neural network model compressionLiang Cheng, Peiyuan Guan, Amir Taherkordi et al.
Variational Autoencoders (VAEs), as a form of deep generative model, have been widely used in recent years, and shown great great peformance in a number of different domains, including image generation and anomaly detection, etc.. This paper aims to explore neural network model compression method based on VAE. The experiment uses different neural network models for MNIST recognition as compression targets, including Feedforward Neural Network (FNN), Convolutional Neural Network (CNN), Recurrent Neural Network (RNN) and Long Short-Term Memory (LSTM). These models are the most basic models in deep learning, and other more complex and advanced models are based on them or inherit their features and evolve. In the experiment, the first step is to train the models mentioned above, each trained model will have different accuracy and number of total parameters. And then the variants of parameters for each model are processed as training data in VAEs separately, and the trained VAEs are tested by the true model parameters. The experimental results show that using the latent space as a representation of the model compression can improve the compression rate compared to some traditional methods such as pruning and quantization, meanwhile the accuracy is not greatly affected using the model parameters reconstructed based on the latent space. In the future, a variety of different large-scale deep learning models will be used more widely, so exploring different ways to save time and space on saving or transferring models will become necessary, and the use of VAE in this paper can provide a basis for these further explorations.
6.0LGJun 2
LLM Compression with Jointly Optimizing Architectural and Quantization choicesHoang-Loc La, Truong-Thanh Le, Amir Taherkordi et al.
Deploying large language models (LLMs) is challenging due to their significant memory and computational requirements. While some methods address this by developing small or tiny language models from scratch, these approaches demand extensive GPU training. Compressing pre-trained LLMs for edge devices offers a compelling alternative. Beyond pruning and quantization, Neural Architecture Search (NAS) enables effective compression, yet prior NAS approaches often limit the search space and decouple architecture from quantization. We introduce a differentiable NAS framework that explores the entire space and jointly optimizes architectural configurations alongside mixed-precision quantization for linear layers of LLMs. Experiments demonstrate superior accuracy-latency trade-offs: our models achieve up to 1.4x faster inference than sequential NAS-then-quantization baselines at comparable accuracy, or up to 6% higher average accuracy across seven reasoning tasks at equivalent latency.
5.3RONov 5, 2021
Digital Twin-Assisted Controlling of AGVs in Flexible Manufacturing EnvironmentsMohammad Azangoo, Amir Taherkordi, Jan Olaf Blech et al.
Digital Twins are increasingly being introduced for smart manufacturing systems to improve the efficiency of the main disciplines of such systems. Formal techniques, such as graphs, are a common way of describing Digital Twin models, allowing broad types of tools to provide Digital Twin based services such as fault detection in production lines. Obtaining correct and complete formal Digital Twins of physical systems can be a complicated and time consuming process, particularly for manufacturing systems with plenty of physical objects and the associated manufacturing processes. Automatic generation of Digital Twins is an emerging research field and can reduce time and costs. In this paper, we focus on the generation of Digital Twins for flexible manufacturing systems with Automated Guided Vehicles (AGVs) on the factory floor. In particular, we propose an architectural framework and the associated design choices and software development tools that facilitate automatic generation of Digital Twins for AGVs. Specifically, the scope of the generated digital twins is controlling AGVs in the factory floor. To this end, we focus on different control levels of AGVs and utilize graph theory to generate the graph-based Digital Twin of the factory floor.
5.9NIJun 13, 2021
Active Learning for Network Traffic Classification: A Technical StudyAmin Shahraki, Mahmoud Abbasi, Amir Taherkordi et al.
Network Traffic Classification (NTC) has become an important feature in various network management operations, e.g., Quality of Service (QoS) provisioning and security services. Machine Learning (ML) algorithms as a popular approach for NTC can promise reasonable accuracy in classification and deal with encrypted traffic. However, ML-based NTC techniques suffer from the shortage of labeled traffic data which is the case in many real-world applications. This study investigates the applicability of an active form of ML, called Active Learning (AL), in NTC. AL reduces the need for a large number of labeled examples by actively choosing the instances that should be labeled. The study first provides an overview of NTC and its fundamental challenges along with surveying the literature on ML-based NTC methods. Then, it introduces the concepts of AL, discusses it in the context of NTC, and review the literature in this field. Further, challenges and open issues in AL-based classification of network traffic are discussed. Moreover, as a technical survey, some experiments are conducted to show the broad applicability of AL in NTC. The simulation results show that AL can achieve high accuracy with a small amount of data.