DC AINov 1, 2024

On the Impact of White-box Deployment Strategies for Edge AI on Latency and Model Performance

Jaskirat Singh, Bram Adams, Ahmed E. Hassan

arXiv:2411.00907v32.31 citationsh-index: 8

Originality Incremental advance

AI Analysis

It provides practical guidance for MLOps engineers on deployment strategies in Edge AI, though it is incremental as it compares existing methods.

This study empirically assesses the accuracy vs latency trade-off of white-box and black-box operators and their combinations in Edge AI setups, finding that combining Distillation and SPTQ operators (DSPTQ) is preferred for lower latency with small to medium accuracy drops, and distilled operators perform better in mobile and edge tiers.

To help MLOps engineers decide which operator to use in which deployment scenario, this study aims to empirically assess the accuracy vs latency trade-off of white-box (training-based) and black-box operators (non-training-based) and their combinations in an Edge AI setup. We perform inference experiments including 3 white-box (i.e., QAT, Pruning, Knowledge Distillation), 2 black-box (i.e., Partition, SPTQ), and their combined operators (i.e., Distilled SPTQ, SPTQ Partition) across 3 tiers (i.e., Mobile, Edge, Cloud) on 4 commonly-used Computer Vision and Natural Language Processing models to identify the effective strategies, considering the perspective of MLOps Engineers. Our Results indicate that the combination of Distillation and SPTQ operators (i.e., DSPTQ) should be preferred over non-hybrid operators when lower latency is required in the edge at small to medium accuracy drop. Among the non-hybrid operators, the Distilled operator is a better alternative in both mobile and edge tiers for lower latency performance at the cost of small to medium accuracy loss. Moreover, the operators involving distillation show lower latency in resource-constrained tiers (Mobile, Edge) compared to the operators involving Partitioning across Mobile and Edge tiers. For textual subject models, which have low input data size requirements, the Cloud tier is a better alternative for the deployment of operators than the Mobile, Edge, or Mobile-Edge tier (the latter being used for operators involving partitioning). In contrast, for image-based subject models, which have high input data size requirements, the Edge tier is a better alternative for operators than Mobile, Edge, or their combination.

View on arXiv PDF

Similar