LGAIJun 19

An Empirical Study of OpenPangu Quantization on Ascend NPUs

arXiv:2606.212575.7
Predicted impact top 72% in LG · last 90 daysOriginality Synthesis-oriented
AI Analysis

Provides an accuracy map for deploying OpenPangu models on Ascend NPUs, but the findings are incremental and confirm known quantization trends.

This paper empirically evaluates post-training quantization methods for OpenPangu models on Ascend NPUs, finding that 8-bit weight-only quantization is lossless, 4-bit is practical for 7B but harmful for 1B, and ultra-low precision (2-bit, binary) collapses.

OpenPangu models are attractive targets for private and domestic large-language-model deployment, yet their robustness under aggressive post-training quantization on Ascend NPUs has not been systematically characterized. This paper conducts a controlled empirical study of OpenPangu 1B and 7B models on Huawei Ascend 910B1 NPUs. We evaluate representative weight-only and weight-activation post-training quantization methods, including RTN, GPTQ, AWQ, SmoothQuant, GPTAQ, BiLLM, and SliM-LLM, under a unified calibration and evaluation protocol. Across 18 evaluation tasks, we find that 8-bit weight-only quantization is effectively lossless for both models, while 4-bit quantization remains practical for the 7B model but is visibly more harmful for the 1B model on reasoning, math, and code tasks. Ultra-low precision remains challenging: most 2-bit and binary settings collapse to near-random behavior, and W4A4 SmoothQuant produces non-finite perplexity in our evaluation. These results provide an NPU-oriented accuracy map for selecting OpenPangu quantization settings and highlight the persistent difficulty of extreme low-bit compression.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes