SmartGR: Hierarchy and Beam-Aware Knowledge Distillation for Generative Recommendation
For practitioners deploying generative recommender systems, this work provides a method to compress large models into lightweight ones with better performance and efficiency, but it is an incremental improvement over existing distillation techniques.
The paper tackles the problem of inefficient knowledge distillation for generative recommendation models, addressing two specific challenges: imbalanced distillation difficulty across the semantic ID hierarchy and incorrect prefix pruning during beam search. The proposed SmartGR framework improves recommendation performance by 8.6% while achieving a 2.39x inference speedup on average across four benchmark datasets.
Generative recommendation (GR) has emerged as a promising paradigm for recommender systems. Scaling up GR models can improve recommendation performance, but it also substantially increases inference cost. Knowledge distillation provides a practical solution by transferring knowledge from a large GR model to a lightweight one. However, existing distillation methods do not account for two GR-specific challenges: imbalanced distillation difficulty across the semantic ID (SID) hierarchy and incorrect prefix pruning during beam search. To address these challenges, we propose SmartGR, a novel distillation framework that utilizes Hierarchy-Aware SID Distillation to transfer the teacher's modeling capability across the hierarchy and leverages Beam-Aware Ranking Distillation to distill the teacher's ranking preferences during beam search. Extensive experiments on four benchmark datasets demonstrate the effectiveness and efficiency of SmartGR, improving the performance by 8.6% while achieving a 2.39$\times$ inference speedup on average.