CLJun 25

Nemotron-TwoTower: Diffusion Language Modeling with Pretrained Autoregressive Context

arXiv:2606.2649322.9Has Code
Predicted impact top 27% in CL · last 90 daysOriginality Incremental advance
AI Analysis

For practitioners needing efficient text generation, this work offers a practical diffusion model that matches AR quality while being faster, though it is incremental over existing diffusion approaches.

Nemotron-TwoTower decouples context representation and iterative denoising in diffusion language models using two towers, achieving 98.7% of autoregressive baseline quality with 2.42X higher generation throughput.

Diffusion language models offer a promising alternative to autoregressive models due to their potential for parallel and iterative generation. However, existing approaches use a single network for both context representation and iterative denoising, forcing one model to serve both roles and limiting its capacity for either role. We propose TwoTower, a block-wise autoregressive diffusion model that decouples these roles into two towers: a frozen AR context tower that causally processes clean tokens, and a trainable diffusion denoiser tower with bidirectional block attention that refines noisy blocks via cross-attention to the context. Built on Nemotron-3-Nano-30B-A3B, an open-weight 30B hybrid Mamba-Transformer MoE model, and trained on approximately 2.1T tokens, Nemotron-TwoTower retains 98.7% of the autoregressive baseline's quality while offering 2.42X higher wall-clock generation throughput. We release the code and model weights at https://huggingface.co/collections/nvidia/nemotron-twotower.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes