ITJun 10

STCC: A Unified Source-Channel Semantic Token Coding Framework for Semantic Communications

arXiv:2606.11819v19.5h-index: 13
Predicted impact top 19% in IT · last 90 daysOriginality Highly original
AI Analysis

For wireless communication systems using foundation models, STCC addresses the cliff effect and random errors by enabling robust transmission of discrete semantic tokens without receiver-side modification.

STCC proposes a unified source-channel semantic token coding framework that transmits discrete semantic tokens over noisy channels using a learned constellation aligned with the semantic embedding space, significantly outperforming traditional systems in low-SNR regimes by converting channel noise into semantic variations.

Deep Joint Source-Channel Coding (JSCC) has emerged as a promising paradigm for overcoming the ``cliff effect" in wireless communications. However, existing Deep JSCC frameworks operate directly on raw analog data such as image pixels rather than the discrete semantic tokens that foundation models require. Moreover, traditional systems employ fixed, hand-designed constellations that treat all tokens equally, leading to catastrophic random errors under channel noise. In this paper, the Semantic Token Codebook Communication (STCC) is proposed as a unified source-channel semantic token coding framework designed to transmit the discrete semantic tokens of foundation models over noisy channels. The core of STCC is the Semantic Token Codec (STC). It accepts discrete tokens as input, which maintains compatibility with foundation models while employing a residual multiple layer perceptron, i.e., MLP-based encoder that learns geometrically structured constellations optimized with a triple-loss objective. This learned mapping forces the channel topology to align with the semantic embedding space, ensuring that channel noise results in topological errors rather than random corruption. This phenomenon is theoretically and empirically characterized, identifying ``Semantic Drift" in symbolic modalities and ``Structural Distortion" in perceptual modalities, where errors shift predictions to semantically or structurally similar tokens. Extensive experiments demonstrate that STCC significantly outperforms traditional systems in low-SNR regimes, effectively converting channel noise into semantic variations without requiring receiver-side modification.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes