Perceptual compensation for tonal context in self-supervised speech models
For researchers in speech processing and phonology, this work highlights the limitations of self-supervised learning in capturing phonological regularities without supervised fine-tuning.
The study tested whether wav2vec2.0 models show perceptual compensation for tonal context in Mandarin Chinese, finding no evidence in the pre-trained model and only partial evidence in the fine-tuned model, failing to replicate human performance.
This study examines the extent to which the wav2vec2.0 architecture exhibits evidence of compensation for phonological context. We conducted a pseudo-replication of a perceptional compensation experiment on Mandarin Chinese tones, and compared the embedding similarities and probing classifier outputs between a purely self-supervised pre-trained model and a model fine-tuned for Mandarin ASR. No evidence of compensation was found in the embedding similarities of the purely pre-trained model. Probing classifiers showed some evidence of compensation in addition to the expected layer-wise improvements in categorization, but failed to replicate human performance on isolated test syllables. Our findings contrast with previous reports of sensitivity to phonological structure emerging through pre-training alone, and suggest that supervised objectives may be necessary to encourage the abstraction of at least some types of phonological regularities.