Velocity Prediction in Automatic Guitar Transcription
This work addresses the lack of velocity prediction in guitar transcription for researchers and practitioners, but the improvement in note transcription is limited and incremental.
The authors present a method for velocity prediction in automatic guitar transcription by pretraining on synthetic data and transferring weights to a model trained on real audio, achieving velocity prediction that outperforms a baseline without pretraining and note transcription comparable to state-of-the-art.
Automatic Music Transcription (AMT) models have achieved a high level of success in polyphonic transcription of various instruments. Velocity, typically a measure of note intensity, is less commonly predicted in these models due to the absence of velocity labels in available datasets and lack of a proper definition for instruments other than piano. We present a methodology and model for velocity prediction in Automatic Guitar Transcription (AGT) which uses virtual instruments to generate synthetic training data with velocity labels. We first pretrain a model on this synthetic data. These weights are then transferred to a different model and trained on real guitar audio, allowing the model to retain the working velocity prediction while also achieving high performance and generalisability from the real training data. The velocity prediction is shown to outperform a baseline model which does not use the pretrained velocity weights, when evaluated on synthetic data. In addition, using the pretrained velocity weights offers a small improvement in note transcription, though the magnitude of this improvement is limited and not always significant depending on the testing data. Overall the model achieves results comparable to the state of the art in guitar transcription, while also successfully predicting velocity.