Reconstruction Meets Prediction: Self-Supervised Learning for Multi-Instrument Music Transcription
Online
19:00 BST – UK LOCAL TIME (18:00 UTC)
Event Duration: 1.5 hours
Automatic music transcription, which converts an audio recording into a symbolic representation such as a score or MIDI, remains a hard problem for multi-instrument music, where overlapping harmonics and diverse timbres must be disentangled. Progress is held back by a data bottleneck: note-level annotation is expensive and requires expert musicians, leaving labelled datasets for transcription very scarce.
This talk explores self-supervised learning (SSL) as a way to sidestep that bottleneck by learning musical representations from unlabelled audio. I will focus on a controlled comparison of two masked-modelling paradigms: reconstruction-based learning with a masked autoencoder (MAE), which reconstructs masked spectrogram regions, and predictive learning with a joint embedding predictive architecture (JEPA), which predicts masked regions directly in representation space. Using identical Transformer encoders, I will show that the two objectives are broadly comparable on transcription performance but offer complementary strengths depending on instrument and texture. Initialising JEPA from an MAE-pretrained encoder produces consistent further gains. I will also discuss what the learned latent spaces reveal about how each objective organises musical information, and why no single geometric metric reliably predicts downstream transcription performance.
Register
Registration is currently closed for this event.
