September 12, 2026
Online 19:00 BST – UK LOCAL TIME (18:00 UTC) Event Duration: 1.5 hours Automatic music transcription, which converts an audio recording into a symbolic representation such as a score or MIDI, remains a hard problem for multi-instrument music, where overlapping harmonics and diverse timbres must be disentangled. Progress is held back by a data bottleneck: note-level annotation is expensive and requires expert musicians, leaving labelled datasets for transcription very scarce. This talk explores self-supervised learning (SSL) as a way to sidestep that bottleneck by learning musical representations from unlabelled audio. I will focus on a controlled comparison of two masked-modelling paradigms: reconstruction-based learning with a masked autoencoder (MAE), which reconstructs masked spectrogram regions, and predictive learning with a joint embedding predictive architecture (JEPA), which predicts masked regions directly in representation space. Using identical Transformer encoders, I will show that the two objectives are broadly comparable on transcription performance but offer complementary strengths depending on instrument and texture. Initialising JEPA from an MAE-pretrained encoder produces consistent further gains. I will also discuss what the learned latent spaces reveal about how each objective organises musical information, and why no single geometric metric reliably predicts downstream transcription performance.