Keep pulling the thread on Yann LeCun.
The Music-JEPA model learns a world model of piano sound by framing music as an action-conditioned system where audio is the state and the pianoroll is the instrument action.
The Music-JEPA model predicts a future audio state based on a given current audio state and a pianoroll action.
Experiments with Music-JEPA demonstrate that the model learns the relationships between musical actions and their resulting sound.
The latent representations learned by Music-JEPA can be used for downstream tasks such as beat tracking, composer identification, and key estimation.
The Music-JEPA model enables piano transcription through a planning process that searches for the actions that best explain a target sound.