“The Music-JEPA model predicts a future audio state based on a given current audio state and a pianoroll action.”