“A scheme has been proposed for learning the unmasking order in masked diffusion models by using an additional lightweight policy network.”