Keep pulling the thread on Yoshua Bengio.
A scheme has been proposed for learning the unmasking order in masked diffusion models by using an additional lightweight policy network.
The proposed adaptive order policy for masked diffusion models outperforms common heuristics on problems sensitive to token ordering, such as combinatorial tasks and protein sequence generation.
The proposed adaptive order policy for masked diffusion models uses a loss function that reweights terms according to policy probabilities, resulting in a policy that favors positions where the denoiser is more likely to be correct.