Keep pulling the thread on Oriol Vinyals.
The FoMo-in-Flux benchmark for continual multimodal pretraining was constructed using 63 datasets and is designed with realistic compute constraints and practical deployment requirements.
Research using the FoMo-in-Flux benchmark investigated continual pretraining by analyzing data mixtures, stream orderings, various update methods including fine-tuning and model merging, meta learning rate schedules, and the effects of model and compute scaling.
The benchmark and code for FoMo-in-Flux are publicly available on GitHub in the ExplainableML/fomo_in_flux repository.