“CrossBERT is a two-part architecture that separates the learning of encoded representations from the token reconstruction process.”