“The misalignment in BERT's representations is caused by its flat design, which couples representation learning to the token reconstruction loss.”