“Recurrent Neural Networks have largely been superseded because they empirically underperform compared to the Transformer architecture.”