“The 2014 autoregressive model by Sutskever, Vinyals, and Lee used a pipelining parallelization strategy across 8 GPUs to achieve a 3.5x speedup.”