Keep pulling the thread on Ilya Sutskever.
The 2014 paper by Sutskever, Vinyals, and Lee presented an early version of the scaling hypothesis, stating that training a very large neural network on a very large dataset guarantees success.
The AI industry has reached "peak data" for pre-training models because the amount of data available on the internet is finite.
The 2014 paper by Ilya Sutskever, Oriol Vinyals, and Kwok Lee was the first to strongly believe that a well-trained autoregressive neural network could achieve any desired outcome, such as machine translation.
The 2014 autoregressive model by Sutskever, Vinyals, and Lee used a pipelining parallelization strategy across 8 GPUs to achieve a 3.5x speedup.
The development of models like GPT-2 and GPT-3 marked the beginning of the "age of pre-training" in AI.
Alec Radford, Jared Kaplan, and Dario Amodei were key contributors to the work that led to the "age of pre-training" in AI.
The O-1 model is a recent, vivid example of a system that utilizes significant inference-time compute.
Hominids exhibit a different brain-to-body mass scaling exponent compared to other non-human primates and mammals.