“A typical Large Language Model is trained on approximately 20 trillion words, which is equivalent to about 10^14 bytes of data.”