“Georgi Gerganov performed a 4-bit quantization of the Llama model, enabling it to run inference on an M1 or M2 processor.”