Keep pulling the thread on Noam Shazeer.
Gemma 3 is a multimodal addition to the Gemma family of lightweight open models.
Gemma 3 introduces vision understanding abilities and wider language coverage.
Gemma 3 supports a context length of at least 128,000 tokens.
The architecture of Gemma 3 was changed to reduce KV-cache memory usage for long contexts.
The Gemma3-4B-IT model is competitive with the Gemma2-27B-IT model across benchmarks.
The Gemma3-27B-IT model is comparable to Gemini-1.5-Pro across benchmarks.
Gemma 3 models are available in sizes ranging from 1 to 27 billion parameters.
Gemma 3 reduces KV-cache memory by increasing the ratio of local to global attention layers and keeping the span on local attention short.
Gemma 3 models achieve superior performance compared to Gemma 2 for both pre-trained and instruction-finetuned versions.
A novel post-training recipe significantly improves Gemma 3's abilities in math, chat, instruction-following, and multilingual tasks.
All Gemma 3 models are released to the community.