“Google's Gemini model can understand text, vision, and audio inputs, but can only generate text as output.”