← Back to the wire

Welcome Gemma 4: Frontier multimodal intelligence on device

AnnouncementModelApr 2, 2026

Google DeepMind released Gemma 4, a family of multimodal models now available on Hugging Face under Apache 2 licenses. The models support image, text, and audio inputs across five sizes, with the 31B dense variant achieving an estimated LMArena score of 1452. Gemma 4 introduces Per-Layer Embeddings and a shared KV cache for efficiency. Hugging Face contributors including merve, Pedro Cuenca, and Sergio Paniego detailed integration with transformers, llama.cpp, MLX, and other frameworks. DiffusionGemma enables text generation via diffusion.

Receipt № 5211 source · awaiting confirmation ◐

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

01highPRIMARY
Gemma 4ModelDiffusionGemmaModelPedro CuencaPersonSergio PaniegoPersonben burtenshawPersonmervePersonNathan HabibPersonGoogle DeepMindCompanySteven ZhengPersonAlvaro BartolomePersonHugging FaceCompany
Canonical: https://huggingface.co/blog/gemma4