Google Gemma 4 12B
MOUNTAIN VIEW, Calif.: Google releases Gemma 4 12B, a multimodal open-weight AI model designed to run locally on laptops with just 16GB of memory.
Google DeepMind announces the model under an Apache 2.0 license. Gemma 4 12B handles text, images, and native audio inputs without separate encoders. The encoder-free architecture cuts latency and memory requirements significantly versus traditional multimodal designs.
The model packs 12 billion parameters and delivers performance close to much larger systems. Google benchmarks show it approaching the Gemma 4 26B mixture-of-experts model on key tasks. This makes it the company’s first mid-sized Gemma model with native audio support.
Also read: Google Launches Two New AI Research Agents via Gemini API
Weights deploy freely on Kaggle and Hugging Face at just under 18GB total. The model runs through Hugging Face Transformers, vLLM, SGLang, MLX, llama.cpp, and LiteRT-LM. Developers can serve it as an OpenAI-compatible local API through the new litert-lm CLI.
Google also launches AI Edge Gallery and AI Edge Eloquent for macOS users. Both apps run Gemma 4 12B fully on-device, processing voice and visual inputs locally. A sandboxed Python execution loop lets users plot scientific charts inside the chat interface.
DRAM prices jumped roughly 90% in Q1 2026 as memory production redirected toward AI data centers. Micron told CNBC at CES it had effectively sold out memory capacity for 2026. A capable 16GB-memory model sidesteps both the hardware crunch and ongoing cloud inference costs.
Gemma models have now crossed 150 million total downloads since launch.
The launch reflects a broader industry pivot toward on-device AI deployment recently. Microsoft pushed Surface Laptop Ultra with RTX Spark earlier this week for local AI workloads. Apple Intelligence, Anthropic, and OpenAI all explore similar device-resident model strategies.
San Francisco: OpenAI has disclosed what it is calling an "unprecedented cyber incident," in which…
New Delhi: OpenAI has announced the launch of OpenAI Presence Corporate Software, a new enterprise-grade…
SAN FRANCISCO: OpenAI has released ChatGPT Work, a new agent mode designed to take a…
California: Google DeepMind has released three new AI models, Gemini 3.6 Flash, Gemini 3.5 Flash-Lite,…
HANGZHOU, China: Alibaba unveils Qwen 3.8, a 2.4 trillion-parameter model its team ranks "second only…
SAN FRANCISCO: Prasanna Sankar, co-founder and former CTO of Rippling, launches Vorflux with $15 million…