EmbeddingGemma 2: open multimodal embedding model for on-device search
Google DeepMind's 740M-parameter open model that maps text, code, images, video, and audio into one 768-dim embedding space, released under Apache 2.0.
Why it matters
First open embedding model to unify all four modalities natively at sub-1B size — fully offline, privacy-first multimodal RAG on a phone or laptop.
What you could build with it
An indie developer could build a fully offline personal media organizer that finds photos, videos, and voice memos from a typed query, with all embeddings computed on the phone. A small dev team could ship private RAG over a company's meeting recordings, screenshots, and code repos without any data leaving the laptop.
Does it hold up?
Day one, but promising: GGUF builds from ggml-org and Unsloth landed within hours, the sentence-transformers quick start works out of the box, and r/LocalLLaMA testers (369-upvote launch thread) are already benchmarking it; independent MTEB comparisons still pending.
Built with EmbeddingGemma 2
- Zubiqo: Google DeepMind Drops EmbeddingGemma 2blog · Independent writeup analyzing the Apache 2.0 release and the MTEB Code jump from 68.76 to 78.68 for local codebase indexing.
- Google AI Edge Gallery: Instant Media Search and Video Moments Finderblog · Google's on-device test app is adding phone photo/video search driven by EmbeddingGemma 2, plus a Mac app (AI Edge Foresight) that searches meeting transcripts and notes.
- Undercode News breakdown of the launchblog · Detailed walkthrough of the unified vector space approach, Matryoshka embeddings, and on-device RAM figures.
- ggml-org/embeddinggemma-2-GGUFgithub · Official GGUF conversion of EmbeddingGemma 2 for llama.cpp local runners (verified live on HF).
- unsloth/embeddinggemma-2-GGUFgithub · Unsloth's GGUF release for fast local inference tooling (verified live on HF).
Learn more
First spotted on hf: source.
More AI models
Mistral Large 4 (Le Chonk): 1.05T-param open-weight MoEMistral's new flagship: a 1.05T-parameter mixture-of-experts multimodal model with 49B active params and 1M-token…models · JEV 0.72ElevenLabs Eleven v4 + v4 Turbo: new emotive TTS architecture with a 100ms real-time variantElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28: a new text-to-speech architecture with inline…models · JEV 0.72VoiceStudio: open-source, fully-local ElevenLabs alternative with voice cloning, dubbing and dictation in 646 languagesVoiceStudio is an open-source, fully-local voice AI studio: voice cloning from a 3-second sample, voice design, video…models · JEV 0.71Mistral Forge: enterprise platform for training and continuously improving proprietary modelsMistral AI launched Forge (Sep 28) — a full-lifecycle model training platform (pre-training, SFT, DPO/ODPO, RL…models · JEV 0.69DeepSeek V4-Flash official API: public beta with upgraded agent capabilities, Responses API and Codex supportDeepSeek's official V4-Flash is now a public beta API with massively upgraded agent capabilities - benchmark scores…models · JEV 0.69kev: open, trainable Jev-like family of small decision models on Qwen3.5/3.8A family of small decision models (0.8B to 27B) built on Qwen3.5/3.8 that reproduces Jev's prefill-only typed-decision…models · JEV 0.68
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.