Embedding Gemma 2: Multimodal Search & RAG (Free Colab)

From the creator

EmbeddingGemma 2 is a new open embedding model from Google DeepMind that puts text, code, images, video and audio into one shared vector space. It has 740M parameters, an Apache 2.0 license, and it's small enough to run on a phone. In this video I explain why multimodal retrieval matters for RAG and search, how people solved it before (caption-and-transcribe pipelines, CLIP, ImageBind, ColPali), and how EmbeddingGemma 2 works under the hood: modality encoders, one shared Gemma 4 backbone, mean pooling, contrastive training, Matryoshka embeddings, and modular loading. Then we walk through a Colab notebook you can run on a free T4 GPU: photo search, multilingual search, voice search with no transcription, sound search, finding moments in a video, searching PDF pages without OCR, and how small you can make the vectors. Let me know in the comments what you'd build with this, and whether you want a follow-up on fine-tuning EmbeddingGemma 2. Colab notebook: https://colab.research.google.com/drive/1mXzvOCo-_y4r1yQ0AqAJJtU8BgU3LlCP Model card: https://huggingface.co/google/embeddinggemma-2 Launch blog: https://blog.google/innovation-and-ai/technology/developers-tools/embeddinggemma-2/ Developer guide: https://developers.googleblog.com/embeddinggemma-2-the-developer-guide/ On-device guide (LiteRT, MediaPipe): https://developers.googleblog.com/google-ai-edge-with-embeddinggemma-2/ Gemini Embedding 2 paper: https://arxiv.org/abs/2605.27295 CLIP paper: https://arxiv.org/abs/2103.00020 ImageBind paper: https://arxiv.org/abs/2305.05665 ColPali paper: https://arxiv.org/abs/2407.01449 My ColPali videos: https://www.youtube.com/watch?v=rhJJynv47Pw https://www.youtube.com/watch?v=DI9Q60T_054 LocalGPT: https://github.com/PromtEngineer/localGPT My voice to text App: whryte.com Website: https://engineerprompt.ai/ RAG Beyond Basics Course: https://prompt-s-site.thinkific.com/courses/rag Signup for Newsletter, localgpt: https://tally.so/r/3y9bb0 💻 Pre-configured localGPT VM: https://bit.ly/localGPT (use Code: PromptEngineering for 50% off). 00:00 - EmbeddingGemma 2: Multimodal Embeddings 00:40 - Demo: Search Photos With Your Voice 01:20 - Why Multimodal Retrieval Matters for RAG 02:25 - Before: Convert Everything to Text 03:21 - Before: CLIP, ImageBind and ColPali 04:14 - One Backbone for Every Modality 04:33 - How It Works: Embeddings, Encoders, Tokens 06:07 - Mean Pooling and the 768-Number Vector 06:29 - Contrastive Training 07:15 - Matryoshka Embeddings: Smaller Vectors 08:05 - Modular Loading: 270M to 740M 09:00 - Google Colab Notebook 19:11 - Verdict Credits: Big Buck Bunny © Blender Foundation (CC BY 3.0). ESC-50 by K. Piczak (CC BY-NC 3.0). Flickr30k test images. Benchmark numbers in the explainer are Google's; results in the notebook section are from my own Colab T4 run.

Choose to Build with AI
Matched to Prompt Engineering

AI Maker Residence at KOKO

The third AI workshop taught by our legendary teacher, Nick Sarafa. In one of the last events we did an asset manager raised an additional £25M on their fund within a space of 9 months. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence. One person did.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence at KOKO
Live event
AI Maker Residence at KOKO
Fri 09 Oct 2026

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.