Training Your Own Embedding Model Is Not As Hard As You Think
Fine-tune Embedding Gemma 2 (EmbeddingGemma 2) on your own data with Unsloth and LoRA in a free Colab notebook. I show how fine-tuning an embedding model differs from fine-tuning an LLM, then train Google's multimodal embedding model on two tasks: environmental sounds (ESC-50 top-1 24.5% to 65.8%) and search over my own YouTube transcripts (66.3% to 75.0% top-1 on videos it never saw), and check what it costs the rest of the model. You'll learn how to build question-passage training pairs, why in-batch negatives (MultipleNegativesRankingLoss) need a no-duplicates sampler, why training and search prompts must match, how LoRA works on a shared multimodal backbone, how to evaluate on unseen data, and a reload bug that silently breaks audio adapters. If you have fine-tuned embedding models for your own applications, share what worked for you in the comments. Links: Colab notebook: https://colab.research.google.com/drive/1xq-84DLZsqteEzL52Qb__47Y_yGckeyB?usp=sharing The first EmbeddingGemma 2 video: [EMBEDDINGGEMMA 2 VIDEO LINK] Unsloth EmbeddingGemma 2 guide: https://unsloth.ai/docs/models/embeddinggemma-2 Unsloth notebooks: https://github.com/unslothai/notebooks Model card: https://huggingface.co/google/embeddinggemma-2 ESC-50 dataset: https://github.com/karolpiczak/ESC-50 Sentence Transformers losses: https://sbert.net/docs/package_reference/sentence_transformer/losses.html 00:00 - Fine-Tuning EmbeddingGemma 2 on Your Own Data 01:03 - Embedding Fine-Tuning vs LLM Fine-Tuning 02:30 - In-Batch Negatives (MultipleNegativesRankingLoss) 03:37 - The Duplicate Trap: No-Duplicates Sampler 04:45 - LoRA on a Shared Multimodal Backbone 06:30 - Notebook: ESC-50 Sounds with Unsloth on an A100 09:32 - Side Effects on Photo, Voice and Text Search 10:06 - Gotcha: Reloading an Audio Adapter 10:50 - Fine-Tuning on YouTube Transcripts 12:04 - Three Things to Get Right