Superfast RAG with Llama 3 and Groq

From the creator

Groq API provides access to Language Processing Units (LPUs) that enable incredibly fast LLM inference. The service offers several LLMs including Meta's Llama 3. In this video, we'll implement a RAG pipeline using Llama 3 70B via Groq, an open source e5 encoder, and the Pinecone vector database. 📌 Code: https://github.com/pinecone-io/examples/blob/master/integrations/groq/groq-llama-3-rag.ipynb 🌟 Build Better Agents + RAG: https://platform.aurelio.ai (use "JBMARCH2025" coupon code for $20 free credits) 👾 Discord: https://discord.gg/c5QtDB9RAP Twitter: https://twitter.com/jamescalam LinkedIn: https://www.linkedin.com/in/jamescalam/ #artificialintelligence #llama3 #groq 00:00 Groq and Llama 3 for RAG 00:37 Llama 3 in Python 04:25 Initializing e5 for Embeddings 05:56 Using Pinecone for RAG 07:24 Why We Concatenate Title and Content 10:15 Testing RAG Retrieval Performance 11:28 Initialize connection to Groq API 12:24 Generating RAG Answers with Llama 3 70B 14:37 Final Points on Why Groq Matters

Choose to Build with AI
Matched to Hugging Face

AI Maker Residence 3

The third AI workshop taught by our legendary teacher, Nick Sarafa. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence 3
Live event
AI Maker Residence 3
Fri 09 Oct 2026

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.