GPT-Fast - blazingly fast inference with PyTorch (w/ Horace He)
Become a Patreon: https://www.patreon.com/theaiepiphany π¨βπ©βπ§βπ¦ Join our Discord community: https://discord.gg/peBrCpheKE Horace He joined us today to talk more about how to make inference fast using just PyTorch native operations! β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬ https://pytorch.org/blog/accelerating-generative-ai-2/ https://github.com/pytorch-labs/gpt-fast β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬ βοΈ Timetable: 00:00 - 00:45 Intro 00:45 - 02:23 HyperStack GPUs! (sponsored) 02:23 - 08:40 What is GPT-Fast? 08:40 - 28:15 PyTorch compile 28:15 - 32:15 int8 quantization 32:15 - 40:12 Speculative Decoding 40:12 - 42:05 Int 4 quantization 42:05 - 45:25 Putting it all together, tensor parallelism 45:25 - 58:10 Bonus optimizations 58:10 - 01:05:04 Outro, questions β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬ π° SPONSOR The AI Epiphany - https://www.patreon.com/theaiepiphany One-time donation - https://www.paypal.com/paypalme/theaiepiphany Huge thank you to these AI Epiphany patreons: Eli Mahler Petar VeliΔkoviΔ β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬ πΌ LinkedIn - https://www.linkedin.com/in/aleksagordic/ π¦ Twitter - https://twitter.com/gordic_aleksa π¨βπ©βπ§βπ¦ Discord - https://discord.gg/peBrCpheKE πΊ YouTube - https://www.youtube.com/c/TheAIEpiphany/ π Medium - https://gordicaleksa.medium.com/ π» GitHub - https://github.com/gordicaleksa π’ AI Newsletter - https://aiepiphany.substack.com/ β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬ #gptfast #inference #pytorch