🤗 Hugging Cast S2E6 - Scale LLMs with Intel Gaudi and Xeon
Hugging Cast is a live show about building AI with open source. In this episode, Regis, Ella, Ilyas and Jeff show you how you can accelerate and scale your Gen AI workloads using the latest Intel AI Accelerators, Gaudi 3 and Xeon CPUs, easily with our open source libraries Optimum Intel, Optimum Habana, TGI Gaudi and more. Last we show you how you can run your own benchmarks easily using Optimum Benchmark. Useful resources: https://github.com/huggingface/optimum-intel https://github.com/huggingface/optimum-habana https://github.com/huggingface/tgi-gaudi https://huggingface.co/docs/optimum/main/en/intel https://huggingface.co/docs/optimum/main/en/habana Discussed examples: https://huggingface.co/blog/intel-starcoder-quantization https://huggingface.co/blog/cost-efficient-rag-applications-with-intel https://huggingface.co/blog/universal_assisted_generation https://huggingface.co/blog/setfit-optimum-intel