JetSpec Locally: Breaking the Speed Ceiling of LLM Inference - Up to 9x
This video locally installs and tests JetSpec's new speculative decoding live: real speedup numbers, no hype. π¬Weekly AI Newsletter: https://fahdmirza.substack.com/ π₯ Get 50% Discount on any A6000 or A5000 GPU rental, use following link and coupon: https://bit.ly/fahd-mirza Coupon code: FahdMirza π₯ Buy Me a Coffee to support the channel: https://ko-fi.com/fahdmirza #jetspec PLEASE FOLLOW ME: βΆ LinkedIn: / fahdmirza βΆ YouTube: / @fahdmirza βΆ Blog: https://www.fahdmirza.com RESOURCES: βΆ https://github.com/hao-ai-lab/JetSpec All rights reserved Β© Fahd Mirza