HF Buckets: Upload Only What Changed with Xet
If you work with large datasets, you're probably re-uploading the same bytes every time the data changes. In this video I upload a 191 MB NYC taxi CSV to a Hugging Face Bucket, append 100 rows, and upload it again: Xet, Hugging Face's storage layer, sends only the new pieces. Then I convert the data to Parquet, prepend 100 rows, and sync again: with PyArrow's use_content_defined_chunking=True, only 8.6 MB of the 52.9 MB file goes over the wire. For data and ML engineers who push datasets to the Hub and keep them up to date. Topics covered: - How Xet uploads only the pieces that changed - Hugging Face Buckets with the hf CLI (create, cp, sync) - Content-defined Parquet blocks with PyArrow 21+ - Upload savings: ~90 MB vs ~6 MB on a 100 MB file - Production gotchas: S3 tools, versioning, sorting, writer settings Resources: - Full walkthrough (commands, code, dataset): https://huggingface.co/blog/prpatel/hands-on-with-buckets - Hugging Face forum: https://discuss.huggingface.co/ Chapters: 0:00 Stop wasting bandwidth 0:15 What you'll see 0:39 Demo: CSV and Parquet uploads to a Bucket 9:11 Xet, in three steps 10:10 The one-line fix 10:47 How much Xet saves 11:13 Before you ship it 12:27 Join the discussion