Video by Hugging Face via YouTube

If you work with large datasets, you’re probably re-uploading the same bytes every time the data changes.
In this video I upload a 191 MB NYC taxi CSV to a Hugging Face Bucket, append 100 rows, and upload it again: Xet, Hugging Face’s storage layer, sends only the new pieces. Then I convert the data to Parquet, prepend 100 rows, and sync again: with PyArrow’s use_content_defined_chunking=True, only 8.6 MB of the 52.9 MB file goes over the wire. For data and ML engineers who push datasets to the Hub and keep them up to date.
Topics covered:
– How Xet uploads only the pieces that changed
– Hugging Face Buckets with the hf CLI (create, cp, sync)
– Content-defined Parquet blocks with PyArrow 21+
– Upload savings: ~90 MB vs ~6 MB on a 100 MB file
– Production gotchas: S3 tools, versioning, sorting, writer settings
Resources:
– Full walkthrough (commands, code, dataset): https://huggingface.co/blog/prpatel/hands-on-with-buckets
– Hugging Face forum: https://discuss.huggingface.co/
Chapters:
0:00 Stop wasting bandwidth
0:15 What you’ll see
0:39 Demo: CSV and Parquet uploads to a Bucket
9:11 Xet, in three steps
10:10 The one-line fix
10:47 How much Xet saves
11:13 Before you ship it
12:27 Join the discussion