Introducing Storage Buckets on the Hugging Face Hub
Introducing Storage Buckets on the Hugging Face Hub Introducing Storage Buckets on the Hugging Face Hub Published March 10, 2026 Update on GitHub Upvote 198 Lucain Pouget Wauplin Eliott Coyac coyotte508 Adrien Carreira XciD Victor Mustar victor Julien Chaumond julien-c Quentin Lhoest lhoestq Pierric Cistac pierric Sylvestre Bcht Sylvestre Hugo Larcher hlarcher Rajat Arya rajatarya Di Xiao seanses Assaf Vayner assafvayner Hugging Face Models and Datasets repos are great for publishing final artifacts. But production ML generates a constant stream of intermediate files (checkpoints, optimizer states, processed shards, logs, traces, etc.) that change often, arrive from many jobs at once, and rarely need version control.
This PriceChange is relevant to the technology intelligence record because it involves Hugging Face, GitHub, Amazon Web Services. The source article should remain the factual reference for follow-up coverage.
- Introducing Storage Buckets on the Hugging Face Hub Published March 10, 2026 Update on GitHub Upvote 198 Lucain Pouget Wauplin Eliott Coyac coyotte508 Adrien Carreira XciD Victor Mustar victor Julien Chaumond julien-c Quentin Lhoest lhoestq Pierric Cistac pierric Sylvestre Bcht Sylvestre Hugo Larcher hlarcher Rajat Arya rajatarya Di Xiao seanses Assaf Vayner assafvayner Hugging Face Models and Datasets repos are great for publishing final artifacts.
- But production ML generates a constant stream of intermediate files (checkpoints, optimizer states, processed shards, logs, traces, etc.) that change often, arrive from many jobs at once, and rarely need version control.
- Storage Buckets are built exactly for this: mutable, S3-like object storage you can browse on the Hub, script from Python, or manage with the hf CLI.
- And because they are backed by Xet , they are especially efficient for ML artifacts that share content across files.
- Why we built Buckets Git starts to feel like the wrong abstraction pretty quickly when you're dealing with: Training clusters writing checkpoints and optimizer states throughout a run Data pipelines processing raw datasets iteratively Agents storing traces, memory, and shared knowledge graphs The storage need in all these cases is the same: write fast, overwrite when needed, sync directories, remove stale files, and keep things moving.
- A Bucket is a non-versioned storage container on the Hub.