What happened?
Hugging Face Code has released the Stack v3 training dataset, built from a direct August 2025 GitHub crawl and containing 173 million repositories and roughly 4.9 trillion tokens.
Dataset scale does not replace provenance and rights management. Model builders should distinguish the dataset license from repository licenses and separately manage benchmark contamination and generated-code licensing risks.
Why does it matter?
The recency, licensing and deduplication of training data strongly affect coding models, so a larger public data layer strengthens the foundation for reproducible open models.
Who should care?
AIZIGOO view
Dataset sizes follow the Hugging Face card and may depend on its repository and token counting methodology.