Open sourceOPEN SOURCE

Hugging Face releases the open Stack v3 code dataset with about 4.9 trillion tokens

Hugging Face Code has released the Stack v3 training dataset, built from a direct August 2025 GitHub crawl and containing 173 million repositories and roughly 4.9 trillion tokens.

07/23/20261 sources reviewed
Quick summary
  • The training set contains 15.9 TB across 713 programming languages and 173 million repositories.
  • It includes repository context, license detection, PII redaction and a process for honoring developer opt-outs.
  • It is distributed under ODC-BY, while original licenses for individual source files still need review.
WHAT HAPPENED

What happened?

Hugging Face Code has released the Stack v3 training dataset, built from a direct August 2025 GitHub crawl and containing 173 million repositories and roughly 4.9 trillion tokens.

Dataset scale does not replace provenance and rights management. Model builders should distinguish the dataset license from repository licenses and separately manage benchmark contamination and generated-code licensing risks.

WHY IT MATTERS

Why does it matter?

The recency, licensing and deduplication of training data strongly affect coding models, so a larger public data layer strengthens the foundation for reproducible open models.

WHO SHOULD CARE

Who should care?

Open-model developersCoding-AI researchersData governance teams
AIZIGOO VIEW

AIZIGOO view

Dataset sizes follow the Hugging Face card and may depend on its repository and token counting methodology.