🌱 AI’s carbon problem is widely acknowledged but rarely measured at scale. Unlike manufacturing or agriculture, the AI industry still lacks standardized accounting methods that extend beyond individual models. Hugging Carbon uses Hugging Face as a public corpus to estimate aggregate training emissions from energy, compute, and model metadata, even when disclosures are incomplete. It estimates nearly 60,000 metric tons of CO₂e across more than 5,000 popular open-source models.
Turn heterogeneous Hugging Face metadata into emission estimates with empirical validation.
| Tier | NLP repos | CV/MM repos | Total repos |
|---|---|---|---|
| Tier 1 | 390 | 352 | 742 |
| Tier 2 | 944 | 1,679 | 2,623 |
| Tier 3 | 3,053 | 220 | 3,273 |
When disclosures are incomplete, we estimate training emissions with log–log regressions fit on Tier 1 models.
Tier 2 (FLOPs → emissions). Training emissions increase strongly with computational demand: a 1% increase in training FLOPs is associated with ~0.83% higher emissions.
Tier 3 (parameters → emissions). When training FLOPs are missing, we instead regress emissions on parameter count.
| Category | Average ATCI (tCO₂e/EFLOP) | Mean (tCO₂e) | Total (10^4 tCO₂e) |
|---|---|---|---|
| Foundation & Individual Models | 0.14 | 12 | 5.6 |
| Finetuned Models | 0.23 | 8 | 0.4 |
| CV & Multi-Modal (MM) Models | 0.16 | 12 | 2.3 |
| NLP Models | 0.13 | 11 | 3.6 |
@inproceedings{wang2026huggingcarbon,
title = {Hugging Carbon: Quantifying the Training Carbon Emissions of AI Models at Scale},
author = {Wang, Xinlei and Ming, Ruibo and Qiu, Jing and Zhao, Junhua and Gu, Jinjin},
booktitle = {Proceedings of the Forty-Third International Conference on Machine Learning},
year = {2026}
}