A Holistic Assessment of the Carbon Footprint of Noor, a Very Large Arabic Language Model
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "A Holistic Assessment of the Carbon Footprint of Noor, a Very Large Arabic Language Model".
Tom: In an era where large language models are becoming ubiquitous, it is crucial to consider their environmental impact due to their extreme size and resource use.
Jane: First, who's behind it and why it matters.
Paper summary: Tom: So, looking at the "A Holistic Assessment of the Carbon Footprint of Noor, a Very Large Arabic Language Model," we see that the authors put together a comprehensive view of the carbon cost for this ambitious project. They’ve clearly laid out how every piece—from data collection to deployment—contributes to the final number.
Jane: It really highlights that while pretraining compute is a major driver, it's not the only factor; those R andD and operational elements add a substantial portion of the environmental load, which is something we need to keep in mind for future AI scaling.
Lu: The implication here is that simply training bigger models isn't an automatically positive environmental outcome; we need to be strategic about where and how we train. They recommend assessing the footprint on a per-project basis, which suggests a more granular approach is needed moving forward.
Meng: I think the real impact lies in their recommendation for systematic assessment across the whole project lifecycle; if we treat it as one big carbon problem instead of just one training event, we can start designing systems that are inherently more efficient from day one.
Lalam: For me, this paper is significant because it provides a framework for accountability. By quantifying these various impacts—flights, storage, R andD—it sets a standard for how we should measure the environmental cost of creating large language models.
Tom: The authors conclude that the development of those four Noor models resulted in an estimated thirty-six point five tons of CO2, with sixty-five percent attributed to training compute. They also emphasized that appropriately selecting the location where calculations are performed can significantly reduce this environmental impact.
Jane: And they finish by stressing that large-scale inference could potentially overtake pretraining costs in terms of carbon impact down the line, so we have to keep an eye on that too.
Lu: It seems like the main message is a call for efficiency across the board, suggesting things like efficient architectures and distillation techniques are necessary if we want to manage this footprint effectively as models get bigger.
Meng: I think focusing on hardware choices, specifically looking at data centers with a PUE of one point one and choosing specific accelerators over others for smaller models, is a very tangible step we can take right now.
Lalam: I think the most important implication for our culture is establishing this kind of full energy consumption and CO2e reporting as standard practice for any major AI project moving forward.
Tom: So, the "A Holistic Assessment of the Carbon Footprint of Noor, a Very Large Arabic Language Model" paper gives us a detailed picture that goes way beyond just training numbers. It shows us that managing the carbon impact requires looking at every stage of development and deployment.
Conclusion: Tom: So, to wrap up this whole discussion about Noor, we’re focusing on what that title actually means for us when we look at its authors and their core message.
Jane: It really highlights that the work isn't just about calculating a single number; it's about building a complete picture of the environmental cost across every step of creating this large language model.
Lu: The authors did a fantastic job structuring this assessment, taking the complexity of building an extreme-scale AI project and breaking it down into manageable, quantifiable parts like data ingestion and future inference.
Meng: It shows that they’ve taken a very broad view, making sure to include things like international travel and power usage efficiency in their total carbon bill calculation.
Lalam: This moves the conversation beyond just looking at the massive compute required for pretraining, which is super common, and instead demands a more complete lifecycle perspective for any large AI initiative.
Tom: Exactly! The authors conclude that understanding this full scope is crucial because it shows that we can’t isolate one part of the process to fix the carbon issue effectively.
Jane: And their main implication for us is a shift in how we think about sustainability in AI development—we need to embed these holistic assessments into our planning from day one.
Lu: It suggests that strategic decisions around hardware and location, as they discussed, become much more powerful levers when you have this comprehensive data showing the entire system's footprint.
Meng: If we can use this kind of detailed tracking to guide our engineering choices today, it could lead to significantly leaner and more sustainable models in the future.
Imad Lakim, Ebtesam Almazrouei, Ibrahim Abu Alhaol, Merouane Debbah, Julien Launay
Technology Innovation Institute in the United Arab Emirates · LightOn
cs.CL
Submitted: 2026-09-22
Updated: 2026-09-22
Comments: 11 pages, 3 figures, 2 tables. Published in Proceedings of BigScience Episode #5 -- Workshop on Challenges & Perspectives in Creating Large Language Models (ACL 2022)
Journal ref: Proceedings of BigScience Episode #5 -- Workshop on Challenges & Perspectives in Creating Large Language Models, pages 84-94, Association for Computational Linguistics, 2022
DOI: 10.18653/v1/2022.bigscience-1.8
Code: https://github.com/kingoflolz/mesh-transformer-jax
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 92/100
The gist: In an era where large language models are becoming ubiquitous, it is crucial to consider their environmental impact due to their extreme size and resource use.
Key concepts
- Pretraining Compute
- This refers to the massive computational power required to train the four different sizes of the Noor models (1.5B to 13B parameters). The study found that this stage is responsible for more than half of the project's total carbon emissions, making it a primary area for emission reduction efforts.
- PUE (Power Usage Effectiveness)
- PUE is a metric used to measure how efficiently a data center uses its energy. A lower PUE score indicates better efficiency, meaning less energy is wasted on overhead instead of actual computation. The paper suggests choosing data centers with low PUE scores, like 1.1, to minimize the carbon impact.
- Quantization
- Quantization is a technique used during inference (when the model is being used) to reduce numerical precision. By reducing the precision needed for calculations, this method allows models to process information faster and with less energy consumption, leading to lower carbon emissions during deployment.
Terminology
Summary
In an era where large language models are becoming ubiquitous, it is crucial to consider their environmental impact due to their extreme size and resource use. This work proposes a holistic assessment of the total carbon footprint of Noor, an extremescale Arabic language model project, by evaluating costs from data collection through future inference estimates.
The gist
We evaluate the total carbon bill of the entire project: starting with data collection and storage costs, including research and development budgets, pretraining costs, future serving estimates, and other exogenous impacts sparked by this international cooperation.
Holistic Assessment Scope
The assessment goes beyond traditional focus on pretraining costs to capture the entire project lifecycle. The evaluation includes:
-
Data collection, curation, and storage costs.
-
Research and development budgets and hyper-parameters tuning budgets.
-
Pretraining costs for the four models (1.5B, 2.7B, 6.7B, and 13B parameters).
-
Future serving estimates for inference costs as the models are put into use in the wild.
-
Other exogenous impacts necessary for this international cooperation (e.g., flights, personnel).
Key Findings on Emissions Drivers
The study identifies several key factors influencing the carbon footprint:
We identify pretraining compute as driving more than half of the emissions of the project.
However, other components still contribute significantly:
"all combined, other R&D, storage, and personnel counts still amount for 35% of the carbon footprint."
The paper notes that in scenarios with a low-impact training electric mix,
costs beyond pretraining may become the main sources of emissions.
Factors Influencing Carbon Footprint
Several factors directly related to the models and their deployment are analyzed:
-
Model size: The compute budget is approximated by the formula,
C = 6ND (Kaplan et al., 2020),
andcompute budget will scale more or less linearly with model size.
-
Hardware characteristics: Throughput per Watt is a critical metric, noting that
More efficient hardware will have more throughput per Watt.
-
Modelling decisions: This includes the number of tokens processed (for training or inference) and hardware throughput, which are impacted by decisions like choosing a
more fertile tokenizer
or implementing specific parallelism techniques. -
Data center efficiency: The Power Usage Effectiveness (PUE) is used to assess overall efficiency, with a worldwide average around 1.8, though some providers report lower figures like Google's 1.11.
-
Electricity mix: This is a crucial factor as it
determines the carbon emissions per kWh of electricity.
Carbon Footprint Breakdown
The total estimated electricity consumption for the Noor project is approximately 59.14 MWh, with data preprocessing accounting for 20% of it.
The breakdown shows:
Nearly a third of the energy consumed (30%) went to tasks outside of main models pretraining.
The carbon footprint distribution is highly dependent on location and electricity mix, as shown in Table 2. For instance, the training on the Noor-HPC in the UAE resulted in a footprint of 23.7 tCO2e.
The study also accounts for exogenous costs like international flights, which account for 18% of the total carbon emission of the whole project.
Pathways to Reduction
The paper suggests several recommendations to lower future footprints:
Efficient architectures,
such as Mixture-of-experts (MoE) models, which can bring significant energy savings during training and inference.
Quantization (Yang et al., 2019)
to reduce numerical precision at inference time and accelerate processing.
Distillation (i.e., training a smaller model from the outputs of a larger one)
as a promising direction for making models leaner for inference.
Additionally, focusing on hardware choices is advised: selecting data centers with a PUE of 1.1
and choosing accelerators like T4s over A100s for smaller models can reduce energy consumption. Finally, the paper recommends that the AI community start reporting the full energy consumption and the CO2e of their projects.
Conclusion
The development of the suite of four Noor models is estimated to have emitted 36.5 tons of CO2, with 65% attributed to training. The main driver remains the carbon intensity of the mix used for model training, emphasizing that Appropriately selecting the location of calculations can significantly reduce the environmental impact.
Furthermore, it is stressed that large-scale inference could also rapidly outtake pretraining costs in terms of carbon impact.
Improvements for AI systems
Based on the provided scientific paper, here are specific improvements for AI systems, categorized by modeling, infrastructure, and operational practices:
)Efficient Architectures (Modelling & Engineering):
-
Implement Mixture-of-Experts (MoE) models for future large language models (LLMs). This will allow for significant energy savings during both training and inference by sparsely activating only the relevant experts.
-
Focus on developing more
lean
architectures specifically optimized for inference, potentially through techniques like model distillation (training a smaller, efficient model from the outputs of a larger one) to reduce token generation requirements and computational load during deployment. -
Invest in research into better embeddings and activation functions that have a non-negligible impact on the overall carbon footprint of models.
)Efficient Implementations (Modelling & Engineering):
-
Develop highly optimized distributed training implementations that minimize idle hardware consumption, potentially by incorporating fine-grained effects such as wave and tile quantization to maximize throughput on existing hardware.
-
Prioritize
efficient scaling
techniques, such as switch transformers using simple and efficient sparsity, for scaling models to trillion-parameter sizes while maintaining efficiency gains over dense models.
)Hardware Selection (Infrastructure):
-
When selecting data centers for training, mandate the choice of platforms with a Power Usage Effectiveness (PUE) significantly lower than the world average (e.g., PUE of 1.1 or better).
-
For inference workloads involving smaller models (<3B parameters), strategically utilize more energy-efficient accelerators, such as T4s instead of A100s, leveraging their superior energy efficiency per FLOP when appropriate for the task.
-
Prioritize locating major training operations in regions with a low-carbon electricity mix (e.g., France's mix reported in the study) to drastically reduce the carbon intensity factor of the total footprint.
)Operational Practices (Exogenous Costs & Lifecycle):
-
Mandate comprehensive, end-to-end carbon cost reporting for all large model development projects, including data storage, R&D budgets, pretraining costs, and future serving/inference estimates. This allows for accurate carbon accounting and potential emissions offsetting.
-
Establish rigorous protocols to minimize the environmental impact of international personnel travel (e.g., limiting flights to essential workshops) as these exogenous costs can represent a significant percentage of the total footprint (up to 18% in the Noor project example).
-
Implement best practices for inference deployment by advising end-users and platforms on efficient inference strategies, such as selecting low-impact cloud regions for serving and choosing hardware optimized for the model size.
This paper suggests that future AI systems should be designed with a holistic lifecycle perspective, moving beyond just pretraining compute to include data sourcing, infrastructure choices (PUE), regional electricity mix, and operational efficiency (quantization/MoE) to achieve a significantly reduced total carbon footprint.
Abstract
As ever larger language models grow more ubiquitous, it is crucial to consider their environmental impact. Characterised by extreme size and resource use, recent generations of models have been criticised for their voracious appetite for compute, and thus significant carbon footprint. Although reporting of carbon impact has grown more common in machine learning papers, this reporting is usually limited to compute resources used strictly for training. In this work, we propose a holistic assessment of the footprint of an extreme-scale language model, Noor. Noor is an ongoing project aiming to develop the largest multi-task Arabic language models -- with up to 13B parameters -- leveraging zero-shot generalisation to enable a wide range of downstream tasks via natural language instructions. We assess the total carbon bill of the entire project: starting with data collection and storage costs, including research and development budgets, pretraining costs, future serving estimates, and other exogenous costs necessary for this international cooperation. Notably, we find that inference costs and exogenous factors can have a significant impact on total budget. Finally, we discuss pathways to reduce the carbon footprint of extreme-scale models.
Sources
- On the Opportunities and Risks of Foundation Models
- Unified Scaling Laws for Routed Language Models
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- Scaling Laws for Autoregressive Generative Modeling
- What Changes Can Large-scale Language Models Bring? Intensive Study on HyperCLOVA: Billions-scale Korean Generative Pretrained Transformers
- Quantifying the Carbon Emissions of Machine Learning
- Scaling Language Models: Methods, Analysis & Insights from Training Gopher
- Compute Trends Across Three Eras of Machine Learning
- Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering