Case TS-6CD3C8CB25 Sept 2026research

AI

“Training GPT-3 used about 1,287 megawatt-hours of electricity.”

Plain restatementThe electricity consumed to train OpenAI's 175-billion-parameter GPT-3 model was approximately 1,287 MWh.

Mostly accurateConfidence High
What this verdict means →

This number checks out. The figure of 1,287 megawatt-hours comes from a 2021 paper by researchers at Google and UC Berkeley, which states exactly that for training GPT-3, along with 552 tonnes of CO2 equivalent. It is worth knowing that this is a calculation rather than a meter reading. OpenAI never published an energy total, so the researchers built the estimate from the training compute figure OpenAI did publish, a per-GPU power draw OpenAI supplied, a GPU count taken from an NVIDIA press release, and an assumption about how efficiently the datacenter ran. The arithmetic reproduces cleanly from those stated inputs. One limit matters: the figure covers the single final training run only, so it leaves out the failed and trial runs beforehand, the energy to build the hardware, and everything spent running the model since. No independent party has ever re-measured it, so every later citation of the number traces back to this one calculation.

The drift / as claimed vs as evidenced

[drifted from the evidence:] Training GPT-3 [drifted from the evidence:] used about 1,287 [drifted from the evidence:] megawatt-hours of electricity.


[added by the neutral restatement:] The electricity consumed to train OpenAI's 175-billion-parameter GPT-3 [added by the neutral restatement:] model was approximately 1,287 [added by the neutral restatement:] MWh.

Red-tinted words in the claim drifted from the evidence. Green-tinted words are what a neutral restatement needs.

The trace / claim to source

Where it appeared
Submitted text
Tertiary sourcemixed
Assorted downstream papers and press repeating the figure (Cerebras blog, "How Hungry is AI?", trade coverage)
Secondary sourcepreprint
Li et al., "Making AI Less Thirsty," arXiv:2304.03271
Primary sourcepreprint / lab technical report, authors include the estimate's originators
Patterson et al., "Carbon Emissions and Large Neural Network Training," arXiv:2104.10350 (Google and UC Berkeley, 2021), main text and Appendix A
Primary sourcelab technical report
Du et al., "GLaM: Efficient Scaling of Language Models with Mixture-of-Experts," arXiv:2112.06905, Appendix F
Primary sourcerefereed/edited venue
Patterson et al., "The Carbon Footprint of Machine Learning Training Will Plateau, Then Shrink," arXiv:2204.05149 / IEEE Computer 55(7), 2022
Primary sourcerefereed journal
Luccioni, Viguier, Ligozat, "Estimating the Carbon Footprint of BLOOM," JMLR vol. 24
● Primary source found
What is true
  • The number 1,287 MWh is not invented. It appears verbatim in Patterson et al. 2021 as GPT-3's estimated training energy consumption, and is repeated unchanged in the authors' 2022 follow-on and in Google's GLaM appendix.
  • The claim's hedge, "about," matches how the source presents the number: the paper calls it an estimate.
  • The figure refers to the right model, GPT-3 at 175 billion parameters, and is not a confusion with another model in the same table.
  • Two of the three key inputs behind the number came from OpenAI directly: the published 3.14E+23 total training FLOPs and a measured 330W per V100 running GPT-3. The number is not pure outside guesswork.
  • The figure includes datacenter overhead, not just chip draw, via an assumed PUE of 1.10.
What is misleading
  • The claim reads as a measurement when the source presents a calculation. OpenAI never published a training energy total for GPT-3. The 1,287 MWh figure was constructed by outside researchers from a FLOP count, a per-GPU wattage, a GPU count taken from an NVIDIA press release, and an assumed datacenter efficiency factor. The word "about" partly covers this, and the inputs are disclosed rather than hidden, so the gap is one of provenance rather than accuracy.
  • The figure covers one training run, not the whole development effort. It excludes the trial runs, failed runs, and hyperparameter searches that preceded the final model, and it excludes the energy embodied in manufacturing the hardware and all energy spent running the model afterward. A reader who takes 1,287 MWh as "what it cost to make GPT-3" is reading more into it than the source supports. The true all-in figure is higher by an unknown factor, not lower.
What is uncertain
  • The GPU count of 10,000 rests on an NVIDIA press release rather than on OpenAI's own disclosure. Energy scales directly with this number in the formula, so an error here propagates in full.
  • The PUE of 1.10 is an assumption about which Microsoft datacenter ran the job and how efficiently it was operating at that time. OpenAI did not identify the facility.
  • The 14.8 day training duration is derived, not reported. It assumes the cluster sustained 24.6 TFLOPS per GPU throughout, with no accounting for idle time, restarts, or stragglers. Real runs are rarely uninterrupted, which would push the true energy figure up.
  • No independent party has re-derived the figure from separate data. Every downstream citation found traces back to this one calculation, so the volume of repetition adds nothing to its reliability.
  • Whether OpenAI or Microsoft holds internal metered figures that would confirm or revise the estimate is not publicly known. Neither has published one.
Evidence summary

The figure is real, specific, and traces to one identifiable origin. The Patterson et al. paper states directly that for GPT-3, its estimated carbon emissions due to training are 552 tCO2e and its energy consumption is 1287 MWh. The paper's Appendix A discloses how the number was built. NVIDIA's press release about GPT-3 suggested OpenAI used 10,000 V100 GPUs, and for training time, OpenAI published the total number of floating point operations to train the model, 3.14E+23, while OpenAI told the authors the V100 runs GPT-3 at 24.6 TeraFLOPS/sec, giving roughly 14.8 days for 10,000 GPUs to compute 3.14E+23 FLOPS. On power draw, OpenAI measured V100s as running GPT-3 at 330W. The authors also state that they used the US average CO2e/KWh for GPT-3 at Microsoft Azure. The energy formula used is hours to train multiplied by number of processors multiplied by average power per processor, with all server components including local memory and network links counted in "processor," plus datacenter energy to power and cool the hardware. The datacenter overhead assumption is stated in a separate primary source: the datacenter PUE was 1.10 at the time of training GPT-3 (Patterson et al., 2021), in the same appendix where Google reports 213 MWh for GLaM, 1/6 of the energy cost of GPT-3, 1287 MWh. Reconstructing the arithmetic from those disclosed inputs: 10,000 GPUs × 0.330 kW × 14.8 days × 24 h × 1.10 PUE ≈ 1,289 MWh, which reproduces the published 1,287 MWh. The figure is internally consistent with its stated inputs. No independent measurement exists. OpenAI has never published a training energy total for GPT-3; the parts it contributed were the FLOP count, a per-GPU throughput figure, and a measured per-GPU wattage, not an energy total.

Complete reasoning
The primary artifact was retrieved and states the number the claim makes, in the same units, for the same model, and the claim's "about" matches the source's own hedging. The arithmetic reproduces from the inputs the paper discloses, and the figure is stable across three primary documents including one from Google's own GLaM team. I considered "Accurate" and rejected it because the claim reads as a measured quantity while the source is a modeled estimate covering only the final training run, a real omission even if a minor one. I considered "Partially accurate but misleading" and rejected it because that omission does not change the operative proposition: the source does say roughly this much electricity was used for training GPT-3, and nothing in the evidence bounds that away. "Unverified" and "False" are both plainly wrong here, since the artifact exists and contains the number. Confidence is High because the deciding document was retrieved and read directly; the residual uncertainty sits in the estimate's own inputs, which is a property of the underlying figure rather than a gap in this investigation.
Use this case

The reply is formatted for pasting into the thread where the claim is circulating.

Compact share page: ai.trueseeker.com/s/6cd3c8cb6549/ywgZNRsrlvLkhZ6WH5C49fLLxL6

Ask this case

Answers come only from the case file above; nothing is added.

Is it true that training GPT-3 used about 1,287 megawatt-hours of electricity?

Yes, this figure is accurate and comes from a 2021 study by researchers at Google and UC Berkeley. It is presented as an estimate, which matches the word 'about' in the claim.

Did OpenAI measure and publish this energy figure themselves?

No. OpenAI never published a training energy total for GPT-3. Outside researchers calculated the 1,287 MWh figure using OpenAI's published FLOP count and measured per-GPU wattage, along with a GPU count from an NVIDIA press release and an assumed datacenter efficiency factor.

Does this number cover everything it took to develop GPT-3?

No. It only covers the single final training run. It leaves out failed and trial runs beforehand, the energy used to manufacture the hardware, and any energy used running the model since training.

How reliable are the inputs used to calculate this figure?

Some inputs are solid, like the FLOP count and measured GPU wattage that OpenAI provided. Others carry more uncertainty, including the GPU count from an NVIDIA press release, an assumed datacenter efficiency rating, and a training duration that assumes no interruptions or restarts.

Has anyone independently verified the 1,287 MWh figure?

No. The case file states that no independent party has re-derived the figure from separate data, and every later citation traces back to this one original calculation.

Similar cases on record