As they1 say,

a mathematician is a machine for turning coffee into theorems.

These days, we can extend this into

an LLM is a machine for turning energy into theorems.

Indeed, let’s find out how much coal OpenAI2 turned into a proof3 of Navier Stokes.

Please note: we’re going to do a Fermi estimate here. I’m making up numbers and citing random sources on the internet. They will be off. But hopefully not by more than an order of magnitude?

Input data Link to heading

Most of the source data comes from OpenAI’s announcement. The numbers in there are (intentionally?) such that it’s quite hard to make sense of them. To summarize some facts from their post:

  • Training since Aug 28.
  • Sept 1: rumours of 2 solved millennium problems. They start running the internal model on all open millennium problems, as well as some other problems.
  • Agents were split into groups of varying size and could communicate within groups. The group that solves NS has 10000 agents. It appears they had at least 4 groups running on NS? For each of versions A/B/C/D.
  • They also had a group working on the simplified “Euler” version. This was solved by 100 agents in 50 hours.
  • Then they shifted (most/all?) agents to work on NS?
  • Different groups tried different approaches. (So multiple groups of ~10k, and only the lucky one that got the winning idea is reported?) With cross-pollination for sharing useful results.
  • The solution came on Sept 5, 88h after the first agents were launched. (Then only 17h for formalization and verification, presumably with fewer agents?)
  • In total, across all problems, 300G output tokens, of which 130G for NS.

The wikipedia page has some sources for the fact that in total, the cost of the proof was a couple million USD. A fortune article and new scientist article mention:

  • At first, 1000 agents worked on the Euler problem for 50 hours.
  • Then, 10000 agents worked on the full Navier-Stokes for 11 hours.
  • OpenAI said “this” would cost $15 million for customers.
  • The fortune article mentions it was 1000x more expensive than previous problems which had cost \$2000, putting the internal price at \$2M.

Electricity cost and coal usage Link to heading

Overall, let’s say the internal cost is indeed $2M. Further, let’s assume that 50% of the cost is energy for both running the GPUs and cooling them (with the rest going to hardware, maintenance, and the developers). So we get $1M of energy.

The US cost of bulk electricity is \$25 to \$120 per MWh. Let’s assume OpenAI pays \$50/MWh? That gives us \$1M / (\$50 / MWh) = 20 GWh.

Coal power plants operate at roughly 33% efficiency, so we have to burn 60 GWh worth of coal. Coal has an energy density of 6.7 KWh/kg, making for 60 GWh / (6.7 KWh/kg) = 3000 tonne. A Panamax bulk carrier (the max size that fits through the Panama canal) has 60000 tonne capacity, so this is 5% of a carrier. Coal has a bulk density around 0.8 kg/m^3, so this is ~4000m^3, which is 1.5 2500 m^3 Olympic-size swimming pools.

Bulk coal costs \$100-150 per tonne. Let’s go with \$125/tonne. Then, \$1M buys around 8000 tonnes of raw coal. Thus, the overhead of turning coal into energy and transmitting the energy is around 2.5x (=8000/3000). That sounds about right.

Energy to GPUs Link to heading

Now lets approach this from a different angle. In total, it seems, they had 10000 agents running for 88 hours. Thus, we get an energy consumption of 20 GWh/88h = 230 MW. Let’s say a third of that is cooling, leaving 150 MW for compute. DGX B200 video cards (consisting of 8x B200 cards) are rated at 14 kW max TDP, so let’s say they use around 10 kW during normal operation? (Although you probably want to pretty much saturate your hardware as much as you can probably? idk.) Then it’s 15000 DGX B200 video cards (120k B200’s) working on this. Each of those would cost $0.5M at retail prices, which would be \$7.5 billion of GPUs. (If we assume they run at the full 14 kW TDP, this goes down to 11000 video cards at \$5.4 billion retail cost.)

Agents per GPU Link to heading

Based on the information available, I’m assuming they had all 10k agents running for the full 88 hours: at first spread over various problems, and then later all focused on Navier-Stokes. Thus, they run 10k agents on 15k chips, or 1.5 DGX per agent. Given each DGX B200 has 1.4 TB of memory, using 1.5 of those would give us a 2.1 TB, which would be 4T parameters at 4-bit quantization.

That’s actually pretty close to Anthropic’s Fable, which is said to have around 5T weights.

The one problem I have here is that I always thought that one normally batches/multiplexes many requests/agents to a single chip. E.g., that 100 agents would run on a single (or 1.5) DGX chip. But the math seems to work out nicely if each chip just runs one agent, so maybe not?

Tokens per second Link to heading

Lastly, we know that they produced a total of 300 billion tokens over 88h. That’s a total of roughly 1 million tokens per second4, or 100 tokens/sec per agent. Astra runs at 50 tps, so this seems a bit on the high end. But we also know from the announcement that the internal model is cheaper to run, so maybe this new model is 2x smaller than Astra?

Either way, this means we get 66 amortized tps per DGX B200.

A DGX B200 has 144 petaFLOPS of (4bit?) inference, so each token costs around 144/66=2.18petaFLOP. That is \(2 \cdot 10^{15}\) floating-point operations for every. single. output token! (Well, as an upper bound. But let’s assume we are mostly saturating the GPU.)

Extrapolating Link to heading

Their “let’s have some fun on the side” cluster used 230 MW during this time window. Let’s assume that this is also their training cluster, and that this is 50% of their total capacity, with the other half serving inference? (10k average active parallel users sounds a bit low maybe? Especially if users are also running many agents? But I guess they mostly don’t use Astra, and Sol is 2x cheaper and Luna 50x, so maybe this works out? Either way it sounds reasonable to have a roughly 50/50 split between training and inference hardware? But maybe they have 10x more inference than training?)

Then, they have a total capacity of roughly 460 MW. Typical power plants are on the order of 1 GW, so this means that OpenAI is burning through roughly half a power plant of energy now.

If we extrapolate the coal usage to a year, we get 4000 GWh of energy used, which corresponds to 0.5% of the US coal consumption for power generation. or 2*3000 tonne / 88h * 1year = 600000 tonne/year. (It turns out that 88h is very close to 1% of a year.) So that’s 300 Olympic swimming pools of coal, or 7.5 bulk carriers of coal.

Putting things in perspective Link to heading

NeurIPS 2025 was in San Diego and had 29000 attendees. Let’s say that on average, each attendee does a New-York to San Diego return flight (with many US attendees having shorter flights, and many non-US attendees having longer flights). The distance is roughly 4000 km, and the fuel consumption per passenger-kilometer of a plane is around 3L per 100km, making for a round-trip usage of 240L or 200kg per person. In total, we get an estimated 5800 tonnes of jet fuel for NeurIPS. Given that jet fuel is nearly twice as energy dense as coal, this burns around 4x more energy than the Navier-Stokes proof.

Conclusion Link to heading

To summarize: assuming that $1 million was spent on electricity, we get the follow Fermi estimates:

  • A total energy usage of 20 GWh over 88 hours, or 230 MW of power.
    • That’s 1.5 Olympic swimming pools of coal, or 5% of a bulk carrier.
    • That’s 25% of what 29000 people burn with their flights to/from NeurIPS.
  • They used the equivalent of 15'000 DGX B200 video cards, or 120'000 B200 cards.
  • Given 10000 agents, that’s 12 B200’s per agent, suggesting a 2TB model.
  • Given 300G output tokens, that’s 100 tps per agent, or 66 tps per DGX B200.

Extrapolating to yearly consumption, assuming a 50/50 split of their GPUs between training and inference, we get:

  • 4000 GWh per year, 0.5% of all US generated coal power;
  • 300 Olympic swimming pools of coal, roughly one per day5;
  • 7.5 bulk carriers of coal per year.

  1. Where by “they”, I mean Alfred Renyi. ↩︎

  2. There’s a whole bunch of dispute around who was first, OpenAI or independent researchers Buckmaster and Alpöge. See wikipedia↩︎

  3. Their solution still requires a smooth external force and thus (as I read somewhere?) does not fully solve the original problem? ↩︎

  4. 0.95M tps to be precise. Honestly, this is close enough that I could imagine they choose the number of agents and GPUs so that they have around 1M tps in total. ↩︎

  5. I.e., the US burns 200 swimming pools of coal per day? Ouch… ↩︎