Discussion about this post

User's avatar
Melvin's avatar

The mistaken assumption in this analysis is that all compute is created equal. The rental prices and use cases of the same hardware can vary wildly depending on how it is configured. Fragmented compute (i.e., clusters of various sizes, generally in the hundreds to low thousands of GPUs) is a commodity that is rented at market rates (these are the H100s that rent for ~$2.50/GPU-hour you mentioned). Frontier-scale coherent clusters (gigawatt-blocks of tens to hundreds of thousands of GPUs) on the other hand are a different product entirely, and are scarce. Finding unassigned frontier-scale coherent clusters available for rent before the end of 2026 is almost impossible. This is why Anthropic signed a contract to pay xAI ~$1.25B per month for access to Colossus 1 (an implied ~$7-8/GPU-hour, roughly 3x commodity rates). These frontier-scale coherent clusters are the crown jewels Meta is reportedly considering renting out.

But if Meta is selling access to its coherent clusters, doesn't that mean Meta is admitting it overbuilt? Not quite. The benefit of the frontier-scale coherent cluster is also one of its biggest drawbacks: it excels at training frontier LLMs, but it's overkill for most other tasks, and training is only one part of the model lifecycle. Outside of AI R&D, Meta's biggest need for compute comes from its recommendation systems (recsys). The compute requirements for recsys are massive, but the workload will happily run on older GPUs and fragmented clusters. In other words, commodity compute works just fine for Meta's workloads outside of LLM training. Think of it like owning a supercar and a Prius. You absolutely need the supercar if you want to race, but a Prius is a much more sensible daily driver.

So if you're Meta, and you're in between training runs, would you rather run recsys on your coherent cluster (the supercar), or run it on fragmented clusters (the Prius) at commodity rates so you can rent out the coherent cluster at 3x market rates? Obviously the latter. But what if your partner is using your Prius? Then you rent another one. This is where the CoreWeave and Nebius deals come in: Meta can offload recsys workloads onto rented commodity capacity and free up the coherent blocks to rent out at a premium.

Given all of this, I don’t think it likely that Meta slows down their capex. After all, what's better than having a supercar? Having two supercars. Or having a supercar, a moped for your kid to do DoorDash in, and a Prius for yourself. Or buying modifications to make your supercar even faster. The point is, you have a lot more optionality now that you're making extra money renting your supercar out to finance influencers on the weekends (I've heard that's a great business in Miami).

Rick Larkin's avatar

Great article. Don’t think it is fair to say Grantham is a permabear. He’s been long more than he’s been short but he did point out how implausible the everlasting AI boom story was.

19 more comments...

No posts

Ready for more?