
Jay Kapoor, a general partner at VSC Ventures, ran the numbers behind Dylan Patel's claim that "anyone can make money off compute today" — and concluded the claim is "only directionally correct." His argument: the business case for becoming your own inference provider reduces to four numbers, and the two hardest of them are the ones the claim quietly assumes.
Patel's line comes from his Dwarkesh Podcast appearance, where it is a premise rather than a business plan. Arguing that OpenAI and Anthropic would need to bid compute up to $25–50 million per megawatt to control 100 GW by 2028, Patel said: "anyone can make money off of $10-15 million per megawatt compute today. I kid you not, it's not that hard. Go get a GB300 rack, go download the Kimi weights, go download vLLM or SGLang, set it up… Go put it on OpenRouter. It's very simple. You'll start generating more revenue than you're paying for the compute." Cheap compute being trivially profitable is, in his telling, exactly why compute cannot stay cheap.
Kapoor's post sets out the inputs: selling price (Kimi K3 at $15 per 1M output tokens), compute cost (CoreWeave's $69/hour on-demand 8x B200 node, or $8.60 per B200 GPU-hour), serving efficiency (output tokens per second from those eight GPUs), and utilization (the share of capacity customers actually pay for). His equations:
He stops at the question rather than the answer — "At $15/M output tokens and $69/hour of compute, you need what % utilization to break even?" — and argues the implicit assumptions of high utilization and high serving efficiency are precisely where labs have the advantage: "They have massive demand and sophisticated infrastructure. You... do not."
The prices are public, so the threshold is arithmetic. CoreWeave's published rate card puts an HGX B200 node at $68.80/hour on-demand ($8.60/GPU-hour) and $34.11/hour spot. At $15 per 1M output tokens, covering the on-demand node requires 4.59M output tokens per hour — about 1,274 tokens per second sustained across the node, or roughly 159 tokens per GPU-second, all of them sold and paid for. On spot capacity that halves to about 79 tokens per GPU-second.
For scale, vLLM's Kimi K3 launch benchmarks describe a Pareto frontier on GB300 NVL72 hardware running "from high-throughput serving at 2K+ TPGS to low-latency serving at 100+ TPS/user," with 118 tokens/second per user at batch size 1 and 370 with DSpark speculative decoding. Aggregate throughput at maximum batch is far above the break-even line; per-user interactive decoding is not. Which regime you serve in — and how much of it is billed — is the whole argument.

vLLM's published throughput/latency frontier for Kimi K3. Credit: vLLM.
The 8-GPU node in Kapoor's example cannot hold the model. vLLM states that Kimi K3 — 2.8 trillion parameters, 16 of 896 experts active, MXFP4 weights — "can barely fit in a single NVIDIA DGX B300 and requires a minimum of 16 NVIDIA B200/GB200 GPUs to serve on that hardware generation." That is two CoreWeave nodes, $137.60/hour, and it is why Patel specified a GB300 rack rather than eight B200s.
The selling price holds up better than Kapoor suggests. Of the five providers llm-stats tracks for Kimi K3, four — Moonshot, Fireworks, Novita and Together — list $15 per 1M output tokens; DeepInfra undercuts at $14.25. The margin pressure is real but so far marginal. The licence is a separate obstacle: model-as-a-service businesses earning more than $20M over 12 consecutive months need a separate agreement with Moonshot.
Both men are describing the same market from opposite ends. Patel's point is that cheap compute plus open weights makes a positive spread easy enough that compute prices must rise; Kapoor's is that the spread is not a business, because filling a node with paying demand at a competitive price is the hard part — and CoreWeave's on-demand rate card is close to theoretical anyway, with the company reported to be largely sold out of 2026 capacity and selling mostly on multi-year contracts. Neither position requires the other to be wrong. The number that decides both is utilization, and nobody publishes it.
Jay Kapoor's postthe clip he replied toDwarkesh Podcast: Dylan PatelvLLM on serving Kimi K3CoreWeave GPU pricing 2026llm-stats: Kimi K3

Moonshot AI posts Kimi-K3 weights on Hugging Face

Kimi K3

Dylan Patel: $11T of AI capex through 2029, $5T of it borrowed
An 11x price spread for the same open-weight model on OpenRouter

Dylan Patel: two labs will own most of the world's compute

How rogue inference providers can game router scoring by under-reporting cache hits