In the PJM interconnection, the grid serving thirteen mid-Atlantic states and the densest concentration of data centers on the planet, the 2024/25 capacity auction cleared at $28.92 per megawatt-day. The 2026/27 auction, run at the same mechanism two cycles later, cleared at $329.17. That is an increase of more than eleven times, and PJM's own independent market monitor attributes 63 percent of the single-year jump between those cycles, roughly $9.3 billion, to one class of customer: data centers.1 In the same twenty months, on the other side of the Pacific, the marginal cost of a comparable inference call fell by roughly two orders of magnitude. Both of those are true at once, they are not unrelated, and neither is the headline. The headline is that American developers noticed the second number faster than American utilities priced in the first one — and rerouted.
Part 3 of Hormuz and the Weighting Game. Part 2 traced who wins the open-weight race DeepSeek forced into existence and left one question hanging: why does absorbing a model convert into shipped advantage faster on one side of the Pacific than the other. The answer is electricity, and this part follows it all the way down: from the auction clearing price in Pennsylvania to the overcapacity that makes a Chinese kilowatt-hour cheaper than an American one, and from there to the inference-routing decisions that quietly move American developer traffic onto Chinese-owned infrastructure because it is the rational thing to do this quarter. Part 4 follows what travels inside those routed requests. This part is about why the requests go there at all.
/ 01The Toll, By the Numbers
Start with a tally, not a verdict: a set of dated, sourced figures on what an inference token costs, on each side of the Pacific, and how fast that cost is moving. Some of these numbers are as clean as a government auction result. Others carry real uncertainty about attribution, flagged where it exists. None of them is asked to prove more than it can.
| Figure | Detail | Window |
|---|---|---|
| OpenRouter, US-origin model token share | OpenAI + Anthropic + Google combined, on a major US developer inference marketplace | 70% → 30% |
| OpenRouter, Chinese-origin model token share | DeepSeek, Tencent, Xiaomi, MiniMax, Qwen combined — overtook the three big US labs' 35.7% | 46.4% (Jun 2026) |
| DeepSeek V4-Pro API pricing | Per million tokens, peak/off-peak, vs. Claude Sonnet 4.6 ($3/$15) and GPT-5.5 ($5/$30) | $3.96 / $1.98 |
| PJM capacity price | Per megawatt-day, 2024/25 to 2026/27 auction; 63% of the increase attributed to data centers | $28.92 → $329.17 |
| US data-center power draw | Share of total US peak summer demand rose from 4.1% to 5.3% in the same year | 31 GW → 41 GW |
| Virginia data-center electricity share | One state, one grid, one load class — the densest data-center corridor in the country | <5% → ~40% |
| China coal consumption vs. any single country's total energy use | China's coal alone (~55 EJ) exceeds the entire US total energy consumption (~47 EJ) | 2025 |
Two of these figures are the same statistic wearing different clothes, and it matters to say so plainly rather than let a Part 4 or Part 6 Tally double-count them later: the $9.3 billion figure is the 63 percent figure, expressed in dollars instead of share. Everything else in the table is a separate measurement. Read top to bottom, the table says one thing: American inference is getting structurally more expensive at almost exactly the moment Chinese inference is getting structurally cheaper, and the mechanism connecting the two halves of that sentence is not mysterious. It is a grid.
/ 02Where the Tokens Actually Go
OpenRouter is not a curiosity. It is a routing marketplace used heavily by independent developers and startups: the exact population of American AI builders who are price-sensitive, latency-tolerant, and free to point their traffic at whichever model clears the bar for the least money. Its own published telemetry is platform data, not a survey: in June 2025, the three large American labs held roughly seventy percent of token volume on the platform. One year later, that had fallen to thirty percent, and Chinese-origin models, DeepSeek chief among them, had risen to 46.4 percent, overtaking the American three combined. DeepSeek alone doubled its own share in the first half of 2026, from nine to eighteen percent, with a single model, V4-Flash, the highest-volume model on the entire platform by tokens processed.2 OpenRouter's own reporting is careful to add a caveat worth carrying forward honestly: revenue share tells a different story than token share, because paying enterprise customers still skew toward the US frontier labs even as raw volume moves. The shift is concentrated in exactly the workloads where price is the deciding variable (coding assistants, agentic loops, bulk generation), not in the premium tier. That distinction matters, and it also does not change the direction of travel for the workloads it does describe.
The price gap explaining that shift is not subtle. DeepSeek's R1, at launch in January 2025, priced input tokens at roughly a twenty-seven-times discount to OpenAI's comparable o1 reasoning model.4 Eighteen months on, even after DeepSeek moved to peak/off-peak pricing that raised its headline rate, V4-Pro at peak still undercuts Claude Sonnet 4.6 and GPT-5.5 on both input and output, and the budget-tier V4-Flash runs on the order of ninety to a hundred times cheaper than frontier US models on a standardized workload comparison.2 Trade coverage puts a concrete number on what that means for a builder: a workload costing $100,000 to $150,000 a month on GPT-4o can run for one to three thousand dollars a month on DeepSeek's API (an order-of-magnitude claim from a vendor-adjacent source, not independently audited, but consistent in direction with the platform-level token-share shift, which is independently measured).5 A developer choosing the cheaper API is not making a geopolitical statement. They are doing arithmetic. That is precisely why the arithmetic itself is the story: nobody had to persuade an American startup to route its traffic through Chinese-hosted infrastructure. The price did it.
One complication belongs here rather than in a footnote, because it splits the routed traffic into two genuinely different streams. Not every American workload running a Chinese model runs it on Chinese infrastructure. Microsoft added DeepSeek to its Azure AI Foundry catalog in early 2025 and DeepSeek V4 in May 2026; Amazon hosts R1 on Bedrock. An enterprise that wants a Chinese open model with SOC2 guarantees, EU data residency, and a contractual SLA typically rents it through exactly those wrappers, which means a US hyperscaler collects the serving margin and the data stays inside American compliance perimeters. The Stack, Part 3, called this the relocated toll booth, and it cuts against the simplest version of this part's claim: for compliance-bound enterprise traffic, Chinese weights are often an American hosting business. The OpenRouter shift this section documents is the other stream, the price-sensitive developer traffic that routes to the cheapest endpoint directly, and it is that stream, not the Foundry-wrapped one, that lands on Chinese-hosted infrastructure with everything Part 4 says travels along with it.19
Every request routed to a Chinese-hosted API also exposes something besides money: prompts, feedback signals, red-team probes, whatever a developer's traffic happens to contain, retained under a privacy policy that grants broad rights to use de-identified content to “develop and improve” the service, on servers reachable under China's 2017 National Intelligence Law.6 That exposure is real, it is why the Pentagon, NASA, the Navy, and the Commerce Department all blocked device-level DeepSeek access within weeks of the R1 shock, and it is Part 4's subject, not this one's. This part is only the toll booth. What crosses it is next.
/ 03The Grid Doesn't Care Whose Model It Is
The American half of this price gap is not a story about American AI labs pricing badly. It is a story about American electricity getting more expensive at the exact rate American AI demand is scaling, and the two curves crossing in public, auditable auction data. Goldman Sachs projects US data-center power draw rising from 31 gigawatts in 2025 to 41 in 2026, on its way to 66 by 2027, a share of total US peak demand climbing from 4.1 to 5.3 to a projected 8.5 percent in three years.7 In Virginia specifically, home to the densest data-center corridor in the country, data centers now draw something close to 40 percent of the state's total electricity, against under 5 percent as recently as 2010.8 None of that demand growth is hypothetical or model-specific. It does not matter whether the training run behind it happened in San Francisco or Hangzhou. What matters is that the grid supplying it is American, permitted under American process, and priced by an American capacity auction that has just posted an eleven-fold increase in two cycles, with independent market analysis pinning the majority of the latest jump specifically on data-center load.1 A watt in PJM territory is worth roughly the same, physically, as a watt anywhere else. It is not priced the same, and the price is the entire mechanism this part is about.
That export curve is the other half of the mechanism, and it did not happen by accident either. China built battery, solar, and EV manufacturing capacity for a growth rate its own domestic market stopped needing years ago, and the resulting overcapacity did not sit idle. It went abroad, at prices that reflect genuinely lower input costs as much as any subsidy. The same overbuild logic that put a battery pack on a container ship to Rotterdam thirty percent below the American price also sits underneath the compute those batteries increasingly feed: cheap, overbuilt Chinese power is not only an export good. It is the input cost structure a Chinese inference provider gets to run on.
/ 04The Other Side of the Meter
Put two numbers next to the PJM auction result and the shape of the gap becomes hard to unsee. First, China's own manufacturing capacity runs roughly two to three times its actual production across every major clean-energy category: battery cells at 2,830 gigawatt-hours of capacity against 890 produced, solar modules at 1,045 gigawatts against 587, wind nacelles at 204 against 96, electric cars at 22 million units of capacity against 12 million built.9 That is not a cyclical inventory overhang correcting itself next quarter. It is a decade of low-cost production capacity, already funded and already built, sitting there whether or not the domestic market wants it: the physical precondition for the export pricing that undercuts American and European competitors on batteries and solar by roughly thirty to thirty-five percent.10
Second, and further upstream: China's coal consumption alone, not its total energy use, just the coal, exceeds the total energy consumption of any other single country on earth, including the United States.11 That is a staggering base of cheap, dispatchable domestic power, and it sits alongside the wind, solar, and battery capacity above, not instead of it. The specific technical explanation for DeepSeek's own serving costs (a mixture-of-experts architecture that activates only a fraction of parameters per token, plus aggressive training-efficiency engineering) is real and independently plausible on its own terms.12 But no cost-accounting teardown has isolated that architectural saving from the power bill underneath it, and the power bill is not a mystery. A country running on the cheapest large-scale electricity on the planet does not need to subsidize inference pricing to make it look cheap. The floor is just lower. The chip side of the ledger shows the same logic running in silicon: Huawei's Ascend 910C loses to Nvidia's best on a per-chip basis by a factor of two to three, and it takes roughly seven racks of 910Cs to match one rack of Nvidia B300s, so Huawei competes at the cluster level instead, deliberately trading efficiency for scale because the electricity feeding the extra racks is cheap enough to make the trade rational. That is not a workaround. It is the energy floor expressing itself as chip strategy.20
The floor
Cheap coal at scale, plus solar/wind/battery manufacturing capacity built two to three times larger than domestic demand requires, sets a power-cost floor no permitting-constrained American grid can currently match.
The compute
Inference is a power-hungry, repeatable process: the cheaper the electricity underneath a data center, the lower the marginal cost of every additional token it serves, independent of any subsidy on the API price itself.
The API
DeepSeek and its peers price accordingly: a real architectural efficiency (mixture-of-experts) stacked on top of a real power-cost advantage, not a subsidy story standing in for engineering.
The routing
Price-sensitive American developers on OpenRouter and similar marketplaces reroute toward the cheaper option, at scale, without anyone framing it as a strategic decision. It's just the lower number in the invoice.
The strain
Meanwhile the American grid absorbs data-center load it was not built for, and the PJM auction prices that strain back onto every ratepayer in the region, making the American side of the same trade more expensive with each cycle.
/ 05The Honest Reckoning: China Is Unwinding This, Not Deepening It
A fair accounting has to sit with an inconvenient fact before it draws its conclusion, and the fact is this: as of 2026, Beijing is dismantling, not building out, the subsidy architecture that produced the export pricing above. VAT export rebates on batteries were cut from nine to six percent between April and December 2026, and are scheduled to disappear entirely on January 1, 2027; solar-product export rebates were eliminated in April 2026; and the 15th Five-Year Plan, released in late 2025, dropped “new energy vehicles” from China's list of strategic industries for the first time in more than a decade, with officials citing “excessive reliance on low-margin exports.”13 If underpricing were a coordinated dependency wedge, this is a strange moment to start dismantling the machinery that supposedly built it. China's own on-record position, stated by Commerce Minister Wang Wentao and Vice Finance Minister Liao Min, is that Western “overcapacity” complaints are “groundless” and that the price gap reflects innovation and complete supply chains, not policy design: a position that deserves to be reported, not waved past.14
The source underneath this entire part does not argue a clean one-way story either. Cembalest's own note frames the choice explicitly as a Polyphemus problem: the one-eyed giant of the Odyssey, who could only look at what was directly in front of him. Read only the energy column and Chinese exports look like an unambiguous gift to every buyer, American developers included. Read only the trade and deindustrialization column and the same exports look like an attack on the industries that can't compete with them. “Both columns are correct. Neither is the whole.” This series is not simply reusing that frame. It is pushing past it, arguing that on the specific energy/inference axis Cembalest's note actually measures, the one-way reading survives its own steelman better than it does on, say, the talent axis. That is a real extension of the source's argument, not a restatement of it, and it should be read as one.
So why does the net direction still favor one-way, once the steelman is fully granted? Because the subsidy-phaseout and the cost-floor argument answer two different questions, and only one of them is load-bearing for this series. Whether Chinese underpricing was designed as a wedge is a question about intent, and the honest answer, per this steelman, is probably not, or at least not primarily. Whether the resulting price gap functions as a one-way transfer of American inference demand regardless of intent is a separate, structural question, and removing a rebate on top of a genuine, structural power-cost floor does not remove the floor. VAT rebates shaved a few points off an export price already built on electricity roughly a fifth to a third the American cost. The subsidies were the icing. The floor was always the cake, and nothing in the 2026 rebate phaseout touches it.
/ 06Sensitivity Is Not Vulnerability
This is the place in the series where the rival hypothesis (that all of this is just a symmetric prisoner's dilemma, both sides feeling the same squeeze) gets its clearest test, because energy dependence is close to a textbook case for the exact distinction Robert Keohane and Joseph Nye built their theory of complex interdependence around. Sensitivity is how quickly and how severely a change on one side produces costly effects on the other, holding existing policy constant. Vulnerability is the cost of adjusting to that effect once a state actually tries to change its own policy in response: the cost of its best available alternative.15 Their central claim is that vulnerability, not sensitivity, is where the real power sits: a state that can pivot cheaply is not vulnerable even if the initial shock is sharp; a state with no good alternative is vulnerable even if the shock looks modest.
Applied here, both countries are plainly sensitive to the same energy-and-inference dynamic: American developers feel the price gap the moment they check an invoice, and China's own exporters feel American tariff and anti-subsidy actions the moment a shipment is delayed at a port. That symmetry is real and belongs in the piece. But sensitivity-symmetry is not vulnerability-symmetry, and the PJM auction result is exactly the entry that shows the difference. China's adjustment path away from its current overcapacity, should it choose one, is a policy dial its own five-year-plan apparatus already turned once, on schedule, in 2026 (a state with a multi-decade planning horizon deciding, on its own timeline, to phase out a rebate). America's adjustment path away from grid strain runs through permitting reform, transmission buildout, and a capacity auction that reprices in real time and takes years to reflect new generation: a process no single administration, operating on an electoral cycle Robert Axelrod's own framework (see Part 2) predicts will keep interrupting it, currently controls end to end. The physical bottlenecks underneath that adjustment path are already documented in this platform's own prior reporting: the United States imports roughly 90 percent of its large power transformers, with lead times stretched toward four years for the biggest units, and gas turbines are sold out through 2030 or 2031, with GE Vernova's backlog at 100 gigawatts and prices rising 10 to 20 points per kilowatt in a single half-year. The Stack, Part 1, walked through those numbers in full; they are the dollars-and-months content of the word "vulnerability" as this section uses it.18 Both countries feel the exchange. Only one of them can absorb the shock of it cheaply. That is the technical, Keohane-and-Nye-precise version of “one-way” this series is actually arguing: not that China feels nothing, but that its felt cost converts to a policy lever it can pull on its own schedule, while America's converts to a multi-year infrastructure problem it cannot.
Structural realism supplies the harder question underneath that distinction: even granting every dollar of it is mutually beneficial in absolute terms (cheaper compute for American startups, real export revenue for Chinese manufacturers), Kenneth Waltz's and John Mearsheimer's relative-gains logic says a rational Washington should still worry about a transfer that measurably strengthens a rival's relative position, independent of whether it also helps the buyer.16 Joseph Grieco's formalization of this concern is precise: a state will limit even a Pareto-improving exchange if it expects the other side to capture a disproportionate share of the relative gain.17 The DeepSeek-shock reaction described in Part 2 (export-control tightening, the AI Action Plan, a joint congressional investigation into US firms' Chinese-model use) is exactly what a relative-gains alarm looks like firing late, years after an absolute-gains-era engagement policy had already let the underlying cost structure tilt. The energy axis this part documents is the clearest place that tilt shows up in hard, auditable numbers, because unlike a researcher's departure or a checkpoint's publication, a grid auction price cannot be spun.
/ 07The Invoice, Read Correctly
Put the numbers back together. PJM's capacity price rose more than eleven-fold in two auction cycles, with data centers named as the majority driver of the latest jump. American data-center power draw grew by a third in a single year and is on pace to nearly double again by 2027. Virginia now runs its data-center corridor on a share of the grid that would have been unthinkable in 2010. On the other side of the Pacific, a coal-and-overcapacity power floor several times cheaper than the American rate underwrites an inference price that fell by roughly two orders of magnitude in twenty months. American developers, doing nothing more strategic than reading an invoice, moved forty points of token share onto it. The steelman is real: China is dismantling its own export subsidies in 2026, not deepening them, and Cembalest's own source material insists on reading both the gift and the injury in the same column. But the subsidy was never the floor. The floor is a grid, and grids don't unwind on anyone's five-year plan. Part 4 follows what actually travels inside a routed request — not the toll, but the cargo.
“The chip embargo had removed the hardware. But the hardware had already been mapped by a mind that no customs officer could detain.”