In June 2026, Microsoft did something none of the first four layers of this stack ever do in public: it looked at its own internal bill for a tool it didn't build, decided the number didn't work, and pulled the plug. It canceled the majority of its internal Claude Code licenses, moved thousands of engineers to its own GitHub Copilot CLI, and left behind a tidy piece of tape-reader's evidence. The tool hadn't gotten worse. The bill had gotten honest: a $200-a-month subscription seat, run the way an engineer actually runs an agentic coding session, can consume compute that costs as much as $8,000 a month to serve at metered rates — a gap SemiAnalysis sizes at roughly forty times the sticker price.12 Parts 1 through 4 asked who holds the shell, the silicon, the rented cloud, the frontier model itself. Here the question turns personal: who eats the token when the seat costs a fraction of what it actually runs through the meter? And for the first time in six layers, the name for this rung of the stack — “the harness” — wasn't coined by a critic. It's the word J.P. Morgan's own house economist uses for it, in the same report where he explains why it's the industry's escape hatch.
Part 5 of The Stack, one rung up from the frontier models. Same two questions, asked of the orchestration layer: who books the equity, and who eats the subsidized loss? The frontier labs sell intelligence by the token; the harness — Claude Code, OpenCode, Cursor, Devin, the RL-as-a-service shops — is the software that decides, request by request, whether that intelligence has to come from the most expensive model available. That decision is worth real money, and 2026 is the year both sides of the transaction started fighting over who pays for the token in between.
/ 01The Two-Hundred-Dollar Seat That Costs Eight Thousand
Start with the product itself, because the arbitrage is built into its pricing tiers on purpose. Claude Code — Anthropic's own coding agent, launched publicly in May 2025 — sells three consumer-facing plans: Pro at $20 a month, Max 5x at $100, and Max 20x at $200, each buying a multiple of Pro's usage ceiling rather than a fixed token allowance.3 Enterprise self-serve is $20 a seat — and once an account grows past the Team plan's 150-seat cap, that $20 buys access, not tokens; usage past that point is fully metered at API rates, with no included allowance at all.1 The flat fee is a door. What's behind it is priced by the token.
SemiAnalysis ran the experiment the plans are daring someone to run: buy the subscription, use it flat-out to its stated ceiling, and price every token consumed at Anthropic's own posted API rates. Applied to a Claude Max 20x seat, the metered-equivalent bill comes to as much as $8,000 a month — about forty times the $200 charged for it.12 SemiAnalysis ran the same methodology against OpenAI's $200-a-month ChatGPT Pro plan and found a metered-equivalent bill of up to $14,000 for a user who fully exhausts its ceiling — a roughly seventy-fold gap, though the two figures aren't perfectly comparable across different models and task mixes.4
Who ate that gap before anyone had to talk about it? The vendor, blended across a much larger base of light users. Anthropic's overall gross margin runs around 44% — about $0.71 of compute cost against every dollar of revenue in Q1 2026, projected to improve to $0.56 in Q2.16 That's a perfectly ordinary SaaS pattern: heavy users cost more than they pay, light users pay more than they cost, and the blend works as long as the heavy tail stays a tail. Two things happened in the same quarter that suggested it might not be staying a tail.
First, Microsoft. It canceled the majority of internal Claude Code licenses across its Experiences + Devices division, effective by June 30, 2026, after the tool's token-metered cost “ran past the annual AI budget months ahead of schedule.”5 One contemporaneous account frames the logic in a single sentence worth keeping: “Cost pulled the trigger. What cost exposed is the real lesson: renting intelligence by the token is a price Microsoft does not control.”7 The same reporting cluster claimed Uber burned through its entire 2026 AI-tools budget on Claude Code and Cursor in just four months — a vivid number, but one that traces to a single secondary aggregator with no primary corroboration located; treat it as a flare, not a fact, until Uber or a first-tier outlet confirms it.5
Second, GitHub. On June 1, 2026 — the same week — GitHub Copilot switched every plan from opaque “premium request” credits to literal, token-metered billing. Headline seat prices didn't move ($10 Pro / $39 Pro+ / $19 Business), but the meter underneath them was exposed to daylight for the first time, and users reported real bills jumping from $39 to over $800 a month, with some reporting increases as high as 100x.8910 Nothing about the underlying compute cost changed on June 1. What changed is that Microsoft stopped absorbing the gap between flat pricing and actual agentic-workflow consumption — on its own product, in the same month it stopped absorbing that gap on someone else's.
The two SemiAnalysis usage-ceiling figures — $200-vs-$8,000 for Claude and $200-vs-$14,000 for ChatGPT Pro — are the single most load-bearing numbers in this section, and neither was read from a directly fetched SemiAnalysis publication in this research pass. The Claude figure comes via J.P. Morgan's reproduction of SemiAnalysis's work; the ChatGPT Pro figure comes via secondary aggregation of the same underlying SemiAnalysis research. Both are consistent with the pricing tiers and the Microsoft/GitHub repricing that followed — the direction and rough scale check out — but a direct pull from SemiAnalysis's own newsletter archive should replace this sourcing before either number is treated as filed fact rather than well-corroborated estimate.
/ 02The Businesses
Five kinds of company sit at this layer, and only one of them is profitable in any sense a public filing would recognize.
Claude Code is Anthropic's own harness — the reference implementation of “orchestration running a closed frontier model” — and by its own account it's working: over $2.5 billion in run-rate revenue as of February 2026, more than doubling since the start of the year, business subscriptions quadrupling since January, and enterprise customers now more than half of Claude Code's revenue.11 OpenCode is its open-source mirror image: the leading model-agnostic harness, plugging into 75-plus model providers — Claude, GPT, Gemini, open weights, local models — with no single vendor to toll it. As of late June 2026 it led Claude Code on GitHub stars (roughly 180,500 to 135,000) despite launching about two months later, with 900-plus contributors and a claimed 7.5 million monthly developers.12 It has no disclosed revenue, because it isn't really a company in the conventional sense — its business model is closer to “commoditize the harness so no one can toll it” than to venture-scale SaaS, which makes it this layer's cleanest structural counter-example to the idea that every good harness eventually gets captured. There's simply no equity to acquire.
Cursor, built by Anysphere, is the layer's fastest-moving equity story, full stop. Watch the valuation staircase:
$9.9B — Series C, early 2025
A well-funded developer tool, not yet a phenomenon.
$29.3B — Series D, November 2025
Post-money, at roughly $1B ARR. Three times the Series C valuation in under a year.
$50B pre-money — in talks, April 17, 2026
ARR had reached $2B — doubled in five months — and the valuation rose roughly 70% to match.13
$60B all-stock — SpaceX acquisition, announced June 16, 2026
Reported ARR by then: roughly $4 billion, up from about $100 million in early 2025.14
Cursor is also, notably, gross-margin-positive at the enterprise-account level as of April 2026 — individual-developer accounts are still loss-making — driven substantially by its own proprietary Composer model handling roughly half of autocomplete requests and routing the rest to cheaper third-party models like Kimi instead of defaulting to frontier tokens.15 That is worth sitting with: the harness business closest to profitable got there specifically by reducing its dependence on the most expensive frontier inference — the strongest evidence in this entire layer that the escape-hatch thesis isn't just a banker's turn of phrase.
Cognition (Devin) sells autonomous, ticket-to-pull-request coding at $20/month plus $2.25 per unit of agent-compute (roughly fifteen minutes of autonomous work), or $500/month with 250 units included.16 Its path here runs through one of the stranger M&A chains in the whole series: OpenAI agreed to buy the IDE maker Windsurf for $3 billion in May 2025; the deal's exclusivity lapsed that July over an IP conflict with Microsoft; Google then paid $2.4 billion for a non-exclusive license to Windsurf's technology and hired away its CEO and top researchers — a “reverse acquihire,” the same structure used on Character.AI and Inflection; and Cognition closed a deal for Windsurf's remaining IP, brand, and roughly 210 employees within 72 hours of Google's move.17 Cognition then raised $400 million at a $10.2 billion valuation two months later, and over $1 billion at $25 billion pre-money in late May 2026, reaching $492 million in annualized run-rate revenue — with Devin reportedly writing 89% of Cognition's own code.1819 No margin figure has been disclosed alongside any of it.
LangChain sits one level up the stack from the coding-specific harnesses — infrastructure for building any agentic application, not a coding assistant itself. It raised $125 million at a $1.25 billion valuation in October 2025, and reports 90 million monthly downloads with adoption across 35% of the Fortune 500.20 Its revenue is the thin spot in an otherwise well-documented layer: a single third-party aggregator put 2025 ARR at roughly $16 million — unverified, company-undisclosed, and if directionally right, a revenue-to-valuation ratio near 78x, an outlier even by 2026 AI-infrastructure standards.20
And then there's the sharpest form of the thesis: RL-as-a-service. Ramp used Prime Intellect's open-source prime-rl stack to post-train a 35-billion-parameter open model that beat Claude Opus 4.6 by several points on Ramp's own retrieval benchmark, running at Haiku-level cost and latency.21 Harvey had Trajectory Labs post-train Nvidia's open-weight Nemotron 3 Ultra on Harvey's own Legal Agent Bench in under 24 hours, reaching “the same band as leading closed models” on legal work at a fraction of frontier token cost — a claim that comes from Trajectory Labs and Harvey themselves and hasn't been independently benchmarked by a third party.22 J.P. Morgan frames the same category of result more starkly still: Opus 4.6 costs $3,700 to run a standard benchmark task set for a score of 56; DeepSeek V4 Pro scores 44 for $186 — roughly twenty times cheaper.1 This isn't routing to a cheaper model mid-request. It's a customer training its own model so it never has to rent the frontier lab's intelligence for that task again.
| Business | Latest disclosed figure | Profitability signal |
|---|---|---|
| Claude Code (Anthropic) | $2.5B+ run-rate, Feb 202611 | Parent's ~44% blended gross margin; no operating profit disclosed1 |
| OpenCode | 7.5M monthly devs (claimed)12 | No revenue disclosed — not a conventional company |
| Cursor (Anysphere) | ~$4B ARR, Jun 202614 | Gross-margin positive, enterprise accounts only15 |
| Cognition (Devin) | $492M ARR run-rate18 | No margin figure disclosed |
| LangChain | ~$16M ARR (single estimate)20 | Too thinly sourced to assess |
None of this is a shell game. Developers using these tools ship measurably faster; Cursor's enterprise gross margin and Ramp's benchmark win aren't marketing claims, they're the kind of number a finance team signs off on before renewing a contract. The harness genuinely does what Cembalest says it does — it lets a well-built orchestration layer extract most of a frontier model's value from a much cheaper one underneath it. The bear case in this section isn't that the product doesn't work. It's that almost none of the businesses selling it have found a durable price for what it costs them to deliver — and the one that has, Cursor, got there by routing away from the frontier tokens the rest of this stack is built to sell.
/ 03The Escape Hatch, in the Banker's Own Words
Here is the sentence that gives this layer its name, and it did not come from a critic of the AI boom. It came from J.P. Morgan's own house economist, describing exactly the mechanism this section has been building toward:
“Agent harnesses (external systems such as memory, tools, safety boundaries and orchestration layers in Claude Code or OpenCode) that can run open models often improves their output quality and reduces the need to rely on the most expensive frontier models.” — Michael Cembalest, “Semiquincententacles,” J.P. Morgan Eye on the Market, June 2026, p.15
Read that sentence again in the context of everything Parts 2 through 4 documented: Nvidia's eroding accelerator share, the frontier labs' circular profit, the $2 trillion of cloud backlog resting on counterparties whose revenue doesn't exist yet. The word “harness” is Cembalest's own coinage for the layer that lets a customer walk away from all of it — and he's not speculating. The migration is dated, documented, and happening in production, not in a slide deck.
Lindy AI, a 25-person startup, moved its entire inference stack off Claude and onto DeepSeek V4 Flash, cutting inference cost on the migrated routes by roughly 90%, in a public postmortem dated June 24, 2026.23 Its CEO's own account is unglamorous in exactly the way that makes it credible: AI costs had exceeded personnel costs at the company, and moving off Claude was, in his words, “a matter of survival for the business.”24 It's worth noting Lindy didn't move everything indiscriminately — it tested Kimi K2.5, GLM 5.1, and Kimi K2.6 against production traffic and rejected all three on real-world “vibe” despite promising offline benchmark scores, and kept Claude Sonnet for the subset of tasks that needed genuinely higher intelligence. The harness didn't just swap the cheapest model in; it did the discernment a flat “switch to open-weight” headline would erase.
Coinbase's CEO made the same bet in public, in numbers. On June 8, 2026, Brian Armstrong posted that Coinbase is “working hard on routing prompts to cheaper models,” predicting that within twelve to eighteen months, 80% of workloads will run on models that are 99% cheaper than today's frontier tier — with only the top 20%, the scientific-breakthrough and high-level-orchestrator work, staying on frontier tokens.25 And this isn't one company's anecdote against the grain: J.P. Morgan's own OpenRouter data shows Chinese open-weight models — Qwen, DeepSeek V4, Kimi — surging through the API-call leaderboard by April 2026, landing within “a few dozen Elo points” of closed frontier models at ten to fifty times lower token cost.1
“The escape hatch from frontier lock-in is documented, dated, and footnoted — by the same bank that underwrites the lock-in.”
It would be easy to read the harness cynically — as marketing dressed up as infrastructure. That reading doesn't survive contact with Lindy's postmortem or Coinbase's public routing bet, both of which are dated, both of which put a real company's own cost line on the record. A well-built agent harness genuinely does decouple “what the developer experiences as quality” from “which lab trained the model underneath it” — that's the mechanical fact this whole section rests on, not a sales pitch. The honest complication isn't that the escape hatch is fake. It's the one the next section takes up: who owns the door.
/ 04Escape, or Relocated Toll Booth?
The single biggest equity event at this layer in the entire research window runs directly against the escape-hatch thesis, and it happened nine days after Cembalest's harness sentence went to print.
SpaceX had just absorbed xAI in a $1.25 trillion merger in February 2026 and completed the largest IPO in history on June 12. On June 16, 2026, it announced it would acquire Cursor — the layer's largest venture-funded, model-agnostic harness — for $60 billion in stock, explicitly to put Grok in front of Cursor's developer base.26 The deal had reportedly been pre-committed as far back as April 2026, ahead of SpaceX's own IPO, structured as either the $60 billion stock acquisition or a $10 billion break-up fee if it fell through — which tells you how much certainty both sides wanted before SpaceX went public.27 Cursor's four cofounders, average age about twenty-five, each became roughly $2.7 billion richer on paper the day the deal was announced.28
Read straight, that's a genuinely different equity story from every other layer in this series. The value didn't flow up to whoever owns the scarce shell, the scarce silicon, the rented cloud, or the frontier model — it flowed to the team that built the software sitting between the user and whichever model happened to be cheapest that week. Cognition's Windsurf-to-$26B trajectory tells a version of the same story with different names. For once in this series, the escape hatch's builders got paid like the escape hatch was the valuable thing.
But sit with what the buyer is. The instant a harness gets large and valuable enough, a frontier-model-adjacent conglomerate bought it outright — folding the layer built to route around expensive frontier models into the balance sheet of a company that also owns one. As of this writing, the deal had not closed (targeted for the third quarter of 2026), and no public statement commits Cursor to Grok-exclusivity or confirms it will keep routing to Claude, GPT, Gemini, and Kimi the way it does today.27 That silence is the whole ballgame. If Cursor keeps routing wherever the model is best and cheapest after the deal closes, this was capital rewarding the escape hatch. If Cursor's routing quietly narrows toward Grok, this was the house buying the exit.
This is the single most consequential unresolved fact in this layer, and it is explicitly the fact Part 7's synthesis needs answered: does the SpaceX–Cursor deal preserve or kill Cursor's model-agnosticism? It is unresolved because the deal itself is unresolved — it had not closed as of the research date behind this article, and Anysphere, as a private company being acquired, has made no post-close routing commitment. Anthropic's own hedge here is instructive: the frontier lab whose token pricing power the harness is supposed to erode still owns the single most commercially successful harness product in the layer, Claude Code, at $2.5B-plus run-rate. If the harness wins regardless of which model sits underneath it, a frontier lab would rather own the harness too than bet everything on being the model.
/ 05Synthesis — Follow the Token
Follow the token: the subsidy at this layer is the cleanest, best-dated one in the whole series — a flat-rate seat priced well below what its heaviest use costs to serve at metered rates, absorbed for as long as the heavy users stayed a minority of the base. Anthropic's blended 44% gross margin is not evidence of a scam; it's evidence the arbitrage was, until recently, an ordinary SaaS subsidy pattern, light users paying for heavy ones. What makes 2026 different is that two of the parties closest to the meter — Microsoft as a buyer, Microsoft/GitHub as a seller — both concluded in the same month that the heavy tail wasn't staying a tail, and both repriced rather than keep absorbing it. The open question the labs haven't disclosed is the one that decides whether this was a one-time correction or the start of a trend: as agentic, multi-step, tool-calling workflows become the default way developers use these products rather than the edge case, does the average subscriber's usage keep drifting toward the ceiling that produces an $8,000-equivalent bill? Microsoft's and GitHub's own internal numbers, whatever they were, said yes loudly enough to act on inside a single quarter.
Follow the equity: this layer's steelman is the most genuine one in the series so far — real developer productivity, a real mechanism (Lindy, Coinbase, the JPM OpenRouter data) by which the harness lets a customer walk away from the most expensive frontier tokens without walking away from quality. The bear case is equally plain: almost nobody at this layer is profitable in a way a filing would confirm, and the one business that is — Cursor, at the enterprise-account level — got there specifically by reducing its own dependence on the frontier tokens the rest of the stack survives on. Whether that business stays the escape hatch or becomes the house's own toll booth is a question the SpaceX deal opened and has not yet answered.
Every layer in this series so far has answered “who holds the risk” with an institution — a builder, a bondholder, a ratepayer, a pension. This is the first layer where the answer is closer to a design choice: a harness that stays model-agnostic genuinely relocates value away from whoever owns the frontier model. A harness that gets bought by one doesn't. The variable that decides which future you're underwriting when you use one of these tools isn't the model. It's who owns the software making the choice for you — and, one rung further up, who owns the data that software is working on. Part 6 follows the token into the ontology.