The Business Engineer · September 20, 2026

Open Escape Velocity

Open-weight AI has reached production escape velocity — but the ecosystem has not yet proved it can pay for the next model.

Analysis by Gennaro Cuofano · The Business Engineer · September 20, 2026
Executive SummaryBy Gennaro Cuofano · The Business Engineer · 2026-09-20

Synopsis🔗

Gennaro Cuofano argues that open-weight AI has reached 'open escape velocity': on major inference gateways open models now carry a majority of token volume, and a wider industrial system — labs training, infrastructure companies serving, chip vendors benefiting, gateways routing, enterprises adapting — gives the ecosystem momentum that no longer depends on any single model company.

He draws two hard lines: gateway token share is not economic value (open models were 56% of Vercel's tokens but only about 14% of its spending), and operational escape velocity — keeping existing models useful — is not developmental escape velocity — financing and training their successors. The unanswered question, and the gap the sponsorship model must close, is who pays for the next model.

View this analysis as a KG entity →
Narrative

The argument, section by section🔗

Nineteen sections trace the thesis from the closed-model assumption through the gateway data, the unbundling of the value chain, and the enterprise playbook, closing on the twelve mental models.

Section 1

The closed-model assumption, revisited🔗

For almost four years the AI industry assumed closed models would dominate the market: OpenAI established the pattern and Anthropic reinforced it. The DeepSeek moment changed the picture — a class of open-weight model could compete on cost, capability, and deployability in ways that mattered to enterprises — and through 2026 that opening became harder to ignore as more capable open models, better serving infrastructure, and easier hosting arrived.

Section 2

The gateway data🔗

In August 2026 open-weight models processed 56% of all tokens on Vercel's AI Gateway, up from 7% in December 2025; OpenRouter reported roughly 60% of US-originating token consumption by open models in the same period. These are not industry market share: gateway data sees workloads routed through those platforms, not direct API contracts, consumer subscriptions, private deployments, or enterprise agreements — and gateways also miss some open-model usage precisely because open weights can run privately. The narrower, still important reading: within production environments designed around model choice, open weights moved from experimentation into real production infrastructure.

Section 3

Training, serving, routing, and delivery separate🔗

A model can now be trained by one organization, served by another, routed through a third, embedded inside a fourth company's product, and specialized by a fifth for an enterprise customer — training, serving, distribution, customization, and application delivery no longer have to sit inside the same company. Many actors have independent reasons to keep the open frontier moving: inference companies want workloads, chip vendors want compute demand, gateways want model diversity, application developers want substitutable intelligence, enterprises want lower costs and greater control, and regional labs want bases they can adapt. This is what the article means by open escape velocity — while insisting that adoption and self-renewal are different things.

Section 4

Why low spending share can be a sign of success🔗

Open-weight models carried 56% of Vercel's token volume but only around 14% of estimated spending, while Anthropic alone represented 64% of spending. The obvious reading — open models absorb cheap commodity work while closed models retain expensive high-value tasks — has some truth, but value needs unpacking. Token share (inference processed), spending share (dollars paid), and economic value (useful outcome relative to total cost) are not interchangeable: a $5 workload completed for $1 on an open model captures fewer dollars while the buyer receives more value. Vercel reported average token prices fell 23.2% in August, and customers leaving the most expensive Anthropic models often moved down inside Anthropic's own family. The market looks like a continuously repriced portfolio of intelligence, and the enterprise question becomes: what is the lowest-cost system that can complete this piece of work at the quality, latency, and risk level we require?

Section 5

Follow the inference dollar🔗

A Chinese-developed open model consumed through an American inference provider splits its economics across model developer, infrastructure provider, inference host, gateway, application layer, and enterprise — spending associated with a model is not automatically revenue for the company that trained it. Open weights unbundle the economics around the model, so hosting, routing, tuning, evaluation, vertical applications, and enterprise deployment can all become independent businesses. But it also produces the hardest economic question in the whole thesis: who pays for the next model? Serving an existing model can be profitable; training its substantially better successor is a different capital problem.

Section 6

'Open' is a set of rights, not one product category🔗

The word 'open' conceals several different arrangements: downloadable weights without reproducible training, hosted fine-tuning without export rights, free internal use with conditions on commercial hosting. For an enterprise openness is a bundle of practical rights: can we obtain the model, modify it, operate it independently, and continue using the adaptations we create under acceptable terms? There is also a technical distinction — legal permission and operational independence are different assets — and a hosted open model can sometimes create more practical optionality than a self-hosted model depending on a stack the enterprise cannot realistically reproduce. Openness matters when it translates into credible control.

Section 7

Why NVIDIA has a reason to fund the commons🔗

A proprietary model company must monetize access to the intelligence it creates; NVIDIA's business benefits whenever more AI is trained and run, regardless of whether the weights are sold, released openly, or served by somebody else — the economics of complementary products. The Nemotron Coalition brings model developers, application companies, tooling providers, and regional AI players around shared model development with training infrastructure contributed through DGX Cloud. The paradox: the model layer can become more open while the infrastructure beneath it remains highly concentrated, increasing the strategic importance of GPU vendors, clouds, and well-capitalized inference providers.

Section 8

Escape velocity has two different meanings🔗

Operational escape velocity: useful models can continue being deployed and improved without depending entirely on the original publisher's hosted service. Developmental escape velocity: several independent organizations can continue producing competitive successor models. Ten hosts serving one model provide operational redundancy, not ten independent sources of new model development; hundreds of fine-tunes can form a productive ecosystem while depending on a small number of organizations to produce the underlying bases. The three engines are usage, sponsorship, and institutional backing — evidence today is strongest on usage, increasingly interesting on sponsorship and institutional support, while operational resilience is arriving faster than developmental independence.

Section 9

Chinese models, American hosts, and a more complicated map🔗

A substantial portion of open-model volume has come from Chinese developers while those same models can be served through American infrastructure, routed through Western platforms, and embedded in enterprise applications elsewhere; in some configurations the original model lab never receives the customer's prompts. Model provenance and data destination are different questions. Enterprise procurement should therefore separate five dimensions: model provenance, inference location, service operator, data access and controls, and contractual and legal regime — a supply chain, not a national flag attached to an API.

Section 10

Institutional support is not a settled policy regime🔗

The July 'Open Weights and American AI Leadership' letter, with published signatories including NVIDIA, Microsoft, Meta, Google, OpenAI, and Amazon, demonstrates substantial organized support for open weights — evidence of organized support, not a settled regulatory regime. Anthropic has stated it has not advocated a blanket ban on open-weight models while arguing more capable systems can require different treatment; Dario Amodei's proposal to pace frontier development and introduce embedded third-party evaluators addresses supervision of the most capable systems while preserving diffusion elsewhere. The useful frame is 'pacing at the frontier, diffusion at the base' — with the boundary between those layers not fixed, since today's frontier capability becomes tomorrow's common infrastructure.

Section 11

Microsoft's more important bet: the enterprise improvement cycle🔗

Microsoft's June announcement of seven MAI models tied model development to enterprise customization: distribution through Foundry and external providers alongside Frontier Tuning intended to adapt models to customer workflows. The more important idea is continuous improvement of a general model against the organization's own operating environment. But not all institutional knowledge belongs inside model weights — balances belong in systems of record, policies need authoritative sources, credentials need explicit controls. The enterprise learning loop is larger than fine-tuning: outcomes, expert corrections, evaluation cases, business context, workflow changes, and tested model improvements — and the meaningful ownership claim is not 'we customized a model' but 'we can continue improving the capability we built.'

Section 12

Where open models fit in an enterprise portfolio🔗

The practical choice is which deployment arrangement fits each workload: a frontier model for difficult reasoning, a cheaper open model for high-volume classification or extraction, a specialized model for a stable internal domain, and a locally deployed model where privacy, latency, or sovereignty creates a specific requirement. These are decisions about workload economics and control, not a ranking of model quality. A common API does not make substitution effortless — tool behavior, context handling, structured outputs, latency, reliability, and failure modes can all change — and the goal is credible optionality: evaluated alternatives, understood switching costs, and known effort for each consequential workload.

Section 13

Free weights can still produce expensive work🔗

Removing a model-license charge does not remove the cost of operating the model: capacity still has to be purchased, maintained, and utilized, and a lightly used self-hosted deployment can be more expensive than a managed service. The relevant comparison is the full cost of useful work — inference, infrastructure, engineering, human review, monitoring, failure remediation, maintenance, and migration. A cheap model requiring substantially more correction can produce expensive work; a more expensive model removing a large review burden can produce cheaper work. The economically useful unit is cost per accepted result: how much the organization spent across the entire system to produce an outcome it was actually prepared to accept.

Section 14

Ownership requires instrumentation🔗

A downloadable model does not enforce permissions, prevent duplicate actions, validate retrieved information, detect drift, or determine who should approve a consequential decision — those controls belong to the operating system around the model. If an enterprise runs the model it needs to know whether it is behaving correctly; if it changes or fine-tunes it, whether the change improved performance and what deteriorated elsewhere; if the system takes actions, what actually happened afterward. Evaluation, tracing, auditing, monitoring, and governance become part of the open-model stack rather than optional extras. Owning weights without outcomes gives you an artifact; outcomes without evaluations give you data; evaluations without the ability to change the system give you visibility. The strategic asset appears when those components form a loop capable of producing measured improvement.

Section 15

The biggest risk: dependency moves rather than disappears🔗

Open weights can remove dependence on one model endpoint while leaving several other dependencies untouched: a particular hardware family, inference provider, proprietary tuning environment, developer framework, cloud platform, or distribution channel. The objective cannot be complete independence from every supplier; the useful distinction is between dependencies that were deliberately accepted and dependencies that became invisible. A hosted open model might be technically portable while the customer's monitoring and data pipelines make migration expensive; an exported adaptation may still depend on the license of its underlying base model; a huge derivative ecosystem may still depend on only a few organizations capable of training new bases. Escape from one provider is not escape from the supply chain.

Section 16

What would confirm escape velocity?🔗

The stronger thesis needs evidence across several dimensions: continued model renewal (multiple independent organizations producing competitive successors), sustainable operating businesses (hosts and application companies covering infrastructure and support costs), practical portability (moving workloads or retaining adaptations without rebuilding the system), workable rights (licenses and contracts supporting needed deployment patterns), and enterprise improvement (retained feedback, evaluations, and proprietary context translating into measurable improvements). Different failures weaken different parts: strong hosting demand with concentrated development shows operational resilience but not developmental independence; broad choice with prohibitively difficult migration shows availability but not practical control. A low spending share would not automatically mean failure — it could mean competition drove the same useful work to a lower price. The serious warning is an ecosystem creating enormous customer value while failing to finance the production and maintenance of its shared foundations.

Section 17

The AI supercycle connection🔗

When useful intelligence is available from more sources, access to one particular model becomes less exclusive and the difficult work moves toward choosing, operating, adapting, evaluating, and applying those models. Open models do not need to capture most of AI industry revenue to change the market — they only need to make intelligence sufficiently substitutable that the scarcity premium of the model layer begins to compress. Once that happens, value shifts toward whatever remains scarce: infrastructure, proprietary context, evaluation data, trusted workflows, customer relationships, distribution, and accumulated operating knowledge. Openness increases the number of parties that can participate; it does not determine which layer ultimately keeps the margin.

Section 18

The compression🔗

Open-weight AI has crossed an important production threshold: on major gateways it carries a substantial share of inference, and around it a wider industrial system is forming whose incentives are no longer concentrated in one place. But operational escape velocity and developmental escape velocity must remain separate — the ecosystem is increasingly capable of keeping existing models useful when one participant changes direction, and has not yet proved enough independent organizations can continuously finance and train their successors. For enterprises the implication is immediate: do not confuse token volume with value, downloadable weights with independence, customization with ownership, or cheap inference with cheap work. Build what survives model churn — context, evaluations, workflow, outcome history, operating knowledge, and the improvement loop. Cheap intelligence is becoming easier to obtain; the durable advantage is learning how to turn it into useful work, and retaining the ability to keep improving that work when the model underneath changes.

Section 19

The mental models🔗

The argument rests on twelve mental models, each travelling beyond open models: Three Measurements; Follow the Inference Dollar; Openness as a Set of Rights; Complements Fund the Commons; Two Escape Velocities; The Supply-Chain Map; Pacing at the Frontier, Diffusion at the Base; Credible Optionality; Cost per Accepted Result; You Cannot Own What You Cannot Measure; Dependency Moves, It Does Not Vanish; and The Scarcity Premium Compresses.

Mental models

Twelve mental models🔗

The argument rests on twelve mental models, each travelling beyond open models: Three Measurements; Follow the Inference Dollar; Openness as a Set of Rights; Complements Fund the Commons; Two Escape Velocities; The Supply-Chain Map; Pacing at the Frontier, Diffusion at the Base; Credible Optionality; Cost per Accepted Result; You Cannot Own What You Cannot Measure; Dependency Moves, It Does Not Vanish; and The Scarcity Premium Compresses.

Mental model 1 of 12

Three Measurements🔗

Mental model: token share, spending share, and economic value measure different things. Applied here: the 56% token share and 14% spending share describe consumption and supplier revenue; neither measures what customers saved.

Mental model 2 of 12

Follow the Inference Dollar🔗

Mental model: spending associated with a model is not revenue received by the company that trained it. Applied here: trace the payment across developer, infrastructure, host, gateway, application, and enterprise before crediting any layer with it.

Mental model 3 of 12

Openness as a Set of Rights🔗

Mental model: obtaining, modifying, operating independently, and keeping adaptations are separate permissions. Applied here: openness counts only when it translates into credible control, not another file on a server.

Mental model 4 of 12

Complements Fund the Commons🔗

Mental model: a company can sponsor a free input when it sells something that grows with it. Applied here: when a model arrives free, ask whose other product benefits — NVIDIA's accelerators and infrastructure, for instance — and whether that incentive will last.

Mental model 5 of 12

Two Escape Velocities🔗

Mental model: keeping released models useful is operational; producing competitive successors is developmental. Applied here: count independent training programs, not hosts or fine-tunes, before assuming a model family will keep improving.

Mental model 6 of 12

The Supply-Chain Map🔗

Mental model: model provenance, inference location, service operator, data access and controls, and the contractual and legal regime are separate questions. Applied here: make procurement check each one instead of letting a national label stand in for all five.

Mental model 7 of 12

Pacing at the Frontier, Diffusion at the Base🔗

Mental model: different capability levels raise different policy questions, and the boundary between them moves. Applied here: treat today's frontier controls as provisional for capabilities that will soon be common infrastructure.

Mental model 8 of 12

Credible Optionality🔗

Mental model: the goal is evaluated alternatives, not a long list of providers. Applied here: for each consequential workload, know what would change in a switch and how much time and work it would take.

Mental model 9 of 12

Cost per Accepted Result🔗

Mental model: the price of a token is one term in the cost of useful work. Applied here: compare arrangements on inference, infrastructure, engineering, human review, monitoring, failure remediation, maintenance, and migration together.

Mental model 10 of 12

You Cannot Own What You Cannot Measure🔗

Mental model: weights alone are an artifact; outcomes, evaluations, and the ability to change the system turn them into an improvement loop. Applied here: the test of ownership is measured improvement that survives a change of supplier.

Mental model 11 of 12

Dependency Moves, It Does Not Vanish🔗

Mental model: removing one dependency often exposes or creates another. Applied here: map dependencies layer by layer, and separate the ones you accepted from the ones that became invisible.

Mental model 12 of 12

The Scarcity Premium Compresses🔗

Mental model: when intelligence becomes substitutable, value shifts to whatever remains scarce. Applied here: invest in context, evaluation data, trusted workflows, customer relationships, and operating knowledge, not in access to one model.

Evidence

The article’s quantitative claims🔗

The article's quantitative and structural claims, collected in one place. Every figure below is reported in the article itself — Vercel and OpenRouter gateway statistics, spending splits, token-price movement, the July policy letter, and Microsoft's MAI announcement — and none of it has been independently verified here. The figures are presented as the author's evidence for the escape-velocity thesis, not as audited market data.

Open models carry a majority of gateway token volume🔗

In August 2026 open-weight models processed 56% of all tokens on Vercel's AI Gateway, up from 7% in December 2025; OpenRouter reported roughly 60% of US-originating token consumption by open models in the same period. The article itself cautions that gateway data is neither a census nor a curiosity: it misses direct API contracts, private deployments, local models, and enterprise agreements. Framed as the article's reported evidence, not independently verified.

7%
Open-weight share of Vercel AI Gateway token volume, December 2025.
56%
Open-weight share of Vercel AI Gateway token volume, August 2026.
60%
Open-model share of US-originating token consumption reported by OpenRouter for the same period.

Open models hold a fraction of the spending🔗

Open-weight models carried 56% of Vercel's token volume but only around 14% of estimated spending, while Anthropic alone represented 64% of spending. The article reads this as partly commodity absorption and partly a measurement problem: spending share is supplier revenue, not customer value. Framed as the article's reported evidence, not independently verified.

64%
Anthropic's share of estimated spending on Vercel's AI Gateway, August 2026.
14%
Open-weight share of estimated spending on Vercel's AI Gateway, August 2026.

Inference prices keep falling🔗

Vercel reported that average token prices fell 23.2% in August, continuing a multi-month decline; among high-volume customers present in both comparison periods, the median decline was considerably smaller — which matters because an aggregate price index should not be treated as the saving experienced by every workload. Framed as the article's reported evidence, not independently verified.

23.2%
Decline in average token prices reported by Vercel, August 2026.

The market is a continuously repriced portfolio of intelligence🔗

The market is settling into a continuously repriced portfolio of intelligence rather than a clean division where open models do the cheap work and one closed frontier model does everything important. Evidence cited: customers moving away from the most expensive Anthropic models often moved down inside Anthropic's own model family rather than leaving the provider. Framed as the author's thesis, not a measured result.

The value chain has unbundled across five different organizations🔗

Training, serving, distribution, customization, and application delivery no longer have to sit inside the same company: a model can be trained by one organization, served by another, routed through a third, embedded in a fourth company's product, and specialized by a fifth for an enterprise customer. The structural separation is presented as what makes this cycle different from earlier open-AI cycles that depended on one or two labs releasing weights. Framed as the author's thesis, not a measured result.

Openness at the model layer coexists with concentrated infrastructure🔗

The model layer can become more open while the infrastructure beneath it remains highly concentrated: enterprises may gain dozens of viable model choices while those models still depend disproportionately on a small number of GPU vendors, clouds, or well-capitalized inference providers. Openness can reduce concentration in one layer while increasing the strategic importance of another. Framed as the author's thesis, not a measured result.

Who pays for the next model remains unproven🔗

Serving an existing model can be a profitable business; training its substantially better successor is a different capital problem. The open ecosystem is becoming increasingly effective at distributing and monetizing models once they exist; what remains less proven is whether enough value flows back toward the organizations capable of producing the next generation of them. Closing that gap is, in the article's framing, what the sponsorship model ultimately has to do. Framed as the author's thesis, not a measured result.

Entities

Core entities🔗

Organizations, people, works, and platforms named or central to the article — every name links to its knowledge-graph entity.

Organizations🔗

Amazon🔗

Signatory of the July Open Weights and American AI Leadership letter.

Anthropic🔗

AI company that reinforced the closed-model pattern; alone represented 64% of estimated spending on Vercel's AI Gateway in August 2026. Has stated it has not advocated a blanket ban on open-weight models while arguing more capable systems can require different treatment.

DeepSeek🔗

Company whose open-weight release — 'the DeepSeek moment' — showed an open-weight model could compete on cost, capability, and deployability in ways that mattered to enterprises, changing the picture the article describes.

Google🔗

Signatory of the July Open Weights and American AI Leadership letter.

Meta🔗

Company whose Llama releases previously carried much of the open-model momentum — the article notes the earlier dependence on whether a large organization continued releasing competitive weights; a signatory of the July Open Weights and American AI Leadership letter.

Microsoft🔗

Company that announced seven MAI models in June 2026 with Foundry distribution and Frontier Tuning, tying model development to enterprise customization and the enterprise improvement cycle; a signatory of the July Open Weights and American AI Leadership letter.

NVIDIA🔗

Chip and infrastructure company with a complementary-products reason to fund the open-model commons; brought together the Nemotron Coalition and contributes training infrastructure through DGX Cloud. A signatory of the July Open Weights and American AI Leadership letter.

OpenAI🔗

AI company that established the closed-model pattern the article opens with; later a signatory of the July Open Weights and American AI Leadership letter.

OpenRouter🔗

Gateway company reporting roughly 60% of US-originating token consumption by open models in the same period as the Vercel figures.

Vercel🔗

Company whose AI Gateway figures anchor the article's evidence: 56% of August 2026 token volume from open-weight models (up from 7% in December 2025), about 14% of estimated spending, 64% of spending to Anthropic alone, and a 23.2% average token price decline in August.

People🔗

Dario Amodei🔗

Anthropic CEO; his proposal to pace frontier development and introduce embedded third-party evaluators is cited as part of the emerging 'pacing at the frontier, diffusion at the base' policy structure.

Gennaro Cuofano🔗

Gennaro Cuofano, author of The Business Engineer newsletter and of 'Open Escape Velocity'.

The Business Engineer🔗

The Business Engineer, Gennaro Cuofano's Substack newsletter in which 'Open Escape Velocity' was published.

How-To

Evaluate open-weight AI for an enterprise workload🔗

The article's implicit enterprise protocol for evaluating open-weight AI against a real workload: ask the portfolio question, separate the three measurements, map the supply chain, test credible optionality, and instrument before owning.

1

Ask the portfolio question

Stop asking 'which model provider are we standardized on?' and ask instead: what is the lowest-cost system that can complete this piece of work at the quality, latency, and risk level we require? Decide per workload whether a frontier model, a cheaper open model, a specialized model, or a locally deployed model fits — these are decisions about workload economics and control, not a ranking of model quality.

2

Separate the three measurements

Measure token share, spending share, and economic value separately, and never let token volume stand in for value. A lower spending share is not automatically a failure — it can mean competition drove the same useful work to a lower price — and an aggregate price index should not be treated as the saving experienced by every workload.

3

Map the supply chain

Check each of the five procurement dimensions separately rather than reducing the choice to a national label: model provenance, inference location, service operator, data access and controls, and the contractual and legal regime. In some configurations the original model lab never receives the customer's prompts — but the model still carries its original architecture, training history, license, and limitations.

4

Test credible optionality

Evaluate the alternatives for each consequential workload: understand what would change in a switch — tool behavior, context handling, structured outputs, latency, reliability, failure modes — and know how much time and work it would take. A common API can connect to another endpoint in minutes while the evaluation to switch safely takes weeks.

5

Instrument before you own

Treat evaluation, tracing, auditing, monitoring, and governance as part of the stack, not optional extras: know whether the model is behaving correctly, whether a change or fine-tune improved performance and what deteriorated elsewhere, and what actually happened after the system took actions. Compare arrangements on cost per accepted result — inference, infrastructure, engineering, human review, monitoring, failure remediation, maintenance, and migration together.

FAQ

Frequently Asked Questions🔗

It means open AI is beginning to develop enough independent sources of demand, infrastructure, capital, and institutional support that its future depends less on the strategy of any single model company. It explicitly does not mean open models have defeated closed models, that every dependency has disappeared, or that the ecosystem is already economically self-sufficient.

Answer entity ↗

No. Closed models still dominate important parts of the market, particularly at the frontier. The article's claim is about independent momentum: open weights have become part of the production architecture itself, with incentives no longer concentrated in one place.

Answer entity ↗

In August 2026 open-weight models processed 56% of tokens on Vercel's AI Gateway (up from 7% in December 2025) but only about 14% of estimated spending, while Anthropic alone was 64% of spending. They are not industry market share: gateway data misses direct API contracts, consumer subscriptions, private deployments, and enterprise agreements, and private hosting hides some open-model usage too. What they do show is that within production environments designed around model choice, open weights moved from experimentation into real production infrastructure.

Answer entity ↗

They measure different things and can move in opposite directions. Token share is how much inference a model processes; spending share is how many dollars are paid for it; economic value is what useful outcome the customer receives relative to total cost. Suppose a workload previously cost $5 and an open model now completes the same accepted work for $1: the open model captures fewer dollars, but the buyer may receive more economic value because most of the surplus stays with the customer.

Answer entity ↗

Operational escape velocity means useful models can continue being deployed and improved without depending entirely on the original publisher's hosted service. Developmental escape velocity means several independent organizations can continue producing competitive successor models. The first does not prove the second: ten hosts serving one model provide operational redundancy, not ten independent sources of new model development. Serving is not training, and adaptation is not frontier development.

Answer entity ↗

Usage: enough real work moves onto open models to support hosts, tooling companies, and applications. Sponsorship: enough organizations have independent economic reasons to finance model development, infrastructure, and ecosystem investment. Institutional backing: enough companies, standards bodies, foundations, and policymakers treat open models as strategically worth maintaining. Evidence today is strongest on usage; operational resilience is arriving faster than developmental independence.

Answer entity ↗

Trace the payment across the separated layers — model developer, infrastructure provider, inference host, gateway, application layer, enterprise — before crediting any layer with it, because spending associated with a model is not automatically revenue received by the company that trained it. Once the layers separate, economics flow toward hardware vendors, clouds, networking providers, gateways, and application vendors, while the model developer monetizes through its own API, licenses, enterprise agreements, support, or adjacent products.

Answer entity ↗

Who pays for the next model? Serving an existing model can be a profitable business; training its substantially better successor is a different capital problem. The open ecosystem is increasingly effective at distributing and monetizing models once they exist; what remains less proven is whether enough value flows back toward the organizations capable of producing the next generation of them.

Answer entity ↗

No. A model can have downloadable weights without disclosing everything needed to reproduce its training; a developer can be allowed to fine-tune a hosted model without being allowed to export it; a model can be free for internal use while imposing additional conditions on commercial hosting. The word 'open' conceals several different arrangements, and the strategic question is whether the organization can continue operating and improving the capability if a supplier changes pricing, terms, or strategy.

Answer entity ↗

Four rights: can we obtain the model, can we modify it, can we operate it independently, and can we continue using the adaptations we create under acceptable terms? A positive answer to one does not settle the others, and legal permission and operational independence are different assets. Openness counts when it translates into credible control, not when it merely produces another file on a server.

Answer entity ↗

Because of complementary-products economics: more capable open models can create more applications, more applications create more inference, and more inference creates greater demand for accelerators, networking, serving software, and compute infrastructure — the parts of the stack NVIDIA sells. The Nemotron Coalition, bringing model developers, application companies, tooling providers, and regional AI players together with training infrastructure through DGX Cloud, makes that logic concrete. It is a more durable sponsorship than depending on a lab's philosophical commitment to openness, though NVIDIA's upside depends on its infrastructure remaining an attractive place to run the workloads openness creates.

Answer entity ↗

An emerging structure in the policy debate rather than a settled consensus: at the frontier, the questions concern capability thresholds, evaluations, security, inspection, and how quickly systems should advance — Dario Amodei's proposal to pace frontier development with embedded third-party evaluators fits here — while below the frontier the economic pressure favors broader access, competition, customization, and enterprise control. The boundary is not fixed: today's frontier capability becomes tomorrow's common infrastructure, so 'open versus closed' may prove too static a policy frame.

Answer entity ↗

Evaluated alternatives, not a long list of providers. Real optionality exists when the enterprise has evaluated alternatives, understands what would change in a switch — tool behavior, context handling, structured outputs, latency, reliability, failure modes — and knows how much time and work the switch would require. The goal is not maximum provider count.

Answer entity ↗

The full cost of useful work: inference, infrastructure, engineering, human review, monitoring, failure remediation, maintenance, and migration together, for an outcome the organization was actually prepared to accept. A cheap model requiring substantially more correction can produce expensive work; a more expensive model removing a large review burden can produce cheaper work. Falling token prices matter because they reduce one cost component; whether they create economic value depends on the rest of the workflow.

Answer entity ↗

Five dimensions: continued model renewal, sustainable operating businesses, practical portability, workable rights, and enterprise improvement. Different failures weaken different parts: strong hosting demand with concentrated model development shows operational resilience but not developmental independence; broad choice with prohibitively difficult migration shows availability but not practical control. A low spending share would not automatically mean failure — competition could simply have driven the same useful work to a lower price. The serious warning is an ecosystem creating enormous customer value while failing to finance the production and maintenance of its shared foundations.

Answer entity ↗
Glossary

Core technical glossary🔗

Terms introduced or defined in the article, including the twelve mental models the argument rests on.

AI gateway🔗

A routing layer across providers and models that handles access, routing, and billing, and may monetize observability, enterprise controls, or other services rather than marking up every token. The article's token and spending statistics come from Vercel's AI Gateway and OpenRouter.

Complements Fund the Commons🔗

Mental model: a company can sponsor a free input when it sells something that grows with it. Applied here: when a model arrives free, ask whose other product benefits — NVIDIA's accelerators and infrastructure, for instance — and whether that incentive will last.

Cost per Accepted Result🔗

Mental model: the price of a token is one term in the cost of useful work. Applied here: compare arrangements on inference, infrastructure, engineering, human review, monitoring, failure remediation, maintenance, and migration together.

Credible Optionality🔗

Mental model: the goal is evaluated alternatives, not a long list of providers. Applied here: for each consequential workload, know what would change in a switch and how much time and work it would take.

Dependency Moves, It Does Not Vanish🔗

Mental model: removing one dependency often exposes or creates another. Applied here: map dependencies layer by layer, and separate the ones you accepted from the ones that became invisible.

Follow the Inference Dollar🔗

Mental model: spending associated with a model is not revenue received by the company that trained it. Applied here: trace the payment across developer, infrastructure, host, gateway, application, and enterprise before crediting any layer with it.

Frontier model🔗

The most capable models at the leading edge, where closed providers still dominate important parts of the market and where policy questions increasingly concern capability thresholds, evaluations, security, inspection, and pacing. Today's frontier capability becomes tomorrow's common infrastructure.

Inference host🔗

The layer that turns model weights into an available service, supplying compute, memory, serving software, networking, and availability. In the unbundled value chain it is distinct from the model developer that trained the weights.

Nemotron Coalition🔗

NVIDIA's coalition bringing model developers, application companies, tooling providers, and regional AI players around shared model development, with training infrastructure contributed through DGX Cloud — the concrete expression of complementary-products economics funding the open-model commons.

Open escape velocity🔗

The article's central claim: open AI developing enough independent sources of demand, infrastructure, capital, and institutional support that its future depends less on the strategy of any single model company. Not that open models have defeated closed models, not that every dependency has disappeared, and not that the ecosystem is already economically self-sufficient.

Open-weight model🔗

A model whose weights can be downloaded and deployed, without necessarily disclosing training data, training process, or everything needed to reproduce the training. Not interchangeable with 'open-source': weights alone are not the same as having an efficient serving stack, suitable hardware, capacity, or an operating team.

Openness as a Set of Rights🔗

Mental model: obtaining, modifying, operating independently, and keeping adaptations are separate permissions. Applied here: openness counts only when it translates into credible control, not another file on a server.

You Cannot Own What You Cannot Measure🔗

Mental model: weights alone are an artifact; outcomes, evaluations, and the ability to change the system turn them into an improvement loop. Applied here: the test of ownership is measured improvement that survives a change of supplier.

Pacing at the Frontier, Diffusion at the Base🔗

Mental model: different capability levels raise different policy questions, and the boundary between them moves. Applied here: treat today's frontier controls as provisional for capabilities that will soon be common infrastructure.

The Scarcity Premium Compresses🔗

Mental model: when intelligence becomes substitutable, value shifts to whatever remains scarce. Applied here: invest in context, evaluation data, trusted workflows, customer relationships, and operating knowledge, not in access to one model.

The Supply-Chain Map🔗

Mental model: model provenance, inference location, service operator, data access and controls, and the contractual and legal regime are separate questions. Applied here: make procurement check each one instead of letting a national label stand in for all five.

Three Measurements🔗

Mental model: token share, spending share, and economic value measure different things. Applied here: the 56% token share and 14% spending share describe consumption and supplier revenue; neither measures what customers saved.

Two Escape Velocities🔗

Mental model: keeping released models useful is operational; producing competitive successors is developmental. Applied here: count independent training programs, not hosts or fine-tunes, before assuming a model family will keep improving.

Value-chain unbundling🔗

The separation of training, serving, distribution, customization, and application delivery across different organizations — a model trained by one, served by another, routed through a third, embedded by a fourth, specialized by a fifth — which creates the opportunity for an independent ecosystem around open models and the question of who pays for the next one.

Knowledge Graph Explorer 139 nodes · 478 links

Interactive graph visualization derived from the companion RDF. Click nodes to resolve, drag to explore. Graph data embedded from companion RDF at generation time.

Open Escape Velocity — Interactive Infographic🔗

Nodes: 0 Links: 0
Click SVG to activate zoom, click outside to release | Drag nodes to pin, double-click to unpin
Classes Properties Instances

SPARQL Workbench 3 sample queries 🔗

Explore Knowledge Graph using SPARQL against this knowledge graph on URIBurner. The editor opens on the canonical SAMPLE entity-type summary. Pick a recipe, edit freely, then run live or copy. Live execution targets the named graph selected below — the companion Turtle can be loaded into the DAV named graph listed in the footer.

Query recipes🔗

Entity type summary

Counts entities in the companion named graph by rdf:type. Default query for the SPARQL workbench.

PREFIX schema: <http://schema.org/>
PREFIX rdfs: <http://www.w3.org/2000/01/rdf-schema#>
SELECT ?typeIri (SAMPLE(?typeLabel) AS ?type) (COUNT(?s) AS ?count)
WHERE {
  GRAPH <https://substack.com/app-link/post?publication_id=594665&post_id=216511056> {
    ?s a ?typeIri .
    OPTIONAL { ?typeIri rdfs:label ?typeLabel }
  }
}
GROUP BY ?typeIri
ORDER BY DESC(?count)

Endpoint: https://linkeddata.uriburner.com/sparql

FAQ questions and answers

Lists every FAQ question in the collection with its accepted answer and both entity IRIs.

PREFIX schema: <http://schema.org/>
SELECT ?questionIri ?question ?answerIri ?answer
WHERE {
  GRAPH <https://substack.com/app-link/post?publication_id=594665&post_id=216511056> {
    ?questionIri a schema:Question ;
                 schema:name ?question ;
                 schema:acceptedAnswer ?answerIri .
    ?answerIri schema:text ?answer .
  }
}
ORDER BY ?question

Endpoint: https://linkeddata.uriburner.com/sparql

Glossary terms and definitions

Lists every glossary term in the collection — including the twelve mental models — with its definition.

PREFIX schema: <http://schema.org/>
SELECT ?termIri ?term ?definition
WHERE {
  GRAPH <https://substack.com/app-link/post?publication_id=594665&post_id=216511056> {
    ?termIri a schema:DefinedTerm ;
              schema:name ?term ;
              schema:description ?definition .
  }
}
ORDER BY ?term

Endpoint: https://linkeddata.uriburner.com/sparql

Live editor🔗

▶ Run live on URIBurner 🔗 Run live SELECT: text/x-html+tr | DESCRIBE/CONSTRUCT: text/x-html-nice-turtle

Live query links are built with encodeURIComponent(query). SELECT results render as text/x-html+tr; DESCRIBE/CONSTRUCT as text/x-html-nice-turtle.