Nineteen sections trace the thesis from the closed-model assumption through the gateway data, the unbundling of the value chain, and the enterprise playbook, closing on the twelve mental models.
Section 1For almost four years the AI industry assumed closed models would dominate the market: OpenAI established the pattern and Anthropic reinforced it. The DeepSeek moment changed the picture — a class of open-weight model could compete on cost, capability, and deployability in ways that mattered to enterprises — and through 2026 that opening became harder to ignore as more capable open models, better serving infrastructure, and easier hosting arrived.
Section 2In August 2026 open-weight models processed 56% of all tokens on Vercel's AI Gateway, up from 7% in December 2025; OpenRouter reported roughly 60% of US-originating token consumption by open models in the same period. These are not industry market share: gateway data sees workloads routed through those platforms, not direct API contracts, consumer subscriptions, private deployments, or enterprise agreements — and gateways also miss some open-model usage precisely because open weights can run privately. The narrower, still important reading: within production environments designed around model choice, open weights moved from experimentation into real production infrastructure.
Section 3A model can now be trained by one organization, served by another, routed through a third, embedded inside a fourth company's product, and specialized by a fifth for an enterprise customer — training, serving, distribution, customization, and application delivery no longer have to sit inside the same company. Many actors have independent reasons to keep the open frontier moving: inference companies want workloads, chip vendors want compute demand, gateways want model diversity, application developers want substitutable intelligence, enterprises want lower costs and greater control, and regional labs want bases they can adapt. This is what the article means by open escape velocity — while insisting that adoption and self-renewal are different things.
Section 4Open-weight models carried 56% of Vercel's token volume but only around 14% of estimated spending, while Anthropic alone represented 64% of spending. The obvious reading — open models absorb cheap commodity work while closed models retain expensive high-value tasks — has some truth, but value needs unpacking. Token share (inference processed), spending share (dollars paid), and economic value (useful outcome relative to total cost) are not interchangeable: a $5 workload completed for $1 on an open model captures fewer dollars while the buyer receives more value. Vercel reported average token prices fell 23.2% in August, and customers leaving the most expensive Anthropic models often moved down inside Anthropic's own family. The market looks like a continuously repriced portfolio of intelligence, and the enterprise question becomes: what is the lowest-cost system that can complete this piece of work at the quality, latency, and risk level we require?
Section 5A Chinese-developed open model consumed through an American inference provider splits its economics across model developer, infrastructure provider, inference host, gateway, application layer, and enterprise — spending associated with a model is not automatically revenue for the company that trained it. Open weights unbundle the economics around the model, so hosting, routing, tuning, evaluation, vertical applications, and enterprise deployment can all become independent businesses. But it also produces the hardest economic question in the whole thesis: who pays for the next model? Serving an existing model can be profitable; training its substantially better successor is a different capital problem.
Section 6The word 'open' conceals several different arrangements: downloadable weights without reproducible training, hosted fine-tuning without export rights, free internal use with conditions on commercial hosting. For an enterprise openness is a bundle of practical rights: can we obtain the model, modify it, operate it independently, and continue using the adaptations we create under acceptable terms? There is also a technical distinction — legal permission and operational independence are different assets — and a hosted open model can sometimes create more practical optionality than a self-hosted model depending on a stack the enterprise cannot realistically reproduce. Openness matters when it translates into credible control.
Section 7A proprietary model company must monetize access to the intelligence it creates; NVIDIA's business benefits whenever more AI is trained and run, regardless of whether the weights are sold, released openly, or served by somebody else — the economics of complementary products. The Nemotron Coalition brings model developers, application companies, tooling providers, and regional AI players around shared model development with training infrastructure contributed through DGX Cloud. The paradox: the model layer can become more open while the infrastructure beneath it remains highly concentrated, increasing the strategic importance of GPU vendors, clouds, and well-capitalized inference providers.
Section 8Operational escape velocity: useful models can continue being deployed and improved without depending entirely on the original publisher's hosted service. Developmental escape velocity: several independent organizations can continue producing competitive successor models. Ten hosts serving one model provide operational redundancy, not ten independent sources of new model development; hundreds of fine-tunes can form a productive ecosystem while depending on a small number of organizations to produce the underlying bases. The three engines are usage, sponsorship, and institutional backing — evidence today is strongest on usage, increasingly interesting on sponsorship and institutional support, while operational resilience is arriving faster than developmental independence.
Section 9A substantial portion of open-model volume has come from Chinese developers while those same models can be served through American infrastructure, routed through Western platforms, and embedded in enterprise applications elsewhere; in some configurations the original model lab never receives the customer's prompts. Model provenance and data destination are different questions. Enterprise procurement should therefore separate five dimensions: model provenance, inference location, service operator, data access and controls, and contractual and legal regime — a supply chain, not a national flag attached to an API.
Section 10The July 'Open Weights and American AI Leadership' letter, with published signatories including NVIDIA, Microsoft, Meta, Google, OpenAI, and Amazon, demonstrates substantial organized support for open weights — evidence of organized support, not a settled regulatory regime. Anthropic has stated it has not advocated a blanket ban on open-weight models while arguing more capable systems can require different treatment; Dario Amodei's proposal to pace frontier development and introduce embedded third-party evaluators addresses supervision of the most capable systems while preserving diffusion elsewhere. The useful frame is 'pacing at the frontier, diffusion at the base' — with the boundary between those layers not fixed, since today's frontier capability becomes tomorrow's common infrastructure.
Section 11Microsoft's June announcement of seven MAI models tied model development to enterprise customization: distribution through Foundry and external providers alongside Frontier Tuning intended to adapt models to customer workflows. The more important idea is continuous improvement of a general model against the organization's own operating environment. But not all institutional knowledge belongs inside model weights — balances belong in systems of record, policies need authoritative sources, credentials need explicit controls. The enterprise learning loop is larger than fine-tuning: outcomes, expert corrections, evaluation cases, business context, workflow changes, and tested model improvements — and the meaningful ownership claim is not 'we customized a model' but 'we can continue improving the capability we built.'
Section 12The practical choice is which deployment arrangement fits each workload: a frontier model for difficult reasoning, a cheaper open model for high-volume classification or extraction, a specialized model for a stable internal domain, and a locally deployed model where privacy, latency, or sovereignty creates a specific requirement. These are decisions about workload economics and control, not a ranking of model quality. A common API does not make substitution effortless — tool behavior, context handling, structured outputs, latency, reliability, and failure modes can all change — and the goal is credible optionality: evaluated alternatives, understood switching costs, and known effort for each consequential workload.
Section 13Removing a model-license charge does not remove the cost of operating the model: capacity still has to be purchased, maintained, and utilized, and a lightly used self-hosted deployment can be more expensive than a managed service. The relevant comparison is the full cost of useful work — inference, infrastructure, engineering, human review, monitoring, failure remediation, maintenance, and migration. A cheap model requiring substantially more correction can produce expensive work; a more expensive model removing a large review burden can produce cheaper work. The economically useful unit is cost per accepted result: how much the organization spent across the entire system to produce an outcome it was actually prepared to accept.
Section 14A downloadable model does not enforce permissions, prevent duplicate actions, validate retrieved information, detect drift, or determine who should approve a consequential decision — those controls belong to the operating system around the model. If an enterprise runs the model it needs to know whether it is behaving correctly; if it changes or fine-tunes it, whether the change improved performance and what deteriorated elsewhere; if the system takes actions, what actually happened afterward. Evaluation, tracing, auditing, monitoring, and governance become part of the open-model stack rather than optional extras. Owning weights without outcomes gives you an artifact; outcomes without evaluations give you data; evaluations without the ability to change the system give you visibility. The strategic asset appears when those components form a loop capable of producing measured improvement.
Section 15Open weights can remove dependence on one model endpoint while leaving several other dependencies untouched: a particular hardware family, inference provider, proprietary tuning environment, developer framework, cloud platform, or distribution channel. The objective cannot be complete independence from every supplier; the useful distinction is between dependencies that were deliberately accepted and dependencies that became invisible. A hosted open model might be technically portable while the customer's monitoring and data pipelines make migration expensive; an exported adaptation may still depend on the license of its underlying base model; a huge derivative ecosystem may still depend on only a few organizations capable of training new bases. Escape from one provider is not escape from the supply chain.
Section 16The stronger thesis needs evidence across several dimensions: continued model renewal (multiple independent organizations producing competitive successors), sustainable operating businesses (hosts and application companies covering infrastructure and support costs), practical portability (moving workloads or retaining adaptations without rebuilding the system), workable rights (licenses and contracts supporting needed deployment patterns), and enterprise improvement (retained feedback, evaluations, and proprietary context translating into measurable improvements). Different failures weaken different parts: strong hosting demand with concentrated development shows operational resilience but not developmental independence; broad choice with prohibitively difficult migration shows availability but not practical control. A low spending share would not automatically mean failure — it could mean competition drove the same useful work to a lower price. The serious warning is an ecosystem creating enormous customer value while failing to finance the production and maintenance of its shared foundations.
Section 17When useful intelligence is available from more sources, access to one particular model becomes less exclusive and the difficult work moves toward choosing, operating, adapting, evaluating, and applying those models. Open models do not need to capture most of AI industry revenue to change the market — they only need to make intelligence sufficiently substitutable that the scarcity premium of the model layer begins to compress. Once that happens, value shifts toward whatever remains scarce: infrastructure, proprietary context, evaluation data, trusted workflows, customer relationships, distribution, and accumulated operating knowledge. Openness increases the number of parties that can participate; it does not determine which layer ultimately keeps the margin.
Section 18Open-weight AI has crossed an important production threshold: on major gateways it carries a substantial share of inference, and around it a wider industrial system is forming whose incentives are no longer concentrated in one place. But operational escape velocity and developmental escape velocity must remain separate — the ecosystem is increasingly capable of keeping existing models useful when one participant changes direction, and has not yet proved enough independent organizations can continuously finance and train their successors. For enterprises the implication is immediate: do not confuse token volume with value, downloadable weights with independence, customization with ownership, or cheap inference with cheap work. Build what survives model churn — context, evaluations, workflow, outcome history, operating knowledge, and the improvement loop. Cheap intelligence is becoming easier to obtain; the durable advantage is learning how to turn it into useful work, and retaining the ability to keep improving that work when the model underneath changes.
Section 19The argument rests on twelve mental models, each travelling beyond open models: Three Measurements; Follow the Inference Dollar; Openness as a Set of Rights; Complements Fund the Commons; Two Escape Velocities; The Supply-Chain Map; Pacing at the Frontier, Diffusion at the Base; Credible Optionality; Cost per Accepted Result; You Cannot Own What You Cannot Measure; Dependency Moves, It Does Not Vanish; and The Scarcity Premium Compresses.