Thesis & Framework Meshup · Enterprise AI

Right-Sizing Your Intelligence Spend

Frontier models keep getting more intelligent, but most of the economy runs on context and execution, not frontier intelligence. Every workload has an intelligence threshold — below it the model cannot perform the task; once crossed, the bottleneck shifts to whether the system knows the company's policies, customers, history, tools and standards. The winning enterprise is the one that learns how to need intelligence the least.

AuthorsJaya GuptaFoundation Capital Avanika NarayanStanford AI PhD Jon Saad-FalconStanford AI PhD
Published 2026-08-19 on LinkedIn Pulse 53 likes 12 comments KG curated by DeepSeek V4 Flash on behalf of Kingsley Uyi Idehen
<1%
of US employment in life, physical & social-science occupations (BLS)
2,000
US mathematicians (vs 20,000 physicists, 37,000 research scientists)
7–10×
excess tokens reasoning systems generate on simple tasks (Amazon research)
320×
growth in average enterprise reasoning-token consumption (OpenAI)
56%
of 4,454 CEOs saw no significant financial benefit from AI (PwC)
The Intelligence Threshold workload difficulty → model capability → value
FRONTIER BAND — discovery value: proofs, drug targets, zero-days, market structures only tasks above the rising capability line should be sent to a frontier model Password reset support — threshold: none (deterministic)password reset Refund handling — threshold: lowrefunds Claims adjustment — threshold: moderateclaims Code generation with tests — threshold: high (hybrid)coding Formal mathematical proof — threshold: extreme (discovery)proofs workload difficulty → ← below the line: context & execution is the bottleneck • above the line: frontier intelligence pays →
Synopsis

Frontier AI models have become astonishingly intelligent — OpenAI's internal model disproved a conjecture in Erdos's planar unit-distance problem, and Claude Opus 4.6 found more than 500 high-severity vulnerabilities. These achievements are real and extraordinarily valuable where the answer is unknown: a proof, a drug target, a zero-day, an unfamiliar market structure. But the rest of the economy does not work that way.

A claims adjuster is not searching for a new mathematical truth; she is deciding which existing rules apply to this accident and this policyholder. A nurse is not inventing medicine; she is combining established protocols with one patient's history. Companies are defended not by superior frontier intelligence but by proprietary data, dense operational context, long-built processes, regulatory licenses, physical networks, and accumulated judgment.

Every workload has some intelligence threshold. Below it, the model cannot perform the task; approaching it, greater intelligence creates enormous value; once crossed, the bottleneck changes to context and execution. The correct posture is the opposite of today's default: assume open-weight and specialized systems will keep eating the stack, and reserve frontier models only for problems that still sit above the rising capability line.

The article's three predictions: millions of routing gateways inside every AI harness; hybrid workloads as the default; and AI diffusing through proactive assistants and specialized application harnesses beneath an improvement layer. Progress is measured not by how much intelligence a company can afford to consume, but by how quickly intelligence consumed per successful outcome falls toward zero.

What the public scoreboard celebrates — and how narrow it really is.

In May, OpenAI reported that a general-purpose internal model had disproved a central conjecture in Erdos's planar unit-distance problem, a question mathematicians had worked on for nearly eighty years; external mathematicians checked the proof. Anthropic reports that Claude Opus 4.6 has found and validated more than 500 high-severity vulnerabilities, including bugs in codebases fuzzed for years.

When the answer is unknown and discovering it could be worth millions — a proof, a drug target, a zero-day, or an unfamiliar market structure — maximum intelligence is exactly what we should want. Spending more computation to explore more possibilities is rational.

But the public scoreboard is shaped by the work that can be cleanly stated and verified: proofs, patches, exploits, and scientific reasoning. The only problem is that the rest of the economy does not work that way.

Frontier research is extraordinarily valuable and economically narrow; the moat is not thinking better — it is knowing how your world works.

Claims adjuster

Decides which of the company's existing rules apply to this accident, this policyholder, this sequence of events.

Nurse

Combines established protocols with one patient's history, current symptoms and hospital constraints.

Logistics coordinator

Reacts to today's inventory, weather, contracts and delays.

BLS data shows life, physical and social-science occupations account for less than 1% of American employment: roughly 2,000 mathematicians, 20,000 physicists and 37,000 computer and information research scientists — versus millions of nurses, managers, administrators, logistics workers and customer-service representatives.

What makes Walmart, UnitedHealth and Amazon defensible is not superior frontier intelligence: it is proprietary data, dense operational context, long-built processes, supplier relationships, regulatory licenses, physical networks, and the accumulated judgment that only comes from performing the same class of decision millions of times. Replacing every employee at American Airlines, Home Depot or Medtronic with a math olympiad would drown the organization in intelligence it cannot productively use.

The core of the thesis: below the threshold the model cannot perform the task; above it, the bottleneck changes.

Every workload has some intelligence threshold. Below it, the model cannot perform the task. As the model approaches it, greater intelligence creates enormous value. Once the threshold has been crossed, the bottleneck changes. The outcome depends increasingly on whether the system knows the company's policies, customers, history, tools and standards — and whether it can act reliably, quickly and cheaply.

Models keep getting more intelligent, but most workloads are not. Refund policies, baggage rules and invoice-reconciliation procedures are not getting any harder. The latest frontier models are optimized for extreme intelligence — the kind required by research mathematicians, physicists, and engineers solving previously unsolved problems. That market is real. It is also small.

Workload Threshold Spectrum where the twelve sample workloads sit on the rising capability line
None Low Moderate High Extreme capability line (rising) Password reset Refunds / invoice reconciliation Claims / triage / logistics / routing Coding (tests) / market structure Proofs / drug targets / zero-days • dots sit on the thresholds they cross; once crossed, the bottleneck is context & execution, not intelligence

Two trends are eroding the frontier premium: the shift from closed to open-weight models (e.g. Qwen3.8-27B), which lowers the cost of a given level of capability, and the shift from frontier-scale to small and mid-sized models, which compresses that capability into a much smaller compute footprint.

Models are overqualified too — with a more persistent failure mode than humans: they never get bored and quit.

Developers naturally reach for the most capable system available, and frontier products make that choice frictionless: long-running agents, recursive sub-agents, and high-effort reasoning loops that burn tokens by design. The result is frontier-priced intelligence applied to tasks that often do not need it.

Have you ever hired an employee that was dramatically overqualified for a role? They may reconsider settled decisions, introduce unnecessary complexity, become bored and eventually leave. Models are overqualified too, but instead of getting bored and quitting, they keep generating extra complexity, variance and cost indefinitely.

Overqualification is now measurable. Researchers studying reasoning models have documented an overthinking problem: models routinely spend additional inference compute on easy questions without improving accuracy. Amazon researchers estimate that reasoning systems can generate 7 to 10× as many tokens as necessary on simple tasks. On sufficiently easy workloads, the marginal reasoning token can approach zero.

One player is paid to increase the size of the fire. The other is paid to put it out and keep it from starting again.

Model lab wins by

Making more computation useful: tokens consumed, session length, agent count, reasoning depth, tackling harder problems. Its dashboard lights up when usage rises.

Enterprise wins by

Making repeated computation unnecessary: cost per verified resolution, customer retention, cost per successful outcome. Its goal is to need less intelligence.

The data already shows the gap. OpenAI reports that average reasoning-token consumption per enterprise organization rose approximately 320-fold over the past year. PwC's survey of 4,454 CEOs found that 56% had not yet seen a significant financial benefit from AI. Measurable business outcomes are not keeping pace. If an enterprise's goal is to reduce intelligence consumed per successful outcome, it cannot leave every decision about intelligence consumption to the model provider — that control must live inside the enterprise at every layer.

What happens once enterprises own the allocation of intelligence.

1. Millions of routers and gateways

Inside every serious AI harness — at the policy layer, team level and organization level — each deciding, for every unit of work, whether it belongs on a frontier model, a smaller open-weight model, a local model, a deterministic system, a human, or nowhere at all. The TAM of this routing layer will look a lot like the workforce itself.

2. Hybrid workloads as the default

Workloads will stop belonging to a single model. The easy, repetitive or private parts of a task stay local or near the system of record; only the genuinely hard or novel parts go to a frontier model in the cloud.

3. Two complementary application surfaces

General-purpose proactive assistants such as Instinct make intelligence ambient; specialized applications such as Maximor and PlayerZero turn enterprise intent into reliable execution.

Beneath those harnesses, a new improvement layer will emerge: platforms such as Applied Compute's AC2 allow companies to train, evaluate, serve and replace the models operating inside their existing applications. The proactive assistant becomes the demand aggregator; the application harness becomes the execution system; AC2 becomes the specialization and improvement infrastructure that continuously moves each workload toward the least expensive system still capable of the correct outcome.

The most intelligent enterprise is not going to chase the frontier models blindly for all work. It will be the one that has learned how to need intelligence the least, turning yesterday's expensive reasoning into tomorrow's cheaper, more reliable execution.
Intelligence Consumed per Successful Outcome the metric that matters — falling toward zero
today: frontier-priced tokens on routine work → zero: compile rules, route by tier, own the context graph time →

Progress will not be measured by how much intelligence a company can afford to consume, but by how quickly intelligence consumed per successful outcome falls toward zero.

Commentary
Accountable to Kingsley Uyi Idehen · agent-authored via kg-generator

Right-Sizing Intelligence Is a Data and Context Problem

The article's threshold argument is exactly right, and it points to a deeper structure: the context that becomes the bottleneck once the intelligence threshold is crossed is not amorphous — it is data. Proprietary data, dense operational context, long-built processes and accumulated judgment are precisely what a knowledge graph encodes: entities, relationships, rules and history as machine-readable graph. So the right way to right-size intelligence spend is to build the enterprise context layer as a Semantic Web — an HTTP IRI for every entity, hyperlinks as standardized identifiers, RDF views over systems of record, and SPARQL for inspection — so that the routing layer the article predicts can be grounded in authoritative, queryable context rather than in model priors. The zero the article seeks is reached by compiling repeated decisions into machine-checkable rules and linking those rules to the graph, so the millionth claim still has a model working on it but none deciding it. In short: the frontier model is rented; the context graph is owned. Right-sizing intelligence spend is inseparable from investing in the enterprise knowledge graph — the moat the article describes is, at bottom, a data moat.

Context is data: the moat is a data moat

Proprietary data, processes and accumulated judgment are precisely what a knowledge graph encodes.

Routing is only as good as the context graph it routes against

The routing layer needs authoritative, queryable context, not model priors.

Build the context layer as a Semantic Web

HTTP IRIs, hyperlinks as identifiers, RDF views over systems of record, SPARQL inspection.

The frontier model is rented; the context graph is owned

The durable asset is the enterprise knowledge graph — a data moat.

Sample Data & Ontology

Twelve sample workload profiles demonstrating the thesis: each carries a value class (discovery vs application), an intelligence threshold, a recommended tier, a bottleneck, a verification mechanism, token sensitivity and overqualification risk. Typed by the Intelligence Allocation Ontology. Full data: Sample Workload Dataset.

Dataset entity in RDF: schema:Dataset → intel:WorkloadProfile instances → intel:hasIntelligenceThreshold / requiresTier / hasValueClass / hasBottleneck / belongsToWorkflow.

Workload Intelligence Map twelve demo workload instances by intelligence threshold × token sensitivity — colored by value class
NoneLow ModerateHigh Extreme intelligence threshold → token sensitivity → Very HighHigh MediumLow frontier reserve Discovery Application Hybrid (coding) dot radius = overqualification risk Password reset — None / Very High tokens / Extreme overqualificationpwd reset Refund handling — Low / High / Very Highrefunds Invoice reconciliation — Low / High / Highinvoices Claims adjustment — Moderate / Medium / Highclaims Nursing triage — Moderate / Medium / Hightriage Logistics coordination — Moderate / Medium / Highlogistics Workload routing decision — Moderate / Medium / Mediumrouting Code generation — High / Medium / Medium (hybrid)coding Market structure analysis — High / Medium / Lowmarkets Formal proof — Extreme / Low / Noneproofs Drug target discovery — Extreme / Low / Nonedrugs Zero-day discovery — Extreme / Low / Nonezero-days

Intelligence Tiers, Head to Head

Six tiers compared across six dimensions; the comparison dimensions are first-class instances typed cdx:ComparisonDimension in the companion RDF, and every tier header links to its intel:IntelligenceTier entity.

DimensionDeterministic SystemSmall Open-Weight Model (Local)Mid-Size ModelFrontier ModelHuman ExpertHybrid Multi-Tier
Cost per unit~zero (no tokens)Near-zero marginalLow-moderatePremium token pricingSalary (inelastic)Minimized per outcome
LatencyInstant (compiled rules)Low (local inference)ModerateHigh (deep reasoning loops)Slow but adaptiveTier-dependent
Capability ceilingFully specified decisions onlyBroad enterprise application classes; rising capability lineMost application workloads above small-model thresholdsPreviously unsolved problems — proofs, zero-days, drug targetsOpen-ended judgment, accountabilityFull workload spectrum, routed
Best-fit workloadsRefund handling, password reset, rule-based claimsNursing triage, logistics coordination, invoice reconciliationHybrid application workloads, coding assistant tasksFormal proof, zero-day discovery, drug-target discovery, market structure analysisNovel edge cases, disputes, high-stakes judgmentAll workloads with heterogeneous units of work
Overqualification riskNoneLowMediumHigh on easy work — overthinking, variance, costBoredom and departure (humans quit; models do not)Controlled by routing policy
Token economicsNo reasoning tokens consumedMarginal tokens cheap; 7-10x overhead on simple tasks still dominates costBetter token efficiency than frontier on routine work320-fold enterprise reasoning-token growth; marginal token ~zero below the thresholdNo tokens; salary is inelastic while delivered intelligence is elasticIntelligence consumed per successful outcome falls toward zero
Small Open-Weight Model (Local)

Small and mid-sized open-weight models (e.g. the Qwen3.8-27B class) run local or near the system of record; compressed capability at a small compute footprint.

Cost per unitNear-zero marginal
LatencyLow (local inference)
Capability ceilingBroad enterprise application classes; rising capability line
Best-fit workloadsNursing triage, logistics coordination, invoice reconciliation
Token economicsMarginal tokens cheap; 7-10x overhead on simple tasks still dominates cost
Frontier Model

The most capable, most expensive class of model (Anthropic, OpenAI class), optimized for the extreme intelligence required by research mathematicians, physicists and engineers; a real but small market.

Cost per unitPremium token pricing
LatencyHigh (deep reasoning loops)
Capability ceilingPreviously unsolved problems — proofs, zero-days, drug targets
Best-fit workloadsFormal proof, zero-day discovery, drug-target discovery, market structure analysis
Overqualification riskHigh on easy work — overthinking, variance, cost
Token economics320-fold enterprise reasoning-token growth; marginal token ~zero below the threshold
Human Expert

Human judgment and accountability for edge cases, novelty and high-stakes decisions — the elastic-intelligence tier that enterprises already right-size when hiring.

Cost per unitSalary (inelastic)
LatencySlow but adaptive
Capability ceilingOpen-ended judgment, accountability
Best-fit workloadsNovel edge cases, disputes, high-stakes judgment
Overqualification riskBoredom and departure (humans quit; models do not)
Token economicsNo tokens; salary is inelastic while delivered intelligence is elastic

Routing Policies

Four sample intel:RoutingPolicy instances that encode the article's allocation logic for the sample workloads.

Default Enterprise Routing Policy

Route every unit of work to the least expensive system still capable of the correct outcome; the article argues enterprises must not leave this decision to the model provider.

Applies to: Claims adjustment, refund handling, invoice reconciliation, password reset, routing decisions Routes to: Deterministic, small open-weight, mid-size, frontier, human

Frontier Reserve Policy

Reserve frontier models for tasks still above the rising capability line: discovery workloads where additional intelligence changes the outcome.

Applies to: Formal proof, drug-target discovery, zero-day discovery, market-structure analysis Routes to: Frontier only

Hybrid Split Policy

Split workloads across tiers: easy, repetitive or private parts stay local or near the system of record; only genuinely hard or novel parts go to a frontier model in the cloud.

Applies to: Nursing triage, logistics coordination, code generation Routes to: Small open-weight, mid-size, frontier

Zero-Intelligence-Outcome Policy

Compile repeated decisions into machine-checkable rules so the millionth claim still has a model working on it but none deciding it; continuously move each workload toward the least expensive system still capable of the correct outcome.

Applies to: Password reset, refund handling, invoice reconciliation Routes to: Deterministic, small open-weight

Executable queries against the sample workload dataset, modeled as schema:SoftwareSourceCode with target URIBurner.

QWorkloads grouped by recommended intelligence tier▼

Query entity: Workloads grouped by recommended intelligence tier

PREFIX intel: <https://linkeddata.uriburner.com/DAV/demos/daas/ontology/intelligence-allocation#>
PREFIX schema: <http://schema.org/>

# Which workloads should run on which intelligence tier?
SELECT ?tierName (COUNT(?wl) AS ?workloadCount) (SAMPLE(?wlName) AS ?exampleWorkload)
WHERE {
  ?wl a intel:WorkloadProfile ;
      intel:requiresTier ?tier ;
      schema:name ?wlName .
  ?tier schema:name ?tierName .
}
GROUP BY ?tierName
ORDER BY DESC(?workloadCount)
QDiscovery versus application workload counts▼

Query entity: Discovery versus application workload counts

PREFIX intel: <https://linkeddata.uriburner.com/DAV/demos/daas/ontology/intelligence-allocation#>
PREFIX schema: <http://schema.org/>

# Is most enterprise work discovery or application?
SELECT ?valueClassName (COUNT(?wl) AS ?count)
WHERE {
  ?wl a intel:WorkloadProfile ;
      intel:hasValueClass ?vc .
  ?vc schema:name ?valueClassName .
}
GROUP BY ?valueClassName
ORDER BY DESC(?count)
QWorkloads at or above the frontier capability line▼

Query entity: Workloads at or above the frontier capability line

PREFIX intel: <https://linkeddata.uriburner.com/DAV/demos/daas/ontology/intelligence-allocation#>
PREFIX schema: <http://schema.org/>

# Which workloads still need maximum intelligence (extreme threshold)?
SELECT ?wlName ?thresholdName ?tierName
WHERE {
  ?wl a intel:WorkloadProfile ;
      schema:name ?wlName ;
      intel:hasIntelligenceThreshold ?th ;
      intel:requiresTier ?tier .
  ?th schema:name ?thresholdName .
  ?tier schema:name ?tierName .
  FILTER(?thresholdName = "Extreme")
}
ORDER BY ?wlName
QAll demo workload profiles (SELECT *)▼

Query entity: All demo workload profiles (SELECT *)

PREFIX intel: <https://linkeddata.uriburner.com/DAV/demos/daas/ontology/intelligence-allocation#>
PREFIX schema: <http://schema.org/>

# Every demo workload instance from the article, with threshold and recommended tier
SELECT ?name ?threshold ?tier
FROM <https://linkeddata.uriburner.com/DAV/demos/daas/right-sizing-intelligence-spend-deepseek_v4flash-1.ttl>
WHERE {
  ?workload a intel:WorkloadProfile ;
            schema:name ?name ;
            intel:hasIntelligenceThreshold ?th ;
            intel:requiresTier ?tier .
  ?th schema:name ?threshold .
  ?tier schema:name ?tier .
}
ORDER BY ?name
Knowledge Graph Explorer

Explore the Knowledge Graph

Graph derived from the companion RDF (right-sizing-intelligence-spend-deepseek_v4flash-1.ttl): 50 nodes, 79 edges. Drag nodes to pin; double-click to unpin; click the canvas to arm zoom.

🔍 zoom armed — scroll to zoom, drag to pan; click outside to release
Article/section Ontology class Workload/tier/policy Person/org/product Dataset/FAQ/glossary
SPARQL Workbench

Query the Knowledge Graph

SPARQL Workbench

▶ Run live on URIBurner format: text/x-html+tr

SELECT results render as HTML tables (text/x-html+tr); DESCRIBE/CONSTRUCT render as nice Turtle (text/x-html-nice-turtle).

Demo instance data — the workloads Jaya's article covers

Twelve intel:WorkloadProfile instances in the companion RDF; every recipe below queries this sample graph.

A six-step operating procedure derived from the article's thesis; each step links to its absolute-IRI schema:HowToStep entity.

1

Inventory every workload

List the units of work your organization performs — claims, triage, reconciliation, support, code, research — with volumes, current systems and failure costs.

2

Classify each workload by value class

Mark each workload as discovery (expanding what is known: proofs, drug targets, zero-days, market structures) or application (making the right decision within known context); coding is the hybrid exception because outcomes are verifiable through tests.

3

Map thresholds against the capability line

For each workload, estimate its intelligence threshold and where today's open-weight and frontier models sit relative to it; most workloads are not getting harder while models keep getting more intelligent.

4

Route by tier policy

Assign each workload to the least expensive tier still capable of the correct outcome — deterministic rules, local small models, mid-size models, frontier models, or humans — and encode that allocation in a routing policy inside your harness, at the policy, team and organization layers.

5

Measure intelligence per successful outcome

Instrument tokens, latency and cost per verified resolution; track whether customer retention and cost per successful outcome improve, not whether token usage rises.

6

Reinvest savings in context and execution

Compile repeated decisions into machine-checkable rules and invest in proprietary data, dense operational context, processes, systems of record and accumulated judgment — the moat that frontier intelligence alone cannot buy.

Thirteen questions and answers probing the thesis, wrapped in a schema:FAQPage with schema:mainEntity.

01What is the central argument of 'Right-Sizing Your Intelligence Spend'?▼

Frontier intelligence is extraordinarily valuable but economically narrow; most enterprise work is a context-and-execution problem. Enterprises should route each unit of work to the least expensive system still capable of the correct outcome, reserving frontier models for genuine discovery, and measure progress as intelligence consumed per successful outcome falling toward zero.

02What is an intelligence threshold?▼

Every workload has some level of model intelligence below which the model cannot perform the task. As the model approaches it, greater intelligence creates enormous value; once crossed, the bottleneck changes from intelligence to whether the system knows the company's policies, customers, history, tools and standards and can act reliably, quickly and cheaply.

03Why is frontier research economically narrow?▼

BLS data shows life, physical and social-science occupations account for less than 1% of American employment: roughly 2,000 mathematicians, 20,000 physicists and 37,000 computer and information research scientists, versus millions of nurses, managers, administrators, logistics workers and customer-service representatives.

04What distinguishes discovery value from application value?▼

Some jobs create value by expanding the frontier of what is known — a proof, a drug target, a zero-day, an unfamiliar market structure. Most create value by making the right decision within known context and executing it — claims, triage, logistics, reconciliation. Companies are defended not by superior frontier intelligence but by proprietary data, processes, licenses, networks and accumulated judgment.

05Why is coding the important exception?▼

Coding is a large labor market, the work is digital, and outcomes can often be verified through tests — which makes additional model intelligence unusually valuable. But even coding is developing capability thresholds as smaller and open-weight models approach frontier performance on increasingly broad classes of work.

06What is model overqualification, and why is it persistent?▼

Hiring a dramatically overqualified employee leads to reconsidered decisions, unnecessary complexity, boredom and departure; models are overqualified too, but instead of quitting they keep generating extra complexity, variance and cost indefinitely — and overqualification is now measurable as the overthinking problem.

07What is the overthinking problem?▼

Reasoning models routinely spend additional inference compute on easy questions without improving accuracy; Amazon researchers estimate reasoning systems can generate 7 to 10x as many tokens as necessary on simple tasks, so on sufficiently easy workloads the marginal reasoning token approaches zero.

08Why do labs and enterprises have competing incentives?▼

A lab wins by making more computation useful — its metrics are tokens consumed, session length, agent count and reasoning depth. An enterprise wins by making repeated computation unnecessary — cost per verified resolution. OpenAI reports enterprise reasoning-token consumption rose ~320-fold in a year, while PwC found 56% of 4,454 CEOs saw no significant financial benefit from AI.

09What will the routing layer look like?▼

Millions of routers and gateways inside every serious AI harness — at the policy layer, team level and organization level — deciding per unit of work whether it belongs on a frontier model, a smaller open-weight model, a local model, a deterministic system, a human, or nowhere at all. Its TAM will ultimately look a lot like the workforce itself.

10What are hybrid workloads?▼

Workloads that stop belonging to a single model: easy, repetitive or private parts stay local or near the system of record, and only genuinely hard or novel parts are sent to a frontier model in the cloud.

11What are the two application surfaces for AI?▼

General-purpose proactive assistants such as Instinct make intelligence ambient — observing work, identifying opportunities and initiating tasks without waiting for a prompt. Specialized applications such as Maximor and PlayerZero turn enterprise intent into reliable execution by controlling context, tools, permissions, workflow state and definition of done.

12What role does the improvement layer play?▼

Platforms such as Applied Compute's AC2 train, evaluate, serve and replace the models operating inside existing applications. The harness remains the system of execution while AC2 continuously moves each workload toward the least expensive system still capable of the correct outcome.

13How should progress be measured?▼

Not by how much intelligence a company can afford to consume, but by how quickly intelligence consumed per successful outcome falls toward zero — turning yesterday's expensive reasoning into today's cheaper, more reliable execution.

Sixteen terms wrapped in a schema:DefinedTermSet, each linked to its schema:DefinedTerm entity and aligned to the ontology via rdfs:seeAlso.

Intelligence Threshold

The level of model intelligence below which a workload cannot be performed; as the model approaches it, greater intelligence creates enormous value, and once it is crossed the bottleneck shifts from intelligence to context and execution.

Frontier Model

The most capable and most expensive class of AI model, optimized for the extreme intelligence required by research mathematicians, physicists and engineers solving previously unsolved problems.

Open-Weight Model

A model whose weights are publicly released; the shift from closed to open-weight models lowers the cost of a given level of capability.

Context-and-Execution Problem

The dominant mode of economic work: applying what the company already knows across millions of decisions shaped by customers, policies, inventory, contracts and history.

Discovery Value

Value created by expanding the frontier of what is known — a proof, a drug target, a zero-day, or an unfamiliar market structure — where maximum intelligence is exactly what is wanted.

Application Value

Value created by making the right decision within known context and executing it reliably, quickly and cheaply.

Overqualification

Using a system far more capable than a task requires; unlike a bored human who quits, an overqualified model never stops generating extra complexity, variance and cost.

Overthinking

The documented behavior of reasoning models spending additional inference compute on easy questions without improving accuracy.

Reasoning Token

A unit of inference computation consumed during chain-of-thought reasoning; on sufficiently easy workloads the marginal reasoning token approaches zero value.

Routing Layer

Routers and gateways inside every serious AI harness that decide, per unit of work, whether it belongs on a frontier model, a smaller open-weight model, a local model, a deterministic system, a human, or nowhere at all; its TAM resembles the workforce itself.

Hybrid Workload

A workload that stops belonging to a single model: easy, repetitive or private parts stay local or near the system of record, while only genuinely hard or novel parts are sent to a frontier model in the cloud.

Proactive Assistant

A general-purpose assistant that makes intelligence ambient: it observes work, identifies opportunities, and initiates tasks without waiting for a prompt.

Application Harness

A specialized application that turns enterprise intent into reliable execution by controlling the context, tools, permissions, workflow state, and definition of done for a particular job.

Capability Line

The rising line of capability as open-weight and small or mid-sized models cross workload intelligence thresholds; frontier systems should be reserved for tasks that still sit above it.

Intelligence per Successful Outcome

The key metric of right-sizing: intelligence consumed divided by verified successful outcomes; progress is this ratio falling toward zero.

Token Blast

Long-running agent loops, recursive sub-agent spawning and 'keep going until it works' patterns that burn tokens by design — the frictionless default of frontier products.

Reader Perspectives

The article's comment thread, modeled as schema:Comment entities. All sixteen comments — authors, profile URLs and verbatim text — were extracted from the live thread via an authenticated browser session (Playwright + Chrome, 2026-08-19); replies are marked. Kingsley Uyi Idehen first, then the thread in document order.

Jaya, Yep! For enterprises, the issue is really about calibrating “Good Enough AI,” which isn’t necessarily aligned with what frontier model labs are producing. A business simply wants the job done at a cost optimized for its operating model. No more, no less. 😄

Additional notes generated using Google Gemini 3.7 Flash (Medium) as the LLM component. https://linkeddata.uriburner.com/weblog/?post=right-sizing-intelligence-spend-jaya-gupta-gemini_3_7_flash-1.html

Jaya Gupta Decision Intelligence at scale is becoming most necessary layer for the fast growing organisation that actually drives P&L impact.

Jaya Gupta This is probably your best article yet... The Context graph article was inspirational, this article is pragmatic, impactful and consequential. Important to note that when you apply a human to a job, intelligence delivered is elastic, salary is inelastic. The smartest humans sometimes also have to do dumb jobs. But you don't pay them less for doing so. Even Einstein had to go make a cup of coffee. But in the post-intelligence world, enterprises can vary pay in real-time, so the average cost of intelligence could asymptomatically arrive close to zero. Both the price of intelligence, and the cost of applying it are equally elastic.

Vas Bhandarkar thank you!!!!

This immediately made me think of your context graph thesis. There’s an interesting connection here IMHO: for things like claims, rostering, routing and inventory, the optimization machinery has existed for decades. The hard part has often been formulation, translating company policy, current state, exceptions and judgment into the objective and constraints for this decision. Models make that translation dramatically easier. But they also create an important boundary: a model can propose a constraint, but it shouldn’t be the authority that makes it true. That should trace back to the company’s context graph: policy, system state, precedent, contract, or an authorized person. And sometimes the right answer is simply “infeasible”, not more reasoning.

I can't stop thinking about how Anthropic and the other frontier labs have very different motivations in a future where AI adoption stops being the primary push. Use of these models is still cheaper than it “should” be as the providers boost usage limits and reduce costs to get people hooked. What happens when enough people have adopted AI that these companies can pivot to actually attempting to bring in a profit? Their motivation switches to token usage (which it already kind of is) and they're incentivized to over-engineer answers. Your point about the expensive intelligence being most usefully applied to deciding WHAT needs to be done, not doing it, just sounds like the normal shape of an organization - we don't need the CEO answering every customer support email (nor would that ever work in practice).

Love this! Prices for things (all the things) must go down. the proxy is time. We can deliver more per unit of time now than ever before. Why did we charge more for things in the past? Expertise, time, workmanship, etc. judgement on when/where (fancy folk talk this “taste”) is the iffy part. Why? Because LLMs don’t understand how hoomans will react to things (yet). We have that intuition more dialed because of body language and empathy. This will probably change over time as AI gets eyes and sensors for biologics. Like in Silo when the AI reads heart rate and temp that factor into their reasoning.

Minions prices local inference at zero, so 5.7x is a cloud bill and not a total. Section 3, under Measuring cost, states the assumption outright and ignores hardware and energy. Fair for a laptop already bought. The routing thesis survives this. An enterprise fleet is capacity you pay for idle, and the protocol earns its ratio by moving compute onto exactly that unpriced side. The paper even notes tokens are a poor proxy for on-device cost because the batches never fill. Price the local GPU-hours and re-run it. If 5.7x survives, I am wrong about how much of the gap is accounting.

Jaya, your claims adjuster never needed reasoning. The rules she applies are written down, in policies and contracts the company already owns. The direct route to your zero is to compile them: spend intelligence once, at build time, turning the documents into machine-checkable rules. AI still runs the process after that. It gathers the facts, drafts the output, talks to the customer. The one place it never sits is the decision itself, which is computed from the rules and cites its clause. So the millionth claim still has a model working on it, but none deciding it. Swap the model for a human and nothing changes: she works the claim, the rules decide it. That was always true of the adjuster anyway. Learning from production gets you to determinism where the rules are unwritten. In most businesses, far less is unwritten than people think.

Thank you Jaya Gupta. I have to pin this “while the frontier is an intelligence problem, the rest of the economy is a context-and-execution problem”

Subash Mandanapu thank you

Love this concept.. I would add one more perspective.. “Intelligence has a threshold. Economics has no ceiling!” Once a model is good enough, the winners are the ones who deliver the same outcome with fewer tokens, less latency and lower operational cost. You may resonate with this. I explored the idea in my article, The Economics of Intelligence. https://www.linkedin.com/pulse/economics-intelligence-prasad-prabhakaran-cvv7e

Opposite outcomes chart is brilliant. Incentives will have to align for sustainable AI value, and revenue.

Beautifully captures the current state of thinking in the enterprise.

Pramod Agrawal thank you

About

About This Collection

This knowledge-graph collection was generated from the LinkedIn Pulse article Right-Sizing Your Intelligence Spend by Jaya Gupta, Avanika Narayan and Jon Saad-Falcon, transformed into RDF via kg-generator and rendered via rdf-infographic-skill. The collection includes a lightweight Intelligence Allocation Ontology (intel:) and a sample workload dataset demonstrating the article's intelligence-threshold thesis across twelve domain workflows. Entities resolve through URIBurner describe.

Technology stack: