Right-Sizing Your Intelligence Spend
Frontier models keep getting more intelligent, but most of the economy runs on context and execution, not frontier intelligence. Every workload has an intelligence threshold — below it the model cannot perform the task; once crossed, the bottleneck shifts to whether the system knows the company's policies, customers, history, tools and standards. The winning enterprise is the one that learns how to need intelligence the least.
Frontier AI models have become astonishingly intelligent — OpenAI's internal model disproved a conjecture in Erdos's planar unit-distance problem, and Claude Opus 4.6 found more than 500 high-severity vulnerabilities. These achievements are real and extraordinarily valuable where the answer is unknown: a proof, a drug target, a zero-day, an unfamiliar market structure. But the rest of the economy does not work that way.
A claims adjuster is not searching for a new mathematical truth; she is deciding which existing rules apply to this accident and this policyholder. A nurse is not inventing medicine; she is combining established protocols with one patient's history. Companies are defended not by superior frontier intelligence but by proprietary data, dense operational context, long-built processes, regulatory licenses, physical networks, and accumulated judgment.
Every workload has some intelligence threshold. Below it, the model cannot perform the task; approaching it, greater intelligence creates enormous value; once crossed, the bottleneck changes to context and execution. The correct posture is the opposite of today's default: assume open-weight and specialized systems will keep eating the stack, and reserve frontier models only for problems that still sit above the rising capability line.
The article's three predictions: millions of routing gateways inside every AI harness; hybrid workloads as the default; and AI diffusing through proactive assistants and specialized application harnesses beneath an improvement layer. Progress is measured not by how much intelligence a company can afford to consume, but by how quickly intelligence consumed per successful outcome falls toward zero.
What the public scoreboard celebrates — and how narrow it really is.
In May, OpenAI reported that a general-purpose internal model had disproved a central conjecture in Erdos's planar unit-distance problem, a question mathematicians had worked on for nearly eighty years; external mathematicians checked the proof. Anthropic reports that Claude Opus 4.6 has found and validated more than 500 high-severity vulnerabilities, including bugs in codebases fuzzed for years.
When the answer is unknown and discovering it could be worth millions — a proof, a drug target, a zero-day, or an unfamiliar market structure — maximum intelligence is exactly what we should want. Spending more computation to explore more possibilities is rational.
But the public scoreboard is shaped by the work that can be cleanly stated and verified: proofs, patches, exploits, and scientific reasoning. The only problem is that the rest of the economy does not work that way.
Frontier research is extraordinarily valuable and economically narrow; the moat is not thinking better — it is knowing how your world works.
Claims adjuster
Decides which of the company's existing rules apply to this accident, this policyholder, this sequence of events.
Nurse
Combines established protocols with one patient's history, current symptoms and hospital constraints.
Logistics coordinator
Reacts to today's inventory, weather, contracts and delays.
BLS data shows life, physical and social-science occupations account for less than 1% of American employment: roughly 2,000 mathematicians, 20,000 physicists and 37,000 computer and information research scientists — versus millions of nurses, managers, administrators, logistics workers and customer-service representatives.
What makes Walmart, UnitedHealth and Amazon defensible is not superior frontier intelligence: it is proprietary data, dense operational context, long-built processes, supplier relationships, regulatory licenses, physical networks, and the accumulated judgment that only comes from performing the same class of decision millions of times. Replacing every employee at American Airlines, Home Depot or Medtronic with a math olympiad would drown the organization in intelligence it cannot productively use.
The core of the thesis: below the threshold the model cannot perform the task; above it, the bottleneck changes.
Every workload has some intelligence threshold. Below it, the model cannot perform the task. As the model approaches it, greater intelligence creates enormous value. Once the threshold has been crossed, the bottleneck changes. The outcome depends increasingly on whether the system knows the company's policies, customers, history, tools and standards — and whether it can act reliably, quickly and cheaply.
Models keep getting more intelligent, but most workloads are not. Refund policies, baggage rules and invoice-reconciliation procedures are not getting any harder. The latest frontier models are optimized for extreme intelligence — the kind required by research mathematicians, physicists, and engineers solving previously unsolved problems. That market is real. It is also small.
Two trends are eroding the frontier premium: the shift from closed to open-weight models (e.g. Qwen3.8-27B), which lowers the cost of a given level of capability, and the shift from frontier-scale to small and mid-sized models, which compresses that capability into a much smaller compute footprint.
Models are overqualified too — with a more persistent failure mode than humans: they never get bored and quit.
Developers naturally reach for the most capable system available, and frontier products make that choice frictionless: long-running agents, recursive sub-agents, and high-effort reasoning loops that burn tokens by design. The result is frontier-priced intelligence applied to tasks that often do not need it.
Have you ever hired an employee that was dramatically overqualified for a role? They may reconsider settled decisions, introduce unnecessary complexity, become bored and eventually leave. Models are overqualified too, but instead of getting bored and quitting, they keep generating extra complexity, variance and cost indefinitely.
Overqualification is now measurable. Researchers studying reasoning models have documented an overthinking problem: models routinely spend additional inference compute on easy questions without improving accuracy. Amazon researchers estimate that reasoning systems can generate 7 to 10× as many tokens as necessary on simple tasks. On sufficiently easy workloads, the marginal reasoning token can approach zero.
One player is paid to increase the size of the fire. The other is paid to put it out and keep it from starting again.
Model lab wins by
Making more computation useful: tokens consumed, session length, agent count, reasoning depth, tackling harder problems. Its dashboard lights up when usage rises.
Enterprise wins by
Making repeated computation unnecessary: cost per verified resolution, customer retention, cost per successful outcome. Its goal is to need less intelligence.
The data already shows the gap. OpenAI reports that average reasoning-token consumption per enterprise organization rose approximately 320-fold over the past year. PwC's survey of 4,454 CEOs found that 56% had not yet seen a significant financial benefit from AI. Measurable business outcomes are not keeping pace. If an enterprise's goal is to reduce intelligence consumed per successful outcome, it cannot leave every decision about intelligence consumption to the model provider — that control must live inside the enterprise at every layer.
What happens once enterprises own the allocation of intelligence.
1. Millions of routers and gateways
Inside every serious AI harness — at the policy layer, team level and organization level — each deciding, for every unit of work, whether it belongs on a frontier model, a smaller open-weight model, a local model, a deterministic system, a human, or nowhere at all. The TAM of this routing layer will look a lot like the workforce itself.
2. Hybrid workloads as the default
Workloads will stop belonging to a single model. The easy, repetitive or private parts of a task stay local or near the system of record; only the genuinely hard or novel parts go to a frontier model in the cloud.
3. Two complementary application surfaces
General-purpose proactive assistants such as Instinct make intelligence ambient; specialized applications such as Maximor and PlayerZero turn enterprise intent into reliable execution.
Beneath those harnesses, a new improvement layer will emerge: platforms such as Applied Compute's AC2 allow companies to train, evaluate, serve and replace the models operating inside their existing applications. The proactive assistant becomes the demand aggregator; the application harness becomes the execution system; AC2 becomes the specialization and improvement infrastructure that continuously moves each workload toward the least expensive system still capable of the correct outcome.
The most intelligent enterprise is not going to chase the frontier models blindly for all work. It will be the one that has learned how to need intelligence the least, turning yesterday's expensive reasoning into tomorrow's cheaper, more reliable execution.
Progress will not be measured by how much intelligence a company can afford to consume, but by how quickly intelligence consumed per successful outcome falls toward zero.
Twelve sample workload profiles demonstrating the thesis: each carries a value class (discovery vs application), an intelligence threshold, a recommended tier, a bottleneck, a verification mechanism, token sensitivity and overqualification risk. Typed by the Intelligence Allocation Ontology. Full data: Sample Workload Dataset.
Dataset entity in RDF:
schema:Dataset → intel:WorkloadProfile instances →
intel:hasIntelligenceThreshold / requiresTier / hasValueClass / hasBottleneck / belongsToWorkflow.
Intelligence Tiers, Head to Head
Six tiers compared across six dimensions; the comparison dimensions are
first-class instances typed cdx:ComparisonDimension in the
companion RDF, and every tier header links to its
intel:IntelligenceTier entity.
| Dimension | Deterministic System | Small Open-Weight Model (Local) | Mid-Size Model | Frontier Model | Human Expert | Hybrid Multi-Tier |
|---|---|---|---|---|---|---|
| Cost per unit | ~zero (no tokens) | Near-zero marginal | Low-moderate | Premium token pricing | Salary (inelastic) | Minimized per outcome |
| Latency | Instant (compiled rules) | Low (local inference) | Moderate | High (deep reasoning loops) | Slow but adaptive | Tier-dependent |
| Capability ceiling | Fully specified decisions only | Broad enterprise application classes; rising capability line | Most application workloads above small-model thresholds | Previously unsolved problems — proofs, zero-days, drug targets | Open-ended judgment, accountability | Full workload spectrum, routed |
| Best-fit workloads | Refund handling, password reset, rule-based claims | Nursing triage, logistics coordination, invoice reconciliation | Hybrid application workloads, coding assistant tasks | Formal proof, zero-day discovery, drug-target discovery, market structure analysis | Novel edge cases, disputes, high-stakes judgment | All workloads with heterogeneous units of work |
| Overqualification risk | None | Low | Medium | High on easy work — overthinking, variance, cost | Boredom and departure (humans quit; models do not) | Controlled by routing policy |
| Token economics | No reasoning tokens consumed | Marginal tokens cheap; 7-10x overhead on simple tasks still dominates cost | Better token efficiency than frontier on routine work | 320-fold enterprise reasoning-token growth; marginal token ~zero below the threshold | No tokens; salary is inelastic while delivered intelligence is elastic | Intelligence consumed per successful outcome falls toward zero |
Compiled rules and deterministic logic; the zero-intelligence tier for stable, high-volume, well-specified decisions such as refund eligibility and password resets.
Small and mid-sized open-weight models (e.g. the Qwen3.8-27B class) run local or near the system of record; compressed capability at a small compute footprint.
Models between small open-weight and frontier scale, compressing a given level of capability into a much smaller compute footprint as the closed-to-open-weight shift proceeds.
The most capable, most expensive class of model (Anthropic, OpenAI class), optimized for the extreme intelligence required by research mathematicians, physicists and engineers; a real but small market.
Human judgment and accountability for edge cases, novelty and high-stakes decisions — the elastic-intelligence tier that enterprises already right-size when hiring.
Routing that allocates each unit of work to the least expensive tier still capable of the correct outcome — the default future state described by the article's second prediction.
Routing Policies
Four sample intel:RoutingPolicy instances that
encode the article's allocation logic for the sample workloads.
Default Enterprise Routing Policy
Route every unit of work to the least expensive system still capable of the correct outcome; the article argues enterprises must not leave this decision to the model provider.
Applies to: Claims adjustment, refund handling, invoice reconciliation, password reset, routing decisions Routes to: Deterministic, small open-weight, mid-size, frontier, human
Frontier Reserve Policy
Reserve frontier models for tasks still above the rising capability line: discovery workloads where additional intelligence changes the outcome.
Applies to: Formal proof, drug-target discovery, zero-day discovery, market-structure analysis Routes to: Frontier only
Hybrid Split Policy
Split workloads across tiers: easy, repetitive or private parts stay local or near the system of record; only genuinely hard or novel parts go to a frontier model in the cloud.
Applies to: Nursing triage, logistics coordination, code generation Routes to: Small open-weight, mid-size, frontier
Zero-Intelligence-Outcome Policy
Compile repeated decisions into machine-checkable rules so the millionth claim still has a model working on it but none deciding it; continuously move each workload toward the least expensive system still capable of the correct outcome.
Applies to: Password reset, refund handling, invoice reconciliation Routes to: Deterministic, small open-weight
Executable queries against the sample workload dataset, modeled as
schema:SoftwareSourceCode with target
URIBurner.
QWorkloads grouped by recommended intelligence tier▼
Query entity: Workloads grouped by recommended intelligence tier
PREFIX intel: <https://linkeddata.uriburner.com/DAV/demos/daas/ontology/intelligence-allocation#>
PREFIX schema: <http://schema.org/>
# Which workloads should run on which intelligence tier?
SELECT ?tierName (COUNT(?wl) AS ?workloadCount) (SAMPLE(?wlName) AS ?exampleWorkload)
WHERE {
?wl a intel:WorkloadProfile ;
intel:requiresTier ?tier ;
schema:name ?wlName .
?tier schema:name ?tierName .
}
GROUP BY ?tierName
ORDER BY DESC(?workloadCount)QDiscovery versus application workload counts▼
Query entity: Discovery versus application workload counts
PREFIX intel: <https://linkeddata.uriburner.com/DAV/demos/daas/ontology/intelligence-allocation#>
PREFIX schema: <http://schema.org/>
# Is most enterprise work discovery or application?
SELECT ?valueClassName (COUNT(?wl) AS ?count)
WHERE {
?wl a intel:WorkloadProfile ;
intel:hasValueClass ?vc .
?vc schema:name ?valueClassName .
}
GROUP BY ?valueClassName
ORDER BY DESC(?count)QWorkloads at or above the frontier capability line▼
Query entity: Workloads at or above the frontier capability line
PREFIX intel: <https://linkeddata.uriburner.com/DAV/demos/daas/ontology/intelligence-allocation#>
PREFIX schema: <http://schema.org/>
# Which workloads still need maximum intelligence (extreme threshold)?
SELECT ?wlName ?thresholdName ?tierName
WHERE {
?wl a intel:WorkloadProfile ;
schema:name ?wlName ;
intel:hasIntelligenceThreshold ?th ;
intel:requiresTier ?tier .
?th schema:name ?thresholdName .
?tier schema:name ?tierName .
FILTER(?thresholdName = "Extreme")
}
ORDER BY ?wlNameQAll demo workload profiles (SELECT *)▼
Query entity: All demo workload profiles (SELECT *)
PREFIX intel: <https://linkeddata.uriburner.com/DAV/demos/daas/ontology/intelligence-allocation#>
PREFIX schema: <http://schema.org/>
# Every demo workload instance from the article, with threshold and recommended tier
SELECT ?name ?threshold ?tier
FROM <https://linkeddata.uriburner.com/DAV/demos/daas/right-sizing-intelligence-spend-deepseek_v4flash-1.ttl>
WHERE {
?workload a intel:WorkloadProfile ;
schema:name ?name ;
intel:hasIntelligenceThreshold ?th ;
intel:requiresTier ?tier .
?th schema:name ?threshold .
?tier schema:name ?tier .
}
ORDER BY ?nameExplore the Knowledge Graph
Graph derived from the companion RDF
(right-sizing-intelligence-spend-deepseek_v4flash-1.ttl):
50 nodes, 79 edges. Drag nodes to pin; double-click to unpin; click the canvas to arm zoom.
Graph Settings
Predicate visibility
Node filters
Query the Knowledge Graph
SPARQL Workbench
SELECT results render as HTML tables (text/x-html+tr); DESCRIBE/CONSTRUCT render as nice Turtle (text/x-html-nice-turtle).
Demo instance data — the workloads Jaya's article covers
Twelve intel:WorkloadProfile instances in the companion RDF; every recipe below queries this sample graph.
A six-step operating procedure derived from the article's thesis; each
step links to its absolute-IRI schema:HowToStep entity.
Inventory every workload
List the units of work your organization performs — claims, triage, reconciliation, support, code, research — with volumes, current systems and failure costs.
Classify each workload by value class
Mark each workload as discovery (expanding what is known: proofs, drug targets, zero-days, market structures) or application (making the right decision within known context); coding is the hybrid exception because outcomes are verifiable through tests.
Map thresholds against the capability line
For each workload, estimate its intelligence threshold and where today's open-weight and frontier models sit relative to it; most workloads are not getting harder while models keep getting more intelligent.
Route by tier policy
Assign each workload to the least expensive tier still capable of the correct outcome — deterministic rules, local small models, mid-size models, frontier models, or humans — and encode that allocation in a routing policy inside your harness, at the policy, team and organization layers.
Measure intelligence per successful outcome
Instrument tokens, latency and cost per verified resolution; track whether customer retention and cost per successful outcome improve, not whether token usage rises.
Reinvest savings in context and execution
Compile repeated decisions into machine-checkable rules and invest in proprietary data, dense operational context, processes, systems of record and accumulated judgment — the moat that frontier intelligence alone cannot buy.
Thirteen questions and answers probing the thesis, wrapped in a
schema:FAQPage with schema:mainEntity.
01What is the central argument of 'Right-Sizing Your Intelligence Spend'?▼
Frontier intelligence is extraordinarily valuable but economically narrow; most enterprise work is a context-and-execution problem. Enterprises should route each unit of work to the least expensive system still capable of the correct outcome, reserving frontier models for genuine discovery, and measure progress as intelligence consumed per successful outcome falling toward zero.
02What is an intelligence threshold?▼
Every workload has some level of model intelligence below which the model cannot perform the task. As the model approaches it, greater intelligence creates enormous value; once crossed, the bottleneck changes from intelligence to whether the system knows the company's policies, customers, history, tools and standards and can act reliably, quickly and cheaply.
03Why is frontier research economically narrow?▼
BLS data shows life, physical and social-science occupations account for less than 1% of American employment: roughly 2,000 mathematicians, 20,000 physicists and 37,000 computer and information research scientists, versus millions of nurses, managers, administrators, logistics workers and customer-service representatives.
04What distinguishes discovery value from application value?▼
Some jobs create value by expanding the frontier of what is known — a proof, a drug target, a zero-day, an unfamiliar market structure. Most create value by making the right decision within known context and executing it — claims, triage, logistics, reconciliation. Companies are defended not by superior frontier intelligence but by proprietary data, processes, licenses, networks and accumulated judgment.
05Why is coding the important exception?▼
Coding is a large labor market, the work is digital, and outcomes can often be verified through tests — which makes additional model intelligence unusually valuable. But even coding is developing capability thresholds as smaller and open-weight models approach frontier performance on increasingly broad classes of work.
06What is model overqualification, and why is it persistent?▼
Hiring a dramatically overqualified employee leads to reconsidered decisions, unnecessary complexity, boredom and departure; models are overqualified too, but instead of quitting they keep generating extra complexity, variance and cost indefinitely — and overqualification is now measurable as the overthinking problem.
07What is the overthinking problem?▼
Reasoning models routinely spend additional inference compute on easy questions without improving accuracy; Amazon researchers estimate reasoning systems can generate 7 to 10x as many tokens as necessary on simple tasks, so on sufficiently easy workloads the marginal reasoning token approaches zero.
08Why do labs and enterprises have competing incentives?▼
A lab wins by making more computation useful — its metrics are tokens consumed, session length, agent count and reasoning depth. An enterprise wins by making repeated computation unnecessary — cost per verified resolution. OpenAI reports enterprise reasoning-token consumption rose ~320-fold in a year, while PwC found 56% of 4,454 CEOs saw no significant financial benefit from AI.
09What will the routing layer look like?▼
Millions of routers and gateways inside every serious AI harness — at the policy layer, team level and organization level — deciding per unit of work whether it belongs on a frontier model, a smaller open-weight model, a local model, a deterministic system, a human, or nowhere at all. Its TAM will ultimately look a lot like the workforce itself.
10What are hybrid workloads?▼
Workloads that stop belonging to a single model: easy, repetitive or private parts stay local or near the system of record, and only genuinely hard or novel parts are sent to a frontier model in the cloud.
11What are the two application surfaces for AI?▼
General-purpose proactive assistants such as Instinct make intelligence ambient — observing work, identifying opportunities and initiating tasks without waiting for a prompt. Specialized applications such as Maximor and PlayerZero turn enterprise intent into reliable execution by controlling context, tools, permissions, workflow state and definition of done.
12What role does the improvement layer play?▼
Platforms such as Applied Compute's AC2 train, evaluate, serve and replace the models operating inside existing applications. The harness remains the system of execution while AC2 continuously moves each workload toward the least expensive system still capable of the correct outcome.
13How should progress be measured?▼
Not by how much intelligence a company can afford to consume, but by how quickly intelligence consumed per successful outcome falls toward zero — turning yesterday's expensive reasoning into today's cheaper, more reliable execution.
Sixteen terms wrapped in a schema:DefinedTermSet,
each linked to its schema:DefinedTerm entity and aligned to the
ontology via rdfs:seeAlso.
Intelligence Threshold
The level of model intelligence below which a workload cannot be performed; as the model approaches it, greater intelligence creates enormous value, and once it is crossed the bottleneck shifts from intelligence to context and execution.
Frontier Model
The most capable and most expensive class of AI model, optimized for the extreme intelligence required by research mathematicians, physicists and engineers solving previously unsolved problems.
Open-Weight Model
A model whose weights are publicly released; the shift from closed to open-weight models lowers the cost of a given level of capability.
Context-and-Execution Problem
The dominant mode of economic work: applying what the company already knows across millions of decisions shaped by customers, policies, inventory, contracts and history.
Discovery Value
Value created by expanding the frontier of what is known — a proof, a drug target, a zero-day, or an unfamiliar market structure — where maximum intelligence is exactly what is wanted.
Application Value
Value created by making the right decision within known context and executing it reliably, quickly and cheaply.
Overqualification
Using a system far more capable than a task requires; unlike a bored human who quits, an overqualified model never stops generating extra complexity, variance and cost.
Overthinking
The documented behavior of reasoning models spending additional inference compute on easy questions without improving accuracy.
Reasoning Token
A unit of inference computation consumed during chain-of-thought reasoning; on sufficiently easy workloads the marginal reasoning token approaches zero value.
Routing Layer
Routers and gateways inside every serious AI harness that decide, per unit of work, whether it belongs on a frontier model, a smaller open-weight model, a local model, a deterministic system, a human, or nowhere at all; its TAM resembles the workforce itself.
Hybrid Workload
A workload that stops belonging to a single model: easy, repetitive or private parts stay local or near the system of record, while only genuinely hard or novel parts are sent to a frontier model in the cloud.
Proactive Assistant
A general-purpose assistant that makes intelligence ambient: it observes work, identifies opportunities, and initiates tasks without waiting for a prompt.
Application Harness
A specialized application that turns enterprise intent into reliable execution by controlling the context, tools, permissions, workflow state, and definition of done for a particular job.
Capability Line
The rising line of capability as open-weight and small or mid-sized models cross workload intelligence thresholds; frontier systems should be reserved for tasks that still sit above it.
Intelligence per Successful Outcome
The key metric of right-sizing: intelligence consumed divided by verified successful outcomes; progress is this ratio falling toward zero.
Token Blast
Long-running agent loops, recursive sub-agent spawning and 'keep going until it works' patterns that burn tokens by design — the frictionless default of frontier products.
Reader Perspectives
The article's comment thread, modeled as schema:Comment
entities. All sixteen comments — authors, profile URLs and verbatim text — were extracted from
the live thread via an authenticated browser session (Playwright + Chrome, 2026-08-19);
replies are marked. Kingsley Uyi Idehen
first, then the thread in document order.
Jaya, Yep! For enterprises, the issue is really about calibrating “Good Enough AI,” which isn’t necessarily aligned with what frontier model labs are producing. A business simply wants the job done at a cost optimized for its operating model. No more, no less. 😄
Additional notes generated using Google Gemini 3.7 Flash (Medium) as the LLM component. https://linkeddata.uriburner.com/weblog/?post=right-sizing-intelligence-spend-jaya-gupta-gemini_3_7_flash-1.html
Jaya Gupta Decision Intelligence at scale is becoming most necessary layer for the fast growing organisation that actually drives P&L impact.
Jaya Gupta This is probably your best article yet... The Context graph article was inspirational, this article is pragmatic, impactful and consequential. Important to note that when you apply a human to a job, intelligence delivered is elastic, salary is inelastic. The smartest humans sometimes also have to do dumb jobs. But you don't pay them less for doing so. Even Einstein had to go make a cup of coffee. But in the post-intelligence world, enterprises can vary pay in real-time, so the average cost of intelligence could asymptomatically arrive close to zero. Both the price of intelligence, and the cost of applying it are equally elastic.
Vas Bhandarkar thank you!!!!
This immediately made me think of your context graph thesis. There’s an interesting connection here IMHO: for things like claims, rostering, routing and inventory, the optimization machinery has existed for decades. The hard part has often been formulation, translating company policy, current state, exceptions and judgment into the objective and constraints for this decision. Models make that translation dramatically easier. But they also create an important boundary: a model can propose a constraint, but it shouldn’t be the authority that makes it true. That should trace back to the company’s context graph: policy, system state, precedent, contract, or an authorized person. And sometimes the right answer is simply “infeasible”, not more reasoning.
I can't stop thinking about how Anthropic and the other frontier labs have very different motivations in a future where AI adoption stops being the primary push. Use of these models is still cheaper than it “should” be as the providers boost usage limits and reduce costs to get people hooked. What happens when enough people have adopted AI that these companies can pivot to actually attempting to bring in a profit? Their motivation switches to token usage (which it already kind of is) and they're incentivized to over-engineer answers. Your point about the expensive intelligence being most usefully applied to deciding WHAT needs to be done, not doing it, just sounds like the normal shape of an organization - we don't need the CEO answering every customer support email (nor would that ever work in practice).
Love this! Prices for things (all the things) must go down. the proxy is time. We can deliver more per unit of time now than ever before. Why did we charge more for things in the past? Expertise, time, workmanship, etc. judgement on when/where (fancy folk talk this “taste”) is the iffy part. Why? Because LLMs don’t understand how hoomans will react to things (yet). We have that intuition more dialed because of body language and empathy. This will probably change over time as AI gets eyes and sensors for biologics. Like in Silo when the AI reads heart rate and temp that factor into their reasoning.
Minions prices local inference at zero, so 5.7x is a cloud bill and not a total. Section 3, under Measuring cost, states the assumption outright and ignores hardware and energy. Fair for a laptop already bought. The routing thesis survives this. An enterprise fleet is capacity you pay for idle, and the protocol earns its ratio by moving compute onto exactly that unpriced side. The paper even notes tokens are a poor proxy for on-device cost because the batches never fill. Price the local GPU-hours and re-run it. If 5.7x survives, I am wrong about how much of the gap is accounting.
Jaya, your claims adjuster never needed reasoning. The rules she applies are written down, in policies and contracts the company already owns. The direct route to your zero is to compile them: spend intelligence once, at build time, turning the documents into machine-checkable rules. AI still runs the process after that. It gathers the facts, drafts the output, talks to the customer. The one place it never sits is the decision itself, which is computed from the rules and cites its clause. So the millionth claim still has a model working on it, but none deciding it. Swap the model for a human and nothing changes: she works the claim, the rules decide it. That was always true of the adjuster anyway. Learning from production gets you to determinism where the rules are unwritten. In most businesses, far less is unwritten than people think.
Thank you Jaya Gupta. I have to pin this “while the frontier is an intelligence problem, the rest of the economy is a context-and-execution problem”
Subash Mandanapu thank you
Love this concept.. I would add one more perspective.. “Intelligence has a threshold. Economics has no ceiling!” Once a model is good enough, the winners are the ones who deliver the same outcome with fewer tokens, less latency and lower operational cost. You may resonate with this. I explored the idea in my article, The Economics of Intelligence. https://www.linkedin.com/pulse/economics-intelligence-prasad-prabhakaran-cvv7e
Opposite outcomes chart is brilliant. Incentives will have to align for sustainable AI value, and revenue.
Beautifully captures the current state of thinking in the enterprise.
Pramod Agrawal thank you
About This Collection
This knowledge-graph collection was generated
from the LinkedIn Pulse article
Right-Sizing Your Intelligence Spend
by Jaya Gupta, Avanika Narayan and Jon Saad-Falcon, transformed into RDF via
kg-generator
and rendered via
rdf-infographic-skill.
The collection includes a lightweight
Intelligence Allocation Ontology
(intel:) and a sample workload dataset demonstrating the
article's intelligence-threshold thesis across twelve domain workflows. Entities resolve
through
URIBurner describe.
Technology stack:
- AI Agent: OpenCode-style harness via DeepSeek Harness
- Skills: kg-generator, rdf-infographic-skill
- Language Model: DeepSeek V4 Flash
- Server Platform: Virtuoso
- Knowledge Graph: URIBurner
Right-Sizing Intelligence Is a Data and Context Problem
The article's threshold argument is exactly right, and it points to a deeper structure: the context that becomes the bottleneck once the intelligence threshold is crossed is not amorphous — it is data. Proprietary data, dense operational context, long-built processes and accumulated judgment are precisely what a knowledge graph encodes: entities, relationships, rules and history as machine-readable graph. So the right way to right-size intelligence spend is to build the enterprise context layer as a Semantic Web — an HTTP IRI for every entity, hyperlinks as standardized identifiers, RDF views over systems of record, and SPARQL for inspection — so that the routing layer the article predicts can be grounded in authoritative, queryable context rather than in model priors. The zero the article seeks is reached by compiling repeated decisions into machine-checkable rules and linking those rules to the graph, so the millionth claim still has a model working on it but none deciding it. In short: the frontier model is rented; the context graph is owned. Right-sizing intelligence spend is inseparable from investing in the enterprise knowledge graph — the moat the article describes is, at bottom, a data moat.
Context is data: the moat is a data moat
Proprietary data, processes and accumulated judgment are precisely what a knowledge graph encodes.
Routing is only as good as the context graph it routes against
The routing layer needs authoritative, queryable context, not model priors.
Build the context layer as a Semantic Web
HTTP IRIs, hyperlinks as identifiers, RDF views over systems of record, SPARQL inspection.
The frontier model is rented; the context graph is owned
The durable asset is the enterprise knowledge graph — a data moat.