Frontier AI intelligence is real but economically narrow; most enterprise work is a context-and-execution problem with its own intelligence threshold, above which additional intelligence adds cost, not value.
Frontier AI intelligence is real but economically narrow; most enterprise work is a context-and-execution problem with its own intelligence threshold, above which additional intelligence adds cost, not value.
Frontier models now solve previously unsolved mathematics and find zero-day vulnerabilities, but that scoreboard reflects a sliver of the economy: life, physical, and social-science research occupations are under 1% of U.S. employment. Most companies — Walmart, UnitedHealth, Amazon, American Airlines, Home Depot, Medtronic — compete on operational context, not raw intelligence. The article's core claim is that every workload has an intelligence threshold: below it a model fails outright; once crossed, the bottleneck shifts to organizational knowledge, reliability, speed, and cost, not additional intelligence. Labs and enterprises face opposing incentives here — a lab wins by making more computation useful, an enterprise wins by making repeated computation unnecessary — which is why architecture is starting to follow the economics: routers and gateways, hybrid workloads that split easy and hard sub-tasks, and dual application surfaces sitting atop a continuous model-improvement layer.
Frontier models now solve previously unsolved mathematics and find zero-day vulnerabilities, but the public scoreboard only reflects cleanly verifiable work.
Research occupations are under 1% of U.S. employment; most companies including Walmart, UnitedHealth, and Amazon compete on operational context, not raw intelligence.
Below a workload's intelligence threshold the model fails; once crossed, the bottleneck shifts to organizational knowledge, reliability, speed, and cost, not raw intelligence.
Labs win by selling more computation; enterprises win by needing less of it. OpenAI reports 320-fold enterprise reasoning-token growth, yet 56% of CEOs report no significant AI financial benefit.
If enterprises must reduce intelligence consumed per outcome, control must move inside the enterprise through routers, hybrid workloads, and specialized application layers.
Sample enterprise and research workloads, each mapped to the intelligence tier the article's argument implies it requires.
Maximum-capability reasoning, optimized for open-ended discovery where the answer is unknown and could be worth millions, e.g. mathematics, drug targets, zero-days.
5 workloadsComputer and information research science work; BLS data places U.S. employment at roughly 37,000.
Finding and validating previously unknown high-severity vulnerabilities in heavily-fuzzed codebases; Claude Opus 4.6 validated more than 500 such vulnerabilities.
Proving previously unsolved conjectures, e.g. an internal OpenAI model disproving a central conjecture in Erdős's planar unit-distance problem.
Physical-science research work; BLS data places U.S. employment at roughly 20,000 physicists.
A large, digital labor market whose outcomes can often be verified through tests, making additional model intelligence unusually valuable even as smaller and open-weight models approach frontier performance.
The minimum capability level, often met by small or open-weight models, sufficient to reliably perform a bounded, well-defined task once the workload's intelligence threshold has been crossed.
4 workloadsApplying standing baggage rules to a passenger's situation; named by the article alongside refund policies and invoice reconciliation as procedures that are not getting any harder.
A simple, bounded task that does not improve when an overqualified agent considers many explanations, launches an unneeded investigation, or composes a personalized essay.
Matching invoices against records via stable, well-defined procedures that are not increasing in difficulty.
Applying refund policies that are stable and are not getting any harder over time.
A capability profile optimized for reliability, speed, and cost within known organizational context rather than raw reasoning depth.
3 workloadsDeciding which of a company's existing rules apply to a specific accident, policyholder, and sequence of events.
Reacting to today's inventory, weather, contracts, and delays rather than discovering a new algorithm.
Combining established clinical protocols with a patient's history, current symptoms, and hospital-specific constraints.
AI models, research systems, and platforms the article cites as evidence for its intelligence right-sizing thesis.
A 27-billion-parameter open-weight model cited as matching or beating Opus 4.6 Max on several coding and agentic benchmarks, illustrating the shift toward smaller compute footprints.
A specialized application cited alongside Maximor as an example of controlling context, tools, permissions, workflow state, and the definition of done for a particular job.
A specialized application cited as an example of turning enterprise intent into reliable execution by controlling context, tools, permissions, and workflow state.
Applied Compute's platform allowing companies to train, evaluate, serve, and replace the models operating inside their existing applications, connecting production experience back to the models.
Anthropic frontier model reported to have found and validated more than 500 high-severity vulnerabilities in heavily-fuzzed codebases.
A general-purpose proactive assistant cited as an example of intelligence made ambient: observing work, identifying opportunities, and initiating tasks without a prompt.
A hybrid-architecture research system in which a frontier cloud model decomposes long-context problems for execution by smaller local models, recovering 97.9% of frontier accuracy at 5.7x lower cost.
The article's named framework: three predictions for how enterprise AI architecture will evolve to route each unit of work to the cheapest system capable of producing the correct outcome.
Every serious AI harness will contain routers and gateways at the policy, team, and organization level, each deciding whether a unit of work belongs on a frontier model, a smaller open-weight model, a local model, a deterministic system, a human, or nowhere at all; organization-level gateways will enforce security, data residency, cost, and budget.
Workloads stop belonging to a single model: easy, repetitive, or private parts of a task stay local or near the system of record, while only genuinely hard or novel parts are sent to a frontier model in the cloud, as demonstrated by the Minions architecture.
AI will diffuse through general-purpose proactive assistants (e.g. Instinct) that make intelligence ambient, and specialized applications (e.g. Maximor, PlayerZero) that turn enterprise intent into reliable execution; beneath both, a platform such as Applied Compute's AC2 continuously specializes and improves the models running inside them.
A sampling of the article's comment thread, led by the knowledge-graph principal's own comment on calibrating 'Good Enough AI' to a business's operating model.
57 reactions · 12 comments on the original LinkedIn article — LinkedIn's guest view surfaces the first four without sign-in; this is that sample.
Interactive graph visualization derived from the companion RDF. Click nodes to resolve, drag to explore. Graph data embedded from companion RDF at generation time.
Query this knowledge graph on URIBurner. The editor opens on the canonical SAMPLE entity-type summary (DAV named graph). Pick a recipe, edit freely, then run live or copy.
The ontology, twelve sample workload instances, and fourteen typed evidence claims make the article’s own material runnable: which intelligence threshold applies to which domain workflow, and what evidence is offered for each argument. Every query targets the DAV-hosted named graph and opens directly on URIBurner.
The thesis in one query: every sample workload joined to the intelligence tier it requires, with its discovery orientation and U.S. employment where the article states them.
PREFIX rs: <https://www.linkedin.com/pulse/right-sizing-your-intelligence-spend-jaya-gupta-7ngre#>
PREFIX schema: <http://schema.org/>
PREFIX rdfs: <http://www.w3.org/2000/01/rdf-schema#>
SELECT ?workflow ?requiredTier ?discoveryOriented ?usEmployment
FROM <https://linkeddata.uriburner.com/DAV/demos/daas/right-sizing-your-intelligence-spend-claude_sonnet_5-1.ttl>
WHERE {
?w a ?class ;
schema:name ?workflow ;
rs:hasIntelligenceThreshold ?tier .
?class rdfs:subClassOf* rs:Workload .
?tier schema:name ?requiredTier .
OPTIONAL { ?w rs:isDiscoveryOriented ?discoveryOriented }
OPTIONAL { ?w rs:hasEmploymentCount ?usEmployment }
}
ORDER BY ?requiredTier ?workflowThe actionable half of the thesis: everything whose threshold sits below the frontier tier is a candidate to route to a smaller, open-weight, local, or deterministic system.
PREFIX rs: <https://www.linkedin.com/pulse/right-sizing-your-intelligence-spend-jaya-gupta-7ngre#>
PREFIX schema: <http://schema.org/>
SELECT ?workflow ?requiredTier ?tokenOverhead ?why
FROM <https://linkeddata.uriburner.com/DAV/demos/daas/right-sizing-your-intelligence-spend-claude_sonnet_5-1.ttl>
WHERE {
?w rs:hasIntelligenceThreshold ?tier ;
schema:name ?workflow ;
schema:description ?why .
?tier schema:name ?requiredTier .
FILTER (?tier != rs:frontierTier)
OPTIONAL { ?w rs:tokenOverheadRatio ?tokenOverhead }
}
ORDER BY ?requiredTier ?workflowRolls the sample workloads up by tier, counting workflows and summing the U.S. employment figures the article cites. Only the research occupations carry stated headcounts, so the zero sums are an artifact of the article citing the rest as 'millions' qualitatively — the narrowness claim, computed rather than asserted.
PREFIX rs: <https://www.linkedin.com/pulse/right-sizing-your-intelligence-spend-jaya-gupta-7ngre#>
PREFIX schema: <http://schema.org/>
SELECT ?requiredTier
(COUNT(DISTINCT ?w) AS ?workflows)
(COUNT(?emp) AS ?workflowsWithStatedFigure)
(SUM(?emp) AS ?statedUsEmployment)
FROM <https://linkeddata.uriburner.com/DAV/demos/daas/right-sizing-your-intelligence-spend-claude_sonnet_5-1.ttl>
WHERE {
?w rs:hasIntelligenceThreshold ?tier .
?tier schema:name ?requiredTier .
OPTIONAL { ?w rs:hasEmploymentCount ?emp }
}
GROUP BY ?requiredTier
ORDER BY DESC(?statedUsEmployment)
# Note: the article gives headcounts only for the research occupations
# (2,000 mathematicians / 20,000 physicists / 37,000 CS researchers) and
# describes the rest qualitatively as "millions". A zero in the sum column
# therefore means "no figure stated", not "no employment" — which is itself
# the article's point about where the measurable scoreboard gets built.Lists the locally minted TBox — the classes and properties introduced to model workflows, tiers, and the threshold relation between them — with their domains and ranges.
PREFIX rs: <https://www.linkedin.com/pulse/right-sizing-your-intelligence-spend-jaya-gupta-7ngre#>
PREFIX rdfs: <http://www.w3.org/2000/01/rdf-schema#>
PREFIX rdf: <http://www.w3.org/1999/02/22-rdf-syntax-ns#>
SELECT ?term ?kind ?label ?domain ?range
FROM <https://linkeddata.uriburner.com/DAV/demos/daas/right-sizing-your-intelligence-spend-claude_sonnet_5-1.ttl>
WHERE {
?term rdfs:isDefinedBy rs: ;
rdf:type ?kind ;
rdfs:label ?label .
OPTIONAL { ?term rdfs:domain ?domain }
OPTIONAL { ?term rdfs:range ?range }
}
ORDER BY ?kind ?termReturns the three named architecture predictions in the order the article makes them, with the technologies each one cites as existing evidence.
PREFIX rs: <https://www.linkedin.com/pulse/right-sizing-your-intelligence-spend-jaya-gupta-7ngre#>
PREFIX schema: <http://schema.org/>
SELECT ?position ?prediction ?evidence
FROM <https://linkeddata.uriburner.com/DAV/demos/daas/right-sizing-your-intelligence-spend-claude_sonnet_5-1.ttl>
WHERE {
?p a rs:ArchitecturePrediction ;
schema:position ?position ;
schema:name ?prediction .
OPTIONAL { ?p schema:mentions ?m . ?m schema:name ?evidence }
}
ORDER BY ?positionEvery figure the article cites, as typed values rather than prose — the measurement, its unit, what it measures, who reported it, and which argument it supports. Ranges (the 7-10x token overhead, the sub-1% employment share, the 500+ vulnerabilities) surface as min/max.
PREFIX rs: <https://www.linkedin.com/pulse/right-sizing-your-intelligence-spend-jaya-gupta-7ngre#>
PREFIX schema: <http://schema.org/>
SELECT ?figure ?value ?min ?max ?unit ?measures ?reportedBy ?supportsArgument
FROM <https://linkeddata.uriburner.com/DAV/demos/daas/right-sizing-your-intelligence-spend-claude_sonnet_5-1.ttl>
WHERE {
?c a rs:EvidenceClaim ;
schema:name ?figure ;
schema:unitText ?unit .
OPTIONAL { ?c schema:value ?value }
OPTIONAL { ?c schema:minValue ?min }
OPTIONAL { ?c schema:maxValue ?max }
OPTIONAL { ?c rs:claimSubject ?s . ?s schema:name ?measures }
OPTIONAL { ?c rs:reportedBy ?o . BIND(REPLACE(STR(?o), "^.*/", "") AS ?reportedBy) }
OPTIONAL { ?c rs:supportsArgument ?a . ?a schema:name ?supportsArgument }
}
ORDER BY ?supportsArgument ?figureThe tier's defining characteristic, made queryable: for every context-and-execution workflow, the specific organizational and situational inputs the article says it must combine. This is the 'the bottleneck is knowing how this company works' claim as data — none of these inputs is a reasoning problem.
PREFIX rs: <https://www.linkedin.com/pulse/right-sizing-your-intelligence-spend-jaya-gupta-7ngre#>
PREFIX schema: <http://schema.org/>
SELECT ?workflow
(GROUP_CONCAT(?input; SEPARATOR=" | ") AS ?contextItMustCombine)
(COUNT(?input) AS ?inputCount)
?definitionOfDone
FROM <https://linkeddata.uriburner.com/DAV/demos/daas/right-sizing-your-intelligence-spend-claude_sonnet_5-1.ttl>
WHERE {
?w rs:hasIntelligenceThreshold rs:contextExecutionTier ;
schema:name ?workflow ;
rs:contextInput ?input .
OPTIONAL { ?w rs:definitionOfDone ?definitionOfDone }
}
GROUP BY ?workflow ?definitionOfDone
ORDER BY DESC(?inputCount) ?workflowFor every bounded workflow above its threshold: the completion condition, whether the task is getting harder, and — where the article gives one — the concrete overqualification failure mode. The password-reset row is the article's own worked example of marginal reasoning value approaching zero.
PREFIX rs: <https://www.linkedin.com/pulse/right-sizing-your-intelligence-spend-jaya-gupta-7ngre#>
PREFIX schema: <http://schema.org/>
SELECT ?workflow ?definitionOfDone ?complexityTrend ?overshootsBy ?tokenOverhead
FROM <https://linkeddata.uriburner.com/DAV/demos/daas/right-sizing-your-intelligence-spend-claude_sonnet_5-1.ttl>
WHERE {
?w rs:hasIntelligenceThreshold rs:thresholdTier ;
schema:name ?workflow .
OPTIONAL { ?w rs:definitionOfDone ?definitionOfDone }
OPTIONAL { ?w rs:complexityTrend ?complexityTrend }
OPTIONAL { ?w rs:overqualificationFailureMode ?overshootsBy }
OPTIONAL { ?w rs:tokenOverheadRatio ?tokenOverhead }
}
ORDER BY ?workflowDetermine whether the workload is discovery-oriented (expanding the frontier of what is known) or context-and-execution-oriented (applying existing organizational knowledge).
Establish the minimum capability level the workload actually requires, distinct from the maximum capability available.
Send work below the threshold to smaller, open-weight, deterministic, or local systems rather than the default frontier model.
Keep maximum-intelligence models for the thin slice of work where the answer is genuinely unknown and additional intelligence still changes the outcome.
Install policy-, team-, and organization-level routers and gateways that enforce security, data-residency requirements, cost budgets, and permitted models per class of work.
Architect hybrid workloads so easy, repetitive, or private parts of a task stay local or near the system of record, sending only genuinely hard or novel parts to a frontier model in the cloud.
Track cost per verified resolution, tokens per processed unit, time to completion, escalation rate, and customer outcome, not raw tokens consumed.
Use an improvement layer, such as AC2, to train, evaluate, serve, and replace the models inside production applications, feeding production experience back into the models.
Frontier-level AI intelligence is real and valuable for open-ended discovery, but most of the economy is a context-and-execution problem: every workload has an intelligence threshold above which additional intelligence adds cost without adding value.
The minimum level of model capability a workload requires. Below it, the model cannot perform the task; as the model approaches it, added intelligence creates enormous value; once crossed, the bottleneck shifts to organizational knowledge, reliability, speed, and cost.
BLS data shows life, physical, and social-science occupations account for less than 1% of American employment (roughly 2,000 mathematicians, 20,000 physicists, 37,000 computer and information research scientists), versus millions of nurses, managers, administrators, logistics workers, and customer-service representatives.
Discovery work expands the frontier of what is known (a proof, a drug target, a zero-day). Context-and-execution work makes the right decision within what an organization already knows, e.g. a claims adjuster applying existing rules or a nurse combining protocols with a patient's history.
Coding is a large, digital labor market where outcomes can often be verified through tests, making additional model intelligence unusually valuable — though smaller and open-weight models, such as Qwen3.8-27B, are increasingly matching frontier performance on coding and agentic benchmarks.
Like an overqualified employee who reconsiders settled decisions and adds unnecessary complexity, an overqualified model generates extra complexity, variance, and cost indefinitely — e.g. a password-reset agent that considers twelve explanations and launches a security investigation instead of resetting the password.
Amazon researchers estimate that reasoning systems can generate 7 to 10 times as many tokens as necessary on simple tasks, meaning the marginal reasoning token on sufficiently easy workloads can approach zero value.
A lab wins by making more computation useful — tokens consumed, session length, agent count, reasoning depth. An enterprise wins by making repeated computation unnecessary — cost per verified resolution, escalation rate, customer outcome. The moment an enterprise eliminates a recurring call, the lab loses a revenue stream.
OpenAI reports that average reasoning-token consumption per enterprise organization rose approximately 320-fold over the past year, yet PwC's survey of 4,454 CEOs found that 56% had not yet seen a significant financial benefit from AI.
Millions of routers and gateways will sit inside every serious AI harness, deciding for each unit of work whether it belongs on a frontier model, a smaller open-weight model, a local model, a deterministic system, a human, or nowhere at all, while organization-level gateways enforce security, data residency, and cost.
A frontier cloud model decomposed complex long-context problems while smaller local models executed the resulting subtasks, recovering 97.9% of frontier accuracy at 5.7x lower cost, or 87.9% at 30.4x lower cost in a more aggressive configuration.
General-purpose proactive assistants (like Instinct) that make intelligence ambient by observing work and initiating tasks, and specialized applications (like Maximor and PlayerZero) that turn enterprise intent into reliable execution by controlling context, tools, permissions, and workflow state.
AC2 lets companies train, evaluate, serve, and replace the models operating inside their existing applications, forming an improvement layer that connects production experience back to the models beneath the application harness.
At just 27 billion parameters, Qwen3.8-27B matches or beats Opus 4.6 Max on several coding and agentic benchmarks, illustrating the shift from frontier-scale to small and mid-sized models that compress capability into a much smaller compute footprint.
The minimum model capability a given workload requires; below it the model fails, above it added intelligence yields diminishing returns.
Work that creates value by expanding the frontier of what is known, such as mathematical proof or vulnerability research.
Work that creates value by making the right decision within an organization's existing data, policies, and operational context.
The failure mode in which a model more capable than a task requires generates extra complexity, variance, and cost instead of a better outcome.
The documented tendency of reasoning models to spend additional inference compute on easy questions without improving accuracy.
A model operating at the maximum-capability edge, optimized for extreme intelligence tasks such as unsolved research problems.
A model whose weights are published, lowering the cost of a given level of capability and enabling smaller, cheaper deployments.
A task split so that easy, repetitive, or private parts stay local or near the system of record while only genuinely hard or novel parts go to a frontier cloud model.
The policy-level infrastructure inside an AI harness that decides which system — frontier model, smaller model, local model, deterministic system, human, or none — handles each unit of work.
A measure of capability delivered per unit of energy; the article cites an 18x improvement in 16 months, enabling comparable capability at dramatically lower compute, energy, and cost.
A hybrid research architecture in which a frontier cloud model decomposes a long-context problem for execution by smaller local models.
A platform enabling enterprises to train, evaluate, serve, and replace the models operating inside their existing application harnesses.
Kingsley Uyi Idehen
Knowledge-graph principalJaya, Yep! For enterprises, the issue is really about calibrating 'Good Enough AI,' which isn't necessarily aligned with what frontier model labs are producing. A business simply wants the job done at a cost optimized for its operating model. No more, no less. 😄
Vas Bhandarkar
Jaya Gupta This is probably your best article yet... The Context graph article was inspirational, this article is pragmatic, impactful and consequential. Important to note that when you apply a human to a job, intelligence delivered is elastic, salary is inelastic. The smartest humans sometimes also have to do dumb jobs. But you don't pay them less for doing so. Even Einstein had to go make a cup of coffee. But in the post-intelligence world, enterprises can vary pay in real-time, so the average cost of intelligence could asymptomatically arrive close to zero. Both the price of intelligence, and the cost of applying it are equally elastic.
Chintan Thakkar
I think we're underestimating how much AI changes organizational design. If one person can now do the work of a small team, the org chart itself starts looking different.
Prasad Prabhakaran
Love this concept.. I would add one more perspective.. 'Intelligence has a threshold. Economics has no ceiling!' Once a model is good enough, the winners are the ones who deliver the same outcome with fewer tokens, less latency and lower operational cost. You may resonate with this. I explored the idea in my article, The Economics of Intelligence.