Medium · Frank Coyle, PhD · Aug 4, 2026 · 8 min read · notes & commentary

Why Agentic Systems Need
Ontologies

“Probabilistic reasoning inside. Logical guardrails outside.” — notes on Frank Coyle's AI Engineer World's Fair talk, meshed with the Linked Data corpus.

The article argues that agentic systems — chains of LLM-driven agents — need ontologies as external, non-probabilistic guardrails. Frank Coyle reframes hallucination as a property of the probabilistic substrate rather than a defect to engineer out, shows that sequential reliabilities multiply across agent chains, and positions RDFS/OWL constraint layers as the symbolic-verification half of neurosymbolic AI..

The mesh adds Kingsley Uyi Idehen's Linked Data theses — LLMs as generic RDF clients, and the Semantic Web as waiting for AI — plus the agent-rdf-memory skill as a working instance of an ontology-grounded agent contract.

SAMPLE · DRIFT ↔ GROUND · VERIFY · PASS / FAILProbabilistic coreLLM · fluency · open-endedSymbolic guardrailontology · constraints · pass/failground · verifyreject · enforce
By Frank Coyle, PhD Medium Source essay ↗
KG curated by kg-generator on behalf of Kingsley Uyi Idehen
01 · article section

Hallucination is a property, not a defect

Hallucination is misnamed: it sounds like a malfunction a better model or stricter prompt will eliminate, but it is a property of the probabilistic substrate.

02 · article section

The arithmetic nobody runs

Sequential reliabilities multiply, they don't average: chaining excellent agents compounds their error rates exponentially.

03 · article section

Ontologies to the rescue

An ontology works in two dimensions: a shared understanding of the domain, and logical constructs that act as guardrails.

04 · article section

The neurosymbolic pairing

Probabilistic generation with symbolic verification is the combination now called neurosymbolic AI.

05 · article section

Wait — experts? Didn't we try this already?

Expert systems hit a wall in the eighties and nineties, and CYC is the monument to it — but the ask has changed from unbounded to bounded.

06 · article section

The pendulum

Neural methods returned at scale as LLMs, and symbolic methods are returning as the guardrails those LLMs need.

The Mesh

Four sources, one argument

The essay supplies the thesis; the mesh sources supply the Linked Data context that makes it operational — the RDF-client thesis, the Semantic Web's yin-yang, and the agent-rdf-memory contract as the argument made concrete.

The working instance

Agent RDF Memory

The agent-rdf-memory skill is a working instance of an ontology-grounded agent contract: a shared RDF store across sessions, LLMs, and environments, with core.ttl (identity and output paths), preferences.ttl (the queryable behavioral contract of standing HowToSteps), index.ttl (session index), ontology.ttl (prompt-intent vocabulary), howto documents, and session files named YYYY-MM-DD-{llm}-{env}.ttl. Learning is cross-LLM and cross-environment: the store is shared and filenames are provenance, not isolation. It embodies the essay's thesis — agents grounded by an explicit symbolic contract — in the same machinery the Semantic Web project built: RDF, dereferenceable IRIs, and queryable structure.

The RDF-client thesis

Large Language Models (LLMs) as Powerful Generic RDF Clients

Idehen's argument that LLMs are powerful generic RDF clients: they translate structured data into natural language without SPARQL, integrate and mesh data across ERP/CRM/HR systems, blogs, forums, and social media, anchor AI outputs to RDF knowledge graphs to mitigate hallucinations, and support exploration of RDF triples. Real-world examples include progressive knowledge-graph creation from note-taking, augmented utility explainers, visually rich HTML content generation, RDF knowledge-graph visualization, and product or API documentation generation — the start of an Agentic Web era where agents are intermediaries between humans and complex structured data.

The Semantic Web yin-yang

The Semantic Web Project Didn't Fail — It Was Waiting for AI (The Yin of its Yang)

The Semantic Web stalled not because the concept was flawed but because usable interfaces were absent; the tipping point came from better UX, with LLMs translating natural language into SPARQL, SQL, or GraphQL and returning charts, tables, or property sheets. Three historical challenges — entity naming, relationship representation formats, and relationship visualization — are addressed by HTTP URLs as identifiers, RDF serializations, and ontologies that act like GPS for semantic navigation. The article closes with business models that now work: rapid knowledge-graph generation, progressive schema enrichment, fine-grained access control, pay-per-query monetization, and open-Web distribution.

Guardrails (ontology-backed)

Explicit, non-probabilistic constraint layers — RDFS/OWL schema plus domain rules — placed outside the model between the agent and the consequences of its output; they either pass or they don't.

Generic RDF client

A client capable of making RDF actionable without specialized expertise; LLMs fill the role the Web's Mosaic and Mozilla/Netscape filled for HTML — translating triples into insight and enabling follow-your-nose navigation.

Agent RDF memory

A queryable behavioral contract for AI agents encoded as RDF-Turtle: core.ttl, preferences.ttl, index.ttl, howto documents, and session files — a working instance of an ontology-grounded agent contract.

Sequential reliability

The arithmetic of chained agents: sequential reliabilities multiply, they don't average, so end-to-end success is the per-step reliability raised to the power of the chain length.

OPAL

OpenLink AI Layer — bridges between LLMs and structured data sources (files, databases, graphs); the middleware that turns natural language into SPARQL, SQL, or GraphQL queries.

Head-to-Head

The neurosymbolic pairing: probabilistic core vs symbolic guardrail

The essay's central division of labour, on five dimensions from the article's own parallel structure.

DimensionProbabilistic core (LLM)Symbolic guardrail (ontology)
Nature of the layerProbabilistic core (LLM)Symbolic guardrail (ontology)
How it operatesSamples from a probability distribution over what comes next.Checks constraints that must hold; passes or fails.
Reliability profileDrifts toward the plausible.Deterministic within a bounded domain.
Division of labourFluency; handles open-ended, messy input.Certainty; sits between agent and consequences.
Failure modeNovel-but-wrong output (hallucination).Rejects output that violates the model.

The pendulum: expert systems (1980s) vs ontology-guarded agents (2026)

Why the old objection died and what changed: bounded scope, collapsed authoring cost, and a guardrail layer agents need today.

DimensionExpert systems (1980s)Ontology-guarded LLM agents (2026)
Scope of the modelUnbounded model of general knowledge (CYC's common sense).Bounded single-domain constraints.
Authoring costDecades of hand-encoding by knowledge engineers.LLM-drafted structure, human-corrected.
Reasoning architecturePure symbolic inference.Neurosymbolic: probabilistic generation + symbolic verification.
Constraint countMillions of axioms toward a complete world model.Dozens or hundreds of constraints per domain.
OutcomeFailed to scale in the eighties and nineties.Returning as the guardrail layer agents need today.
How-To

Ground an agentic system with an ontology-backed guardrail layer

A seven-step walkthrough that turns the essay's thesis into practice.

VOCABULARY & AUTHORINGentities · relationships · what exists and how it connectsCONSTRAINTS · RDFS / OWL GUARDRAILSdomain · range · cardinality · disjointness · pass / failGROUNDING, AUTHORING & CLIENTvalidate at origin · LLM-drafted constraints · natural-language SPARQL1Agree shared domain vocabulary2Encode constraints in RDFS/OWL3Place guardrail outside the model4Ground output at producing step5LLM drafts; human corrects6Wire LLM as generic RDF client7Persist as agent RDF memoryread before tasks · write session memory back after · the contract is the ontology made operational
  • 1

    Agree a shared domain vocabulary

    Before anyone writes a prompt, force the team to agree on what exists and how it connects: the entities (customer, claim, policy), the relationships among them, and the vocabulary of the domain. A surprising amount of an ontology's value shows up during the writing, in the arguments it starts.

  • 2

    Encode the constraints in RDFS and OWL

    Specify what must hold: domain and range restrictions, cardinality, disjointness, and class hierarchies with inferential teeth. The deliverable is a bounded set of rules and OWL constructs — dozens or hundreds of constraints, not a model of the world.

  • 3

    Place the guardrail layer outside the model

    Keep the ontology external to the LLM's generative core: it doesn't weigh evidence and can't be talked into anything. It sits between the agent and the consequences of its output, and it either passes or it doesn't.

  • 4

    Ground every output at the step that produced it

    Validate each agent step's output against the ontology at the moment it appears — not at the end of the pipeline, where a bad entity that entered at step two has already been faithfully reasoned over for a hundred steps. End-of-pipeline checking is forensics, not validation.

  • 5

    Let the LLM draft the ontology; let humans correct it

    The old objection to ontologies was authoring cost; the new technology dissolves it. Use an LLM to read the corpus, propose structure, and draft the tedious parts, then have the domain expert correct — that is arguably what expertise is.

  • 6

    Wire the LLM as a generic RDF client

    Expose the ontology's knowledge graph through natural language: let the agent translate questions into SPARQL, SQL, or GraphQL and return charts, tables, or property sheets, anchoring outputs to grounded data and supporting follow-your-nose navigation across triples.

  • 7

    Persist the contract as agent RDF memory

    Store the behavioral contract as queryable RDF-Turtle — core identity, standing preferences, session index, and howto documents — and require every agent to read it before tasks and write session memory back after them. The memory store is the ontology made operational.

  • FAQ

    Frequently Asked Questions

    Thirteen questions this meshup answers, each backed by a schema:Question entity in the companion RDF.

    Q1

    No. Hallucination is a property of the probabilistic substrate, not a defect to engineer out: a language model samples from a probability distribution over what comes next, and 'novel and correct' is reasoning while 'novel and wrong' is hallucination — same machinery, same operation. A model constrained to emit only what it has literally seen is a lookup table; you cannot remove the capacity to be wrong without removing the capacity to imagine a better world. You design around properties; you don't apologize for them.

    Q2

    Because both are the same operation judged by outcome: sampling from a probability distribution in multi-dimensional space over what comes next. The only difference between reasoning and hallucination is whether the output happened to land inside the truth. The capacity to combine things never combined — a sentence no one has written, an argument no one has made — is the entire reason these systems are worth building.

    Q3

    Sequential reliabilities multiply, they don't average: an agent right 99% of the time chained five-deep lands at about 95.1% end to end, and the exponent is the length of the workflow. Nobody's intuition works this way: we estimate an excellent pipeline as excellent, when the correct operation is exponentiation.

    Q4

    0.99 to the hundredth is about 0.366 — a hundred-step workflow of 99%-per-step agents succeeds end to end about 36.6% of the time, failing nearly twice as often as it succeeds. At 95% per step it collapses to about 0.6%, roughly one success in 170 attempts. It doesn't degrade; it collapses.

    Q5

    Halving the per-step error rate from 1% to 0.5% moves a hundred-step workflow from 36.6% to about 60.6% end-to-end success — nearly doubling the outcome from a change most teams would wave off as marginal. That is the exponential working in your favor, and the argument for spending real engineering money on step-level reliability instead of hoping a retry loop covers it.

    Q6

    At the step that produced the error, not at the end. Validation that lives only at the end of the pipeline is forensics: a bad entity enters at step two and every later step faithfully reasons over garbage. The cheapest place to catch an error is the moment it appears.

    Q7

    Two dimensions worth separating: a shared understanding of the domain (entities, relationships, the vocabulary of what exists and how it connects — forcing the team to agree on what a customer or a claim is before anyone writes a prompt), and logical constructs that act as guardrails (RDFS/OWL domain and range, cardinality, disjointness, class hierarchies with inferential teeth) that agent output is grounded against.

    Q8

    The guardrail layer doesn't weigh evidence and can't be talked into anything: it sits outside the model, acting as a guardrail between the agent and the consequences, and it either passes or it doesn't. That determinism within a bounded domain is exactly what makes it a complement to a probabilistic generative core.

    Q9

    The pairing of probabilistic generation with symbolic verification: LLMs supply fluency and the ability to handle open-ended, messy input, while symbolic systems supply certainty within a bounded domain. Neither does the other's job well, which is precisely why the combination works — the architectural answer to the chain-reliability arithmetic.

    Q10

    Expert systems failed because the thing being modeled was unbounded: CYC spent decades hand-encoding common sense in a complete symbolic model of everything a person knows, and it didn't scale. What changed is the scope of the ask — a bounded set of rules and OWL constructs for one domain, constraints in the dozens or hundreds rather than the millions — and authoring cost, since an LLM can read a corpus, propose structure, and draft the tedious parts for a human to correct.

    Q11

    They translate structured data into natural-language insight without SPARQL, integrate and mesh data across disparate sources, anchor AI outputs to RDF knowledge graphs (mitigating hallucinations), and support follow-your-nose navigation across triples. Mosaic and Mozilla/Netscape made the document Web usable; LLMs are the analogous client that makes RDF actionable.

    Q12

    Its core ideas were never flawed — entity naming via HTTP URLs was solved from the start, and relationship representation formats matured — but the absence of usable interfaces stalled it. LLMs and tools like OPAL supply the missing half: natural language translated into SPARQL, SQL, or GraphQL, with answers as charts, tables, or property sheets. The Semantic Web didn't fail; it needed AI as its yin-yang complement.

    Q13

    A queryable behavioral contract for AI agents encoded as RDF-Turtle: a shared store of core.ttl, preferences.ttl, index.ttl, howto documents, and session files that agents must read before tasks and write back after them. It is a working instance of an ontology-grounded agent contract — the essay's 'symbolic guardrails outside the probabilistic core' — built on the Semantic Web's own machinery: RDF, dereferenceable IRIs, and queryable structure.

    Glossary

    Defined Terms

    Terms from the essay and the mesh sources, each linked to its knowledge-graph entity.

    A queryable behavioral contract for AI agents encoded as RDF-Turtle: core.ttl, preferences.ttl, index.ttl, howto documents, and session files — a working instance of an ontology-grounded agent contract.

    AI systems that act as agents — chaining models, tools, and memory toward goals; the systems whose reliability arithmetic motivates the essay's call for ontology guardrails.

    The decades-long project to hand-encode common sense as a complete symbolic model of everything a person knows; the monument to how high the expert-system wall was.

    A rule-based symbolic system that encodes a domain expert's knowledge; hit a wall in the eighties and nineties because the thing being modeled (general knowledge) was unbounded.

    A client capable of making RDF actionable without specialized expertise; LLMs fill the role the Web's Mosaic and Mozilla/Netscape filled for HTML — translating triples into insight and enabling follow-your-nose navigation.

    Explicit, non-probabilistic constraint layers — RDFS/OWL schema plus domain rules — placed outside the model between the agent and the consequences of its output; they either pass or they don't.

    Output that is novel but wrong: the same sampling machinery that produces novel-and-correct reasoning, judged only by whether it landed inside the truth. A property of the probabilistic substrate, not a defect to engineer out.

    A JSON-based notation for Linked Data used to embed RDF metadata in web pages; one of the formats that made the Semantic Web's adoption real but unlabeled.

    A graph of entities (things) and their relationships, typically RDF-based; the substrate LLMs anchor their outputs to, mitigating hallucinations through loose coupling with semantically harmonized data.

    Structured data published using dereferenceable HTTP URIs and RDF so that entities connect across the open Web; the mesh sources' shared substrate.

    Large Language Model — a neural network trained on large corpora to understand and generate text; the probabilistic generative core of agentic systems and, in Idehen's framing, the powerful generic RDF client the Semantic Web was waiting for.

    Linked Open Data — publicly accessible structured datasets using Linked Data principles; with Schema.org, integral to the decentralized, knowledge-centric Web.

    Model Context Protocol — a mechanism for brokering structured data access by LLMs; part of the tooling that made the Semantic Web usable, alongside Virtuoso and OPAL.

    The pairing of probabilistic generation with symbolic verification: LLMs supply fluency and open-ended input handling; symbolic systems supply certainty within a bounded domain.

    A formal model of classes, properties, and relationships shared by a community of interest — the schema layer that makes knowledge graphs machine-computable and reasonable-over; in the essay, the non-probabilistic guardrail layer outside the LLM's generative core.

    OpenLink AI Layer — bridges between LLMs and structured data sources (files, databases, graphs); the middleware that turns natural language into SPARQL, SQL, or GraphQL queries.

    The W3C Web Ontology Language for formal class/property definitions and reasoning structures atop RDF — domain and range, cardinality, disjointness, and class hierarchies.

    The Resource Description Framework — a loosely coupled language for structured data representation as subject-predicate-object triples, with multiple notations and serialization formats.

    RDF Schema: the vocabulary description language for RDF, providing classes, properties, domain and range restrictions, and subclass/subproperty hierarchies with inferential teeth.

    The W3C vision of a web of linked data — a Giant Global Graph — where information is given well-defined meaning via RDF so machines and humans can cooperate; 'didn't fail, it was waiting for AI'.

    The arithmetic of chained agents: sequential reliabilities multiply, they don't average, so end-to-end success is the per-step reliability raised to the power of the chain length.

    The declarative query language for RDF data; with LLMs as generic RDF clients, natural language can be translated into SPARQL without specialized expertise.

    Knowledge Graph

    KG Explorer

    The mesh as a graph: 92 nodes, 181 links, zero orphans. Drag a node to pin it (double-click to unpin); click a node or an edge label to open its resolver description.

    92 nodes / 181 links
    Physics
    Predicates
    Nodes & Literals
    Resolver & Arrows
    SPARQL

    Query the Knowledge Graph

    Run recipes against the meshup graph, or edit the query and open it live in URIBurner.

    SPARQL workbench
    Result formats: SELECT → text/x-html+tr · DESCRIBE/CONSTRUCT → text/x-html-nice-turtle
    About

    About This Page

    This knowledge graph overview was generated from the companion Turtle (717 triples) using kg-generator and rdf-infographic-skill. Frank Coyle's Medium essay was transformed into RDF along with the three mesh sources, then rendered as this HTML infographic powered by DeepSeek V4 Flash, running on Virtuoso.

    AI Agent: OpenCode (DeepSeek Harness)

    Skills: kg-generator, rdf-infographic-skill

    Language Model: DeepSeek V4 Flash

    Server Platform: Virtuoso

    Knowledge Graph: URIBurner

    People

    Artur d'Avila GarcezDouglas LenatFrank Coyle, PhDGary MarcusHenry KautzKingsley Uyi IdehenLéon BottouMarvin MinskyR. V. GuhaSeymour PapertYuxi Liu

    Organizations

    CycorpGitHubLinkedInMediumOpenLink SoftwareSouthern Methodist UniversityUniversity of BolognaUniversity of California, BerkeleyWorld Wide Web Consortium (W3C)