“Probabilistic reasoning inside. Logical guardrails outside.” — notes on Frank Coyle's AI Engineer World's Fair talk, meshed with the Linked Data corpus.
The article argues that agentic systems — chains of LLM-driven agents — need ontologies as external, non-probabilistic guardrails. Frank Coyle reframes hallucination as a property of the probabilistic substrate rather than a defect to engineer out, shows that sequential reliabilities multiply across agent chains, and positions RDFS/OWL constraint layers as the symbolic-verification half of neurosymbolic AI..
The mesh adds Kingsley Uyi Idehen's Linked Data theses — LLMs as generic RDF clients, and the Semantic Web as waiting for AI — plus the agent-rdf-memory skill as a working instance of an ontology-grounded agent contract.
Hallucination is misnamed: it sounds like a malfunction a better model or stricter prompt will eliminate, but it is a property of the probabilistic substrate.
Sequential reliabilities multiply, they don't average: chaining excellent agents compounds their error rates exponentially.
An ontology works in two dimensions: a shared understanding of the domain, and logical constructs that act as guardrails.
Probabilistic generation with symbolic verification is the combination now called neurosymbolic AI.
Expert systems hit a wall in the eighties and nineties, and CYC is the monument to it — but the ask has changed from unbounded to bounded.
Neural methods returned at scale as LLMs, and symbolic methods are returning as the guardrails those LLMs need.
The essay supplies the thesis; the mesh sources supply the Linked Data context that makes it operational — the RDF-client thesis, the Semantic Web's yin-yang, and the agent-rdf-memory contract as the argument made concrete.
The agent-rdf-memory skill is a working instance of an ontology-grounded agent contract: a shared RDF store across sessions, LLMs, and environments, with core.ttl (identity and output paths), preferences.ttl (the queryable behavioral contract of standing HowToSteps), index.ttl (session index), ontology.ttl (prompt-intent vocabulary), howto documents, and session files named YYYY-MM-DD-{llm}-{env}.ttl. Learning is cross-LLM and cross-environment: the store is shared and filenames are provenance, not isolation. It embodies the essay's thesis — agents grounded by an explicit symbolic contract — in the same machinery the Semantic Web project built: RDF, dereferenceable IRIs, and queryable structure.
Idehen's argument that LLMs are powerful generic RDF clients: they translate structured data into natural language without SPARQL, integrate and mesh data across ERP/CRM/HR systems, blogs, forums, and social media, anchor AI outputs to RDF knowledge graphs to mitigate hallucinations, and support exploration of RDF triples. Real-world examples include progressive knowledge-graph creation from note-taking, augmented utility explainers, visually rich HTML content generation, RDF knowledge-graph visualization, and product or API documentation generation — the start of an Agentic Web era where agents are intermediaries between humans and complex structured data.
The Semantic Web stalled not because the concept was flawed but because usable interfaces were absent; the tipping point came from better UX, with LLMs translating natural language into SPARQL, SQL, or GraphQL and returning charts, tables, or property sheets. Three historical challenges — entity naming, relationship representation formats, and relationship visualization — are addressed by HTTP URLs as identifiers, RDF serializations, and ontologies that act like GPS for semantic navigation. The article closes with business models that now work: rapid knowledge-graph generation, progressive schema enrichment, fine-grained access control, pay-per-query monetization, and open-Web distribution.
Explicit, non-probabilistic constraint layers — RDFS/OWL schema plus domain rules — placed outside the model between the agent and the consequences of its output; they either pass or they don't.
A client capable of making RDF actionable without specialized expertise; LLMs fill the role the Web's Mosaic and Mozilla/Netscape filled for HTML — translating triples into insight and enabling follow-your-nose navigation.
A queryable behavioral contract for AI agents encoded as RDF-Turtle: core.ttl, preferences.ttl, index.ttl, howto documents, and session files — a working instance of an ontology-grounded agent contract.
The arithmetic of chained agents: sequential reliabilities multiply, they don't average, so end-to-end success is the per-step reliability raised to the power of the chain length.
OpenLink AI Layer — bridges between LLMs and structured data sources (files, databases, graphs); the middleware that turns natural language into SPARQL, SQL, or GraphQL queries.
The essay's central division of labour, on five dimensions from the article's own parallel structure.
| Dimension | Probabilistic core (LLM) | Symbolic guardrail (ontology) |
|---|---|---|
| Nature of the layer | Probabilistic core (LLM) | Symbolic guardrail (ontology) |
| How it operates | Samples from a probability distribution over what comes next. | Checks constraints that must hold; passes or fails. |
| Reliability profile | Drifts toward the plausible. | Deterministic within a bounded domain. |
| Division of labour | Fluency; handles open-ended, messy input. | Certainty; sits between agent and consequences. |
| Failure mode | Novel-but-wrong output (hallucination). | Rejects output that violates the model. |
Why the old objection died and what changed: bounded scope, collapsed authoring cost, and a guardrail layer agents need today.
| Dimension | Expert systems (1980s) | Ontology-guarded LLM agents (2026) |
|---|---|---|
| Scope of the model | Unbounded model of general knowledge (CYC's common sense). | Bounded single-domain constraints. |
| Authoring cost | Decades of hand-encoding by knowledge engineers. | LLM-drafted structure, human-corrected. |
| Reasoning architecture | Pure symbolic inference. | Neurosymbolic: probabilistic generation + symbolic verification. |
| Constraint count | Millions of axioms toward a complete world model. | Dozens or hundreds of constraints per domain. |
| Outcome | Failed to scale in the eighties and nineties. | Returning as the guardrail layer agents need today. |
A seven-step walkthrough that turns the essay's thesis into practice.
Before anyone writes a prompt, force the team to agree on what exists and how it connects: the entities (customer, claim, policy), the relationships among them, and the vocabulary of the domain. A surprising amount of an ontology's value shows up during the writing, in the arguments it starts.
Specify what must hold: domain and range restrictions, cardinality, disjointness, and class hierarchies with inferential teeth. The deliverable is a bounded set of rules and OWL constructs — dozens or hundreds of constraints, not a model of the world.
Keep the ontology external to the LLM's generative core: it doesn't weigh evidence and can't be talked into anything. It sits between the agent and the consequences of its output, and it either passes or it doesn't.
Validate each agent step's output against the ontology at the moment it appears — not at the end of the pipeline, where a bad entity that entered at step two has already been faithfully reasoned over for a hundred steps. End-of-pipeline checking is forensics, not validation.
The old objection to ontologies was authoring cost; the new technology dissolves it. Use an LLM to read the corpus, propose structure, and draft the tedious parts, then have the domain expert correct — that is arguably what expertise is.
Expose the ontology's knowledge graph through natural language: let the agent translate questions into SPARQL, SQL, or GraphQL and return charts, tables, or property sheets, anchoring outputs to grounded data and supporting follow-your-nose navigation across triples.
Store the behavioral contract as queryable RDF-Turtle — core identity, standing preferences, session index, and howto documents — and require every agent to read it before tasks and write session memory back after them. The memory store is the ontology made operational.
Thirteen questions this meshup answers, each backed by a schema:Question entity in the companion RDF.
No. Hallucination is a property of the probabilistic substrate, not a defect to engineer out: a language model samples from a probability distribution over what comes next, and 'novel and correct' is reasoning while 'novel and wrong' is hallucination — same machinery, same operation. A model constrained to emit only what it has literally seen is a lookup table; you cannot remove the capacity to be wrong without removing the capacity to imagine a better world. You design around properties; you don't apologize for them.
Because both are the same operation judged by outcome: sampling from a probability distribution in multi-dimensional space over what comes next. The only difference between reasoning and hallucination is whether the output happened to land inside the truth. The capacity to combine things never combined — a sentence no one has written, an argument no one has made — is the entire reason these systems are worth building.
Sequential reliabilities multiply, they don't average: an agent right 99% of the time chained five-deep lands at about 95.1% end to end, and the exponent is the length of the workflow. Nobody's intuition works this way: we estimate an excellent pipeline as excellent, when the correct operation is exponentiation.
0.99 to the hundredth is about 0.366 — a hundred-step workflow of 99%-per-step agents succeeds end to end about 36.6% of the time, failing nearly twice as often as it succeeds. At 95% per step it collapses to about 0.6%, roughly one success in 170 attempts. It doesn't degrade; it collapses.
Halving the per-step error rate from 1% to 0.5% moves a hundred-step workflow from 36.6% to about 60.6% end-to-end success — nearly doubling the outcome from a change most teams would wave off as marginal. That is the exponential working in your favor, and the argument for spending real engineering money on step-level reliability instead of hoping a retry loop covers it.
At the step that produced the error, not at the end. Validation that lives only at the end of the pipeline is forensics: a bad entity enters at step two and every later step faithfully reasons over garbage. The cheapest place to catch an error is the moment it appears.
Two dimensions worth separating: a shared understanding of the domain (entities, relationships, the vocabulary of what exists and how it connects — forcing the team to agree on what a customer or a claim is before anyone writes a prompt), and logical constructs that act as guardrails (RDFS/OWL domain and range, cardinality, disjointness, class hierarchies with inferential teeth) that agent output is grounded against.
The guardrail layer doesn't weigh evidence and can't be talked into anything: it sits outside the model, acting as a guardrail between the agent and the consequences, and it either passes or it doesn't. That determinism within a bounded domain is exactly what makes it a complement to a probabilistic generative core.
The pairing of probabilistic generation with symbolic verification: LLMs supply fluency and the ability to handle open-ended, messy input, while symbolic systems supply certainty within a bounded domain. Neither does the other's job well, which is precisely why the combination works — the architectural answer to the chain-reliability arithmetic.
Expert systems failed because the thing being modeled was unbounded: CYC spent decades hand-encoding common sense in a complete symbolic model of everything a person knows, and it didn't scale. What changed is the scope of the ask — a bounded set of rules and OWL constructs for one domain, constraints in the dozens or hundreds rather than the millions — and authoring cost, since an LLM can read a corpus, propose structure, and draft the tedious parts for a human to correct.
They translate structured data into natural-language insight without SPARQL, integrate and mesh data across disparate sources, anchor AI outputs to RDF knowledge graphs (mitigating hallucinations), and support follow-your-nose navigation across triples. Mosaic and Mozilla/Netscape made the document Web usable; LLMs are the analogous client that makes RDF actionable.
Its core ideas were never flawed — entity naming via HTTP URLs was solved from the start, and relationship representation formats matured — but the absence of usable interfaces stalled it. LLMs and tools like OPAL supply the missing half: natural language translated into SPARQL, SQL, or GraphQL, with answers as charts, tables, or property sheets. The Semantic Web didn't fail; it needed AI as its yin-yang complement.
A queryable behavioral contract for AI agents encoded as RDF-Turtle: a shared store of core.ttl, preferences.ttl, index.ttl, howto documents, and session files that agents must read before tasks and write back after them. It is a working instance of an ontology-grounded agent contract — the essay's 'symbolic guardrails outside the probabilistic core' — built on the Semantic Web's own machinery: RDF, dereferenceable IRIs, and queryable structure.
Terms from the essay and the mesh sources, each linked to its knowledge-graph entity.
A queryable behavioral contract for AI agents encoded as RDF-Turtle: core.ttl, preferences.ttl, index.ttl, howto documents, and session files — a working instance of an ontology-grounded agent contract.
AI systems that act as agents — chaining models, tools, and memory toward goals; the systems whose reliability arithmetic motivates the essay's call for ontology guardrails.
The decades-long project to hand-encode common sense as a complete symbolic model of everything a person knows; the monument to how high the expert-system wall was.
A rule-based symbolic system that encodes a domain expert's knowledge; hit a wall in the eighties and nineties because the thing being modeled (general knowledge) was unbounded.
A client capable of making RDF actionable without specialized expertise; LLMs fill the role the Web's Mosaic and Mozilla/Netscape filled for HTML — translating triples into insight and enabling follow-your-nose navigation.
Explicit, non-probabilistic constraint layers — RDFS/OWL schema plus domain rules — placed outside the model between the agent and the consequences of its output; they either pass or they don't.
Output that is novel but wrong: the same sampling machinery that produces novel-and-correct reasoning, judged only by whether it landed inside the truth. A property of the probabilistic substrate, not a defect to engineer out.
A JSON-based notation for Linked Data used to embed RDF metadata in web pages; one of the formats that made the Semantic Web's adoption real but unlabeled.
A graph of entities (things) and their relationships, typically RDF-based; the substrate LLMs anchor their outputs to, mitigating hallucinations through loose coupling with semantically harmonized data.
Structured data published using dereferenceable HTTP URIs and RDF so that entities connect across the open Web; the mesh sources' shared substrate.
Large Language Model — a neural network trained on large corpora to understand and generate text; the probabilistic generative core of agentic systems and, in Idehen's framing, the powerful generic RDF client the Semantic Web was waiting for.
Linked Open Data — publicly accessible structured datasets using Linked Data principles; with Schema.org, integral to the decentralized, knowledge-centric Web.
Model Context Protocol — a mechanism for brokering structured data access by LLMs; part of the tooling that made the Semantic Web usable, alongside Virtuoso and OPAL.
The pairing of probabilistic generation with symbolic verification: LLMs supply fluency and open-ended input handling; symbolic systems supply certainty within a bounded domain.
A formal model of classes, properties, and relationships shared by a community of interest — the schema layer that makes knowledge graphs machine-computable and reasonable-over; in the essay, the non-probabilistic guardrail layer outside the LLM's generative core.
OpenLink AI Layer — bridges between LLMs and structured data sources (files, databases, graphs); the middleware that turns natural language into SPARQL, SQL, or GraphQL queries.
The W3C Web Ontology Language for formal class/property definitions and reasoning structures atop RDF — domain and range, cardinality, disjointness, and class hierarchies.
The Resource Description Framework — a loosely coupled language for structured data representation as subject-predicate-object triples, with multiple notations and serialization formats.
RDF Schema: the vocabulary description language for RDF, providing classes, properties, domain and range restrictions, and subclass/subproperty hierarchies with inferential teeth.
The W3C vision of a web of linked data — a Giant Global Graph — where information is given well-defined meaning via RDF so machines and humans can cooperate; 'didn't fail, it was waiting for AI'.
The arithmetic of chained agents: sequential reliabilities multiply, they don't average, so end-to-end success is the per-step reliability raised to the power of the chain length.
The declarative query language for RDF data; with LLMs as generic RDF clients, natural language can be translated into SPARQL without specialized expertise.
The mesh as a graph: 92 nodes, 181 links, zero orphans. Drag a node to pin it (double-click to unpin); click a node or an edge label to open its resolver description.
Run recipes against the meshup graph, or edit the query and open it live in URIBurner.
text/x-html+tr · DESCRIBE/CONSTRUCT → text/x-html-nice-turtleThis knowledge graph overview was generated from the companion Turtle (717 triples) using kg-generator and rdf-infographic-skill. Frank Coyle's Medium essay was transformed into RDF along with the three mesh sources, then rendered as this HTML infographic powered by DeepSeek V4 Flash, running on Virtuoso.
AI Agent: OpenCode (DeepSeek Harness)
Skills: kg-generator, rdf-infographic-skill
Language Model: DeepSeek V4 Flash
Server Platform: Virtuoso
Knowledge Graph: URIBurner