Navigation

An RDF reading of the OpenLink forum essay

LLMs & Language

Language is the map. The territory is everything it leaves out.

A large language models (LLMs) doesn't know language. It emulates it — a Language Emulation Machine, a Langulator. Signs become tokens, syntax becomes learned regularity, semantics becomes contextual association, pragmatics becomes runtime context. And every bit of it is implicit: numerically encoded, probabilistic, unavailable to logic. A Semantic Web works the other way — explicit, identified, declared, machine-computable. The essay's verdict: LLMs don't make explicit semantics obsolete. They make explicit semantics easier to use.

4
Language components
5
Division layers
9
Pipeline steps
12
FAQ questions
Langulator — the Language Emulation Machine THE LANGULATOR LLM · implicit semantics tokens probabilities A Semantic Web — explicit, machine-computable semantics A SEMANTIC WEB explicit semantics URIs logic MAP ⇆ TERRITORY
Written byKingsley IdehenPublisherOpenLink Software
CurationKG curated by kg-generator, rdf-infographic-skill, and Muse Spark on behalf of Kingsley Idehen.
The argument

Overview

Language is the map. The territory is everything it leaves out.

01 · The thesis

A machine that emulates language

Language is the systematic use of signs, syntax, and semantics to encode and decode information — with pragmatics supplying the runtime context that fixes meaning. That makes a large language model something more specific than its name suggests: not a machine that knows language, but a machine that emulates it. A Language Emulation Machine. A Langulator.

02 · The map

Four components, four counterparts

Each component of language has a counterpart inside the model. Signs — the carriers of meaning — become tokens and their numerical representations. Syntax — the structural rules — becomes learned regularities absorbed from training text. Semantics — what expressions mean — becomes contextual associations, also learned, also numerical. Pragmatics — meaning in situation — becomes the runtime context: the prompt, the conversation, the task at hand. Nothing in the model is declared; everything is encoded.

01 · Component

Signs

Provides: denotation — the carriers of meaning

In the model: Tokens and numerical representations

Signs are the perceptible forms that carry meaning: words, marks, sounds. Inside a large language model they become tokens — sub-word units mapped to numerical vectors — the raw material everything else is computed over.

02 · Component

Syntax

Provides: structure — the combinatorial rules

In the model: Learned structural regularities

Syntax is the system of rules for combining signs into well-formed expressions. The model never receives a grammar; it absorbs structural regularities statistically from the shape of training text, and reproduces them fluently.

03 · Component

Semantics

Provides: meaning — what expressions denote

In the model: Learned contextual associations

Semantics is what signs and structures mean — reference, sense, truth conditions. In the model, meaning is not declared anywhere; it is encoded implicitly as contextual associations between numerical representations, recoverable only by running the model.

04 · Component

Pragmatics

Provides: situation — meaning in context of use

In the model: Runtime context

Pragmatics is how situation fixes meaning: who is speaking, to whom, about what, right now — the difference between a financial bank and a river bank. For the model, that situation is the runtime context: the prompt, the conversation history, the task.

03 · The territory

What the map leaves out

Hence the central image: language is the map, and what it describes is the territory. A map is useful precisely because it leaves things out — and dangerous when it is mistaken for the ground itself. An LLM's semantics are implicit: numerically encoded, probabilistic, and unavailable to deterministic inference. A Semantic Web's semantics are the opposite: explicit relationships, identified by URIs, declared in ontologies, and machine-computable by logic. The conclusion follows from the contrast: LLMs don't make explicit semantics obsolete — they make explicit semantics easier to use. Language for the fuzzy parts. Semantic Web for the precise parts. Loose coupling between them.

“Language is the map.”
The central image — the map/territory distinction
“Language for the fuzzy parts. Semantic Web for the precise parts. Loose coupling between them.”
The division-of-labor rule
Infographic: LLMs, Semantic Web, ontology, and Linked Data symbiosis
LLMs, Semantic Web, ontology, and Linked Data symbiosis: the accompanying infographic — implicit semantics meeting explicit semantics (click to enlarge).
04 · The labor

The division of labor

Five layers keep the fuzzy and the precise apart — each with a job the other cannot do.

Layer 01

LLM — the Langulator

Probabilistic natural-language handling

The model emulates language: it parses intent from prose, holds the fuzzy, contextual, ambiguous parts of a task, and renders results back into natural language. Its semantics stay implicit — that is what makes it fluent, and what makes it unreliable for precision.

Layer 02

RDF + Linked Data + Ontologies

Explicit, machine-computable semantics

Relationships identified by URIs, declared in shared vocabularies, traversable over HTTP, and available to deterministic inference. Where the model's meaning is encoded and approximate, here it is declared and exact.

Layer 03

Logic

Deterministic inference

Rules applied to declared relationships yield conclusions that follow necessarily — no sampling, no temperature. The precise parts of a task land here.

Layer 04

Agent harness

Controlled action and tool selection

The scaffolding that lets an agent act: choosing skills and tools, sequencing calls, enforcing policy. It couples the fluent layer to the precise layer without confusing the two.

Layer 05

Data spaces

Databases, knowledge bases, filesystems, APIs

Where the ground truth lives. Documents, tables, graphs, and services — the territory the map describes, reachable through governed interfaces.

05 · The sources

Cited works

Head to head

Two Kinds of Semantics

The model's way against a Semantic Web way — six aspects, one verdict: keep the fuzzy and the precise loosely coupled.

Aspect The model's way — Langulatorimplicit semantics A Semantic Web wayexplicit semantics
SemanticsImplicit — numerically encoded in model weights, learned from text.Explicit — identified by URIs, declared as RDF, shared via ontologies.
RepresentationTokens and embeddings — meaning by statistical association.URIs and triples — meaning by declaration.
IdentityNone — entities blur into proximity in vector space.Stable identity through IRIs: the same thing is the same URI everywhere.
InferenceProbabilistic generation — fluent, approximate, uncheckable.Deterministic logic over declared relationships — computed, not guessed.
CouplingFused inside the model — the fuzzy and the precise collapse together.Loose — language for the fuzzy parts, a Semantic Web for the precise parts.
Failure modeFluent output mistaken for ground truth.Statements machines can check, ground, and compose.

The model’s way — Langulator

implicit semantics
SemanticsImplicit — numerically encoded in model weights, learned from text.
RepresentationTokens and embeddings — meaning by statistical association.
IdentityNone — entities blur into proximity in vector space.
InferenceProbabilistic generation — fluent, approximate, uncheckable.
CouplingFused inside the model — the fuzzy and the precise collapse together.
Failure modeFluent output mistaken for ground truth.

A Semantic Web way

explicit semantics
SemanticsExplicit — identified by URIs, declared as RDF, shared via ontologies.
RepresentationURIs and triples — meaning by declaration.
IdentityStable identity through IRIs: the same thing is the same URI everywhere.
InferenceDeterministic logic over declared relationships — computed, not guessed.
CouplingLoose — language for the fuzzy parts, a Semantic Web for the precise parts.
Failure modeStatements machines can check, ground, and compose.
Pipeline

HowTo: The Division-of-Labor Pipeline

Nine steps of the division-of-labor pipeline — human to human, with the fuzzy and the precise parts kept apart.

01

Start from the human

Every task begins with a person and an intent. The pipeline's job is to carry that intent through machinery without losing it — which is why the first and last steps are human.

02

Express the intent in natural language

The human states what they want in ordinary prose. Natural language is the highest-bandwidth interface a person has; the pipeline meets them there instead of demanding a query language.

03

Let the Langulator parse the fuzzy parts

The LLM — the Language Emulation Machine — handles everything probabilistic: disambiguating intent, filling in context, tolerating ambiguity. Its implicit semantics are a feature here, not a bug.

04

Ground entities in RDF, Linked Data, and ontologies

Translate the parsed intent into identified things and declared relationships: URIs for entities, RDF for the relations, ontologies for shared meaning. This is where implicit becomes explicit.

05

Operate over data spaces

Run the precise parts against databases, knowledge bases, filesystems, and APIs — the territory, not the map. Governed access decides what the agent may touch.

06

Apply logic for deterministic results

Where the answer must follow necessarily, use inference over declared relationships — not sampling. Precision is computed, not guessed.

07

Return through the Langulator

Hand the precise results back to the model for rendering: summaries, explanations, prose shaped for the human's situation. Fluency is the model's home turf.

08

Deliver in natural language

The human receives the outcome as language — the same medium the intent arrived in. The machinery in the middle stays invisible.

09

Keep the coupling loose

The standing rule: language for the fuzzy parts, a Semantic Web for the precise parts, and only loose coupling between them. Never let the probabilistic layer pretend to be the precise one.

Debrief

FAQ

Twelve named questions, twelve named answers — each question heading is its entity.

The author's coinage for a large language model, short for Language Emulation Machine. The name is the argument: the model does not possess language the way a speaker does — it emulates the systematic use of signs, syntax, semantics, and pragmatics, reproducing their outward behavior from learned numerical patterns.

Signs (the carriers of meaning), syntax (the combinatorial structure), semantics (what expressions mean), and pragmatics (how situation fixes meaning). Each has a counterpart inside the model: tokens and numerical representations, learned structural regularities, learned contextual associations, and runtime context.

That meaning is fixed by situation, not by the sign alone. The word “bank” denotes a financial institution in one context and the edge of a river in another. Pragmatics — the runtime context — selects the reading. The model does the same work with its context window that a speaker does with a situation.

Signs become tokens and numerical representations; syntax becomes learned structural regularities absorbed from training text; semantics becomes learned contextual associations between those representations; pragmatics becomes the runtime context — the prompt, the conversation, the task. Nothing is declared; everything is encoded.

Nothing is false about it, but it undersells what the prediction machinery achieves. Next-token prediction is the training objective; language emulation — reproducing the systematic behavior of signs, syntax, semantics, and pragmatics — is what that objective produces at scale. “Langulator” names the product, not the mechanism.

Meaning that is encoded but not declared. In an LLM, what a token “means” exists only as patterns across numerical representations — recoverable by running the model, unavailable to inspection, and closed to deterministic inference. Fluent, approximate, and fundamentally probabilistic.

Meaning that is identified and declared: entities named by URIs, relationships stated as RDF triples, vocabularies shared through ontologies. Explicit semantics are machine-computable — a logic engine can apply rules to them and derive conclusions that follow necessarily.

The difference between the two kinds of semantics. An LLM associates “Paris” with “France” through learned contextual association — strong, fluent, and implicit. Explicit semantics instead declares the relationship (Paris is the capital of France) as an identified, declared, machine-computable fact that logic can build on.

The central image: language describes the world the way a map describes terrain. A map is useful precisely because it leaves things out — and dangerous when mistaken for the ground itself. Confusing an LLM's fluent map for the territory is the characteristic failure mode.

Five layers, loosely coupled: the LLM handles probabilistic natural-language work; RDF, Linked Data, and ontologies supply explicit machine-computable semantics; logic supplies deterministic inference; the agent harness supplies controlled action and skill/tool selection; data spaces — databases, knowledge bases, filesystems, APIs — supply ground truth.

Agents act, and acting on implicit semantics alone means acting on approximations. Grounding the agent's entities and relationships in explicit, identified, declared semantics gives it something deterministic to stand on: the fluent layer interprets intent, the precise layer constrains what may be concluded and done.

Not smarter models — the point is that raw capability was never the gap. The missing piece was the explicit semantic layer that lets machine intelligence be checked, grounded, and composed: identified entities, declared relationships, and logic. LLMs don't make explicit semantics obsolete; they make explicit semantics easier to use.

Lexicon

Glossary

Twelve defined terms, each linked to its canonical home.

Sign

The perceptible carrier of meaning — a word, mark, or sound. In this mapping, signs become tokens and numerical representations inside the language model.

Syntax

The combinatorial structure of a language: the rules by which signs combine into well-formed expressions. The model absorbs syntax as learned structural regularities, never as a declared grammar.

Semantics

What linguistic expressions mean — reference, sense, truth conditions. Implicit in the model (encoded associations); explicit in a Semantic Web (identified, declared relationships).

Pragmatics

How situation fixes meaning: speaker, audience, time, place, purpose. The “bank” example — financial institution vs. river edge — is decided by pragmatics, which the model approximates as runtime context.

Langulator

The author's coinage for a large language model: a Language Emulation Machine. It emulates the systematic use of signs, syntax, semantics, and pragmatics without possessing language as a speaker does.

Token

The model's unit of text: sub-word pieces mapped to numerical vectors. Tokens are what signs become inside the Langulator — the raw material of all computation.

Implicit semantics

Meaning encoded numerically and recoverable only by running the model — fluent and approximate, closed to deterministic inference.

Explicit semantics

Meaning identified by URIs, declared as RDF relationships, and shared via ontologies — machine-computable and open to logic.

Denotation

What a sign points to in the world — the “provides” role assigned to signs. The map–territory distinction is a claim about denotation: the sign is not the thing.

Distributional hypothesis

The linguistic principle that words occurring in similar contexts have similar meanings — the theoretical ancestor of the contextual associations a language model learns.

Transformer

The neural architecture (Vaswani et al., 2017) whose attention mechanism lets a model weigh context when computing representations — the machinery behind learned contextual association.

Data space

The term for where ground truth lives: databases, knowledge bases, filesystems, and APIs — the territory the language-map describes, reached through governed interfaces.

Knowledge Graph Explorer
Knowledge Graph

KG Explorer

Interactive visualization of this collection's knowledge graph. Node labels sit below their circles; click a node to open it in the resolver.

— nodes / — links
Click outside to release zoom
Types:
Classes
Instances

⚙ Advanced Settings

-400
90px
Toggle predicates
SPARQL Workbench
Query

SPARQL Workbench

Run live SPARQL queries against the companion knowledge graph. Recipes open closed by default.

Sample queries from the graph

Entity-type summary (canonical)

SAMPLE-based per-type census of the collection graph: type, one sample entity, its label, and the count — ordered by frequency.

Run live query (query entity)

PREFIX rdf: <http://www.w3.org/1999/02/22-rdf-syntax-ns#>
PREFIX rdfs: <http://www.w3.org/2000/01/rdf-schema#>
PREFIX schema: <http://schema.org/>

SELECT
    ?type
    (SAMPLE(?s) AS ?sampleEntity)
    (SAMPLE(?label) AS ?sampleLabel)
    (COUNT(?s) AS ?entityCount)
WHERE {
    GRAPH <https://linkeddata.uriburner.com/DAV/demos/daas/llms-and-language-muse.ttl> {
        ?s rdf:type ?type .
        OPTIONAL { ?s rdfs:label ?label }
    }
}
GROUP BY ?type
ORDER BY DESC(?entityCount)
The four language components and their LLM realizations

Each LanguageComponent with the role it provides and what it becomes inside the Langulator.

Run live query (query entity)

PREFIX schema: <http://schema.org/>
PREFIX : <https://community.openlinksw.com/t/llms-and-language/6504#>

SELECT ?component ?provides ?realization
WHERE {
    GRAPH <https://linkeddata.uriburner.com/DAV/demos/daas/llms-and-language-muse.ttl> {
        ?component a :LanguageComponent ;
                   schema:name ?name ;
                   :providesRole ?provides ;
                   :llmRealization ?realization .
    }
}
ORDER BY ?name
The division-of-labor pipeline in order

The nine HowTo steps from human intent to human-readable outcome, with positions.

Run live query (query entity)

PREFIX schema: <http://schema.org/>

SELECT ?step ?title ?position
WHERE {
    GRAPH <https://linkeddata.uriburner.com/DAV/demos/daas/llms-and-language-muse.ttl> {
        ?step a schema:HowToStep ;
              schema:name ?title ;
              schema:position ?position .
    }
}
ORDER BY ?position
FAQ questions with their answers

All twelve named questions joined to their accepted answers.

Run live query (query entity)

PREFIX schema: <http://schema.org/>

SELECT ?question ?answer
WHERE {
    GRAPH <https://linkeddata.uriburner.com/DAV/demos/daas/llms-and-language-muse.ttl> {
        ?q a schema:Question ;
           schema:name ?question ;
           schema:acceptedAnswer ?a .
        ?a schema:text ?answer .
    }
}
ORDER BY ?question