Overview
Language is the map. The territory is everything it leaves out.
A machine that emulates language
Language is the systematic use of signs, syntax, and semantics to encode and decode information — with pragmatics supplying the runtime context that fixes meaning. That makes a large language model something more specific than its name suggests: not a machine that knows language, but a machine that emulates it. A Language Emulation Machine. A Langulator.
Four components, four counterparts
Each component of language has a counterpart inside the model. Signs — the carriers of meaning — become tokens and their numerical representations. Syntax — the structural rules — becomes learned regularities absorbed from training text. Semantics — what expressions mean — becomes contextual associations, also learned, also numerical. Pragmatics — meaning in situation — becomes the runtime context: the prompt, the conversation, the task at hand. Nothing in the model is declared; everything is encoded.
Signs
In the model: Tokens and numerical representations
Signs are the perceptible forms that carry meaning: words, marks, sounds. Inside a large language model they become tokens — sub-word units mapped to numerical vectors — the raw material everything else is computed over.
Syntax
In the model: Learned structural regularities
Syntax is the system of rules for combining signs into well-formed expressions. The model never receives a grammar; it absorbs structural regularities statistically from the shape of training text, and reproduces them fluently.
Semantics
In the model: Learned contextual associations
Semantics is what signs and structures mean — reference, sense, truth conditions. In the model, meaning is not declared anywhere; it is encoded implicitly as contextual associations between numerical representations, recoverable only by running the model.
Pragmatics
In the model: Runtime context
Pragmatics is how situation fixes meaning: who is speaking, to whom, about what, right now — the difference between a financial bank and a river bank. For the model, that situation is the runtime context: the prompt, the conversation history, the task.
What the map leaves out
Hence the central image: language is the map, and what it describes is the territory. A map is useful precisely because it leaves things out — and dangerous when it is mistaken for the ground itself. An LLM's semantics are implicit: numerically encoded, probabilistic, and unavailable to deterministic inference. A Semantic Web's semantics are the opposite: explicit relationships, identified by URIs, declared in ontologies, and machine-computable by logic. The conclusion follows from the contrast: LLMs don't make explicit semantics obsolete — they make explicit semantics easier to use. Language for the fuzzy parts. Semantic Web for the precise parts. Loose coupling between them.
“Language is the map.”
“Language for the fuzzy parts. Semantic Web for the precise parts. Loose coupling between them.”
The division of labor
Five layers keep the fuzzy and the precise apart — each with a job the other cannot do.
LLM — the Langulator
The model emulates language: it parses intent from prose, holds the fuzzy, contextual, ambiguous parts of a task, and renders results back into natural language. Its semantics stay implicit — that is what makes it fluent, and what makes it unreliable for precision.
RDF + Linked Data + Ontologies
Relationships identified by URIs, declared in shared vocabularies, traversable over HTTP, and available to deterministic inference. Where the model's meaning is encoded and approximate, here it is declared and exact.
Logic
Rules applied to declared relationships yield conclusions that follow necessarily — no sampling, no temperature. The precise parts of a task land here.
Agent harness
The scaffolding that lets an agent act: choosing skills and tools, sequencing calls, enforcing policy. It couples the fluent layer to the precise layer without confusing the two.
Data spaces
Where the ground truth lives. Documents, tables, graphs, and services — the territory the map describes, reachable through governed interfaces.
Cited works
Two Kinds of Semantics
The model's way against a Semantic Web way — six aspects, one verdict: keep the fuzzy and the precise loosely coupled.
| Aspect | The model's way — Langulatorimplicit semantics | A Semantic Web wayexplicit semantics |
|---|---|---|
| Semantics | Implicit — numerically encoded in model weights, learned from text. | Explicit — identified by URIs, declared as RDF, shared via ontologies. |
| Representation | Tokens and embeddings — meaning by statistical association. | URIs and triples — meaning by declaration. |
| Identity | None — entities blur into proximity in vector space. | Stable identity through IRIs: the same thing is the same URI everywhere. |
| Inference | Probabilistic generation — fluent, approximate, uncheckable. | Deterministic logic over declared relationships — computed, not guessed. |
| Coupling | Fused inside the model — the fuzzy and the precise collapse together. | Loose — language for the fuzzy parts, a Semantic Web for the precise parts. |
| Failure mode | Fluent output mistaken for ground truth. | Statements machines can check, ground, and compose. |
The model’s way — Langulator
A Semantic Web way
HowTo: The Division-of-Labor Pipeline
Nine steps of the division-of-labor pipeline — human to human, with the fuzzy and the precise parts kept apart.
Start from the human
Every task begins with a person and an intent. The pipeline's job is to carry that intent through machinery without losing it — which is why the first and last steps are human.
Express the intent in natural language
The human states what they want in ordinary prose. Natural language is the highest-bandwidth interface a person has; the pipeline meets them there instead of demanding a query language.
Let the Langulator parse the fuzzy parts
The LLM — the Language Emulation Machine — handles everything probabilistic: disambiguating intent, filling in context, tolerating ambiguity. Its implicit semantics are a feature here, not a bug.
Ground entities in RDF, Linked Data, and ontologies
Translate the parsed intent into identified things and declared relationships: URIs for entities, RDF for the relations, ontologies for shared meaning. This is where implicit becomes explicit.
Operate over data spaces
Run the precise parts against databases, knowledge bases, filesystems, and APIs — the territory, not the map. Governed access decides what the agent may touch.
Apply logic for deterministic results
Where the answer must follow necessarily, use inference over declared relationships — not sampling. Precision is computed, not guessed.
Return through the Langulator
Hand the precise results back to the model for rendering: summaries, explanations, prose shaped for the human's situation. Fluency is the model's home turf.
Deliver in natural language
The human receives the outcome as language — the same medium the intent arrived in. The machinery in the middle stays invisible.
Keep the coupling loose
The standing rule: language for the fuzzy parts, a Semantic Web for the precise parts, and only loose coupling between them. Never let the probabilistic layer pretend to be the precise one.
FAQ
Twelve named questions, twelve named answers — each question heading is its entity.
The author's coinage for a large language model, short for Language Emulation Machine. The name is the argument: the model does not possess language the way a speaker does — it emulates the systematic use of signs, syntax, semantics, and pragmatics, reproducing their outward behavior from learned numerical patterns.
Signs (the carriers of meaning), syntax (the combinatorial structure), semantics (what expressions mean), and pragmatics (how situation fixes meaning). Each has a counterpart inside the model: tokens and numerical representations, learned structural regularities, learned contextual associations, and runtime context.
That meaning is fixed by situation, not by the sign alone. The word “bank” denotes a financial institution in one context and the edge of a river in another. Pragmatics — the runtime context — selects the reading. The model does the same work with its context window that a speaker does with a situation.
Signs become tokens and numerical representations; syntax becomes learned structural regularities absorbed from training text; semantics becomes learned contextual associations between those representations; pragmatics becomes the runtime context — the prompt, the conversation, the task. Nothing is declared; everything is encoded.
Nothing is false about it, but it undersells what the prediction machinery achieves. Next-token prediction is the training objective; language emulation — reproducing the systematic behavior of signs, syntax, semantics, and pragmatics — is what that objective produces at scale. “Langulator” names the product, not the mechanism.
Meaning that is encoded but not declared. In an LLM, what a token “means” exists only as patterns across numerical representations — recoverable by running the model, unavailable to inspection, and closed to deterministic inference. Fluent, approximate, and fundamentally probabilistic.
Meaning that is identified and declared: entities named by URIs, relationships stated as RDF triples, vocabularies shared through ontologies. Explicit semantics are machine-computable — a logic engine can apply rules to them and derive conclusions that follow necessarily.
The difference between the two kinds of semantics. An LLM associates “Paris” with “France” through learned contextual association — strong, fluent, and implicit. Explicit semantics instead declares the relationship (Paris is the capital of France) as an identified, declared, machine-computable fact that logic can build on.
The central image: language describes the world the way a map describes terrain. A map is useful precisely because it leaves things out — and dangerous when mistaken for the ground itself. Confusing an LLM's fluent map for the territory is the characteristic failure mode.
Five layers, loosely coupled: the LLM handles probabilistic natural-language work; RDF, Linked Data, and ontologies supply explicit machine-computable semantics; logic supplies deterministic inference; the agent harness supplies controlled action and skill/tool selection; data spaces — databases, knowledge bases, filesystems, APIs — supply ground truth.
Agents act, and acting on implicit semantics alone means acting on approximations. Grounding the agent's entities and relationships in explicit, identified, declared semantics gives it something deterministic to stand on: the fluent layer interprets intent, the precise layer constrains what may be concluded and done.
Not smarter models — the point is that raw capability was never the gap. The missing piece was the explicit semantic layer that lets machine intelligence be checked, grounded, and composed: identified entities, declared relationships, and logic. LLMs don't make explicit semantics obsolete; they make explicit semantics easier to use.
Glossary
Twelve defined terms, each linked to its canonical home.
Sign
The perceptible carrier of meaning — a word, mark, or sound. In this mapping, signs become tokens and numerical representations inside the language model.
Syntax
The combinatorial structure of a language: the rules by which signs combine into well-formed expressions. The model absorbs syntax as learned structural regularities, never as a declared grammar.
Semantics
What linguistic expressions mean — reference, sense, truth conditions. Implicit in the model (encoded associations); explicit in a Semantic Web (identified, declared relationships).
Pragmatics
How situation fixes meaning: speaker, audience, time, place, purpose. The “bank” example — financial institution vs. river edge — is decided by pragmatics, which the model approximates as runtime context.
Langulator
The author's coinage for a large language model: a Language Emulation Machine. It emulates the systematic use of signs, syntax, semantics, and pragmatics without possessing language as a speaker does.
Token
The model's unit of text: sub-word pieces mapped to numerical vectors. Tokens are what signs become inside the Langulator — the raw material of all computation.
Implicit semantics
Meaning encoded numerically and recoverable only by running the model — fluent and approximate, closed to deterministic inference.
Explicit semantics
Meaning identified by URIs, declared as RDF relationships, and shared via ontologies — machine-computable and open to logic.
Denotation
What a sign points to in the world — the “provides” role assigned to signs. The map–territory distinction is a claim about denotation: the sign is not the thing.
Distributional hypothesis
The linguistic principle that words occurring in similar contexts have similar meanings — the theoretical ancestor of the contextual associations a language model learns.
Transformer
The neural architecture (Vaswani et al., 2017) whose attention mechanism lets a model weigh context when computing representations — the machinery behind learned contextual association.
Data space
The term for where ground truth lives: databases, knowledge bases, filesystems, and APIs — the territory the language-map describes, reached through governed interfaces.
Knowledge Graph Explorer
KG Explorer
Interactive visualization of this collection's knowledge graph. Node labels sit below their circles; click a node to open it in the resolver.
⚙ Advanced Settings
SPARQL Workbench
SPARQL Workbench
Run live SPARQL queries against the companion knowledge graph. Recipes open closed by default.
Sample queries from the graph
Entity-type summary (canonical)
SAMPLE-based per-type census of the collection graph: type, one sample entity, its label, and the count — ordered by frequency.
PREFIX rdf: <http://www.w3.org/1999/02/22-rdf-syntax-ns#>
PREFIX rdfs: <http://www.w3.org/2000/01/rdf-schema#>
PREFIX schema: <http://schema.org/>
SELECT
?type
(SAMPLE(?s) AS ?sampleEntity)
(SAMPLE(?label) AS ?sampleLabel)
(COUNT(?s) AS ?entityCount)
WHERE {
GRAPH <https://linkeddata.uriburner.com/DAV/demos/daas/llms-and-language-muse.ttl> {
?s rdf:type ?type .
OPTIONAL { ?s rdfs:label ?label }
}
}
GROUP BY ?type
ORDER BY DESC(?entityCount)The four language components and their LLM realizations
Each LanguageComponent with the role it provides and what it becomes inside the Langulator.
PREFIX schema: <http://schema.org/>
PREFIX : <https://community.openlinksw.com/t/llms-and-language/6504#>
SELECT ?component ?provides ?realization
WHERE {
GRAPH <https://linkeddata.uriburner.com/DAV/demos/daas/llms-and-language-muse.ttl> {
?component a :LanguageComponent ;
schema:name ?name ;
:providesRole ?provides ;
:llmRealization ?realization .
}
}
ORDER BY ?nameThe division-of-labor pipeline in order
The nine HowTo steps from human intent to human-readable outcome, with positions.
PREFIX schema: <http://schema.org/>
SELECT ?step ?title ?position
WHERE {
GRAPH <https://linkeddata.uriburner.com/DAV/demos/daas/llms-and-language-muse.ttl> {
?step a schema:HowToStep ;
schema:name ?title ;
schema:position ?position .
}
}
ORDER BY ?positionFAQ questions with their answers
All twelve named questions joined to their accepted answers.
PREFIX schema: <http://schema.org/>
SELECT ?question ?answer
WHERE {
GRAPH <https://linkeddata.uriburner.com/DAV/demos/daas/llms-and-language-muse.ttl> {
?q a schema:Question ;
schema:name ?question ;
schema:acceptedAnswer ?a .
?a schema:text ?answer .
}
}
ORDER BY ?question