In brief
Large language models (LLMs) are trained on billions of words in natural language, and yet the engineers who drive them in production use XML, YAML, and structured Markdown extensively — not ordinary sentences. This paradox is not an implementation detail: it reflects a structural property of LLMs that mass-market communication often obscures. Understanding why structured language outperforms free prose is the foundation of serious context engineering.
Imagine an interpreter who has listened to millions of hours of conversations in every language on Earth, but who understands none of these exchanges — they only know which sentence statistically follows which sentence. When you speak to them in free prose, they must guess among an ocean of possible continuations. When you give them a form to fill in — name here, date there, reason in the middle — you remove the ambiguity: at each field, they only have to complete what that kind of field usually calls for. LLMs work this way, and structured language plays the role of the form.
The myth of “speaking naturally”
In plain terms : speaking to an LLM like a human works for casual conversation, not for driving it. The more precise the task, the more the form beats the conversation.
The most visible interface of LLMs — a text box where you type in French or English — has established the idea that it is enough to address the model as you would a human. This intuition is partially true for simple conversational uses. It becomes false as soon as one seeks precise, reproducible, and composable behavior.
An LLM does not “understand” an instruction in the human sense. It predicts the statistically probable continuation of a token sequence. When you write “respond in JSON”, the model has learned, from its training data, that this phrase is often followed by a JSON block. The structure of the request activates completion patterns, not an independent comprehension mechanism.
The practical consequence: a vague instruction produces vague behavior. A structured instruction produces structured behavior — not because the model is “more comfortable”, but because structure reduces the space of possible completions.
Why structured language performs better
In plain terms : XML, YAML, and Markdown set up partitions that the model has seen a thousand times during its training. Each partition reduces the number of plausible continuations the model can produce — and therefore its margin of error.
Unambiguous delimitation
Natural language is rich in contextual ambiguity. “Summarize the following text in 3 points” leaves open the question of what constitutes “the following text” if the prompt contains multiple blocks. An XML tag like <text_to_summarize>...</text_to_summarize> delimits without ambiguity.
Anthropic explicitly documents that XML improves performance on structured tasks, particularly when the prompt contains multiple sections (context, constraints, task, examples). The reason: tags constitute separators that the model has learned to recognize as strong semantic boundaries during its training on code and technical documentation.
Lossless compression
A concrete test: a set of rules expressed in free prose over 425 lines can be reformulated in structured Markdown (fixed sections, lists, tables) over 119 lines — with the same constraint compliance rate measured in use [NOT VERIFIED on an independent corpus]. The gain does not come from removing information but from eliminating noise: transitions, reformulations, implicit repetitions typical of prose style.
Structured Markdown forces explicit granularity. Each bullet point is a distinct unit. Each heading delimits a domain. The model treats these separators as strong signals, which continuous prose does not provide.
The effect of structured examples
Among the most effective patterns for driving an LLM: a structured example that decomposes the process into named steps (Need → Research → Result → Action → Verification) rather than a narrative example (“here is how I solved a similar problem…”). Analysis of 116 real AI workflow configuration files shows that 93% use fixed-section structures — none uses free prose as the primary format.
The likely explanation: the named structure creates a constrained completion space. After “Result:”, the model knows it must produce a result, not a question or a transition. Prose leaves this choice open.
What the 116 files reveal
In plain terms : when you look at what practitioners actually do — not what they say they do — 93% drive LLMs in structured Markdown. The convention wasn’t imposed, it emerged.
A systematic analysis of 116 AI workflow configuration files (skills, system instructions, orchestrator prompts) reveals a ground-level consensus:
| Format | Frequency | Typical use |
|---|---|---|
| Structured Markdown (sections + lists) | 93% | Instructions, rules, constraints |
| YAML frontmatter | ~80% | Metadata, classification |
| XML tags | ~40% | Block delimitation, inputs/outputs |
| Free prose | <5% | Introduction or context only |
Median size: 240 lines. Optimal observed size for an instruction file: 100–300 lines. Beyond that, information density per token decreases and the model dilutes its conformity to constraints.
This corpus represents active practitioners, not theoretical recommendations. The near-unanimous choice of structured Markdown is the result of empirical iteration, not a top-down decision.
Limitations and nuances
In plain terms : the form wins for controlled tasks; it stifles free creation, repels non-technical users, and varies from one model to another. A well-tagged XML never excuses a poorly thought-out instruction.
This picture is not complete. Structured language is not universally superior:
- For open creative tasks, a strong structural constraint can limit generation. A novelist driving an LLM in YAML risks getting rigid prose.
- For non-technical users, XML and YAML create a real barrier to entry. Consumer-facing interfaces simplify appropriately.
- Optimization is model-dependent: a format highly effective on GPT-4 may behave differently on Claude or Mistral. “Gain” measurements often circulate without specifying the model and version [NOT VERIFIED across multiple models].
- Structure can mask defective prompts: perfectly tagged XML with a poorly defined task remains a bad prompt.
Key takeaways
- LLMs generate by statistical token completion: structure reduces the space of possible responses and improves conformity.
- XML, YAML, and structured Markdown outperform free prose for complex, multi-section, or repeatable instructions — not by magic, but because these formats appear in mass in training data with consistent patterns.
- Compressing prose → structured format can reduce length by a factor of 3 without functional loss; experience shows that the constraint compliance rate remains stable or even improves.
- Analysis of public AI configuration repositories confirms that practitioners converge on structured Markdown independently of official guidelines.
- “Speaking naturally” remains valid for conversational uses; it becomes an anti-pattern as soon as one seeks reproducibility and fine-grained behavioral control.