In short

Context engineering is the discipline of deciding, for each call to a model, which information to place in the context window, in what order, and in what form. The term was popularized in June 2025, notably by Andrej Karpathy and Tobi Lütke (Shopify). The distinction from prompt engineering is clear-cut: the prompt is an instruction; the context is the complete informational state passed to the model.


Why the term emerged

In plain terms : the word “prompt” has become synonymous with “what you type into ChatGPT”. Engineers needed a term to designate everything else — history, tools, connected data — which often weighs far more than the instruction itself.

In June 2025, Tobi Lütke (CEO of Shopify) tweeted that “context engineering” describes the core skill better than “prompt engineering”: providing all the context needed for the task to be solvable by the model. Andrej Karpathy followed suit: in any production LLM application, the real work is “the delicate art and science of filling the context window with just the right information for the next step”.

Simon Willison, who had defended the term “prompt engineering”, acknowledged the problem: the public had associated “prompt” with simply typing into a chatbot. The term “context engineering” would benefit from a public definition closer to the actual technical intent.

This terminological shift reflects a practical evolution: in complex agent systems, the system prompt represents only a fraction of what the model receives. Conversation history, tool results, data retrieved in real time, dynamically selected examples — all of this forms a context whose construction is an engineering job in its own right.


What context engineering covers

In plain terms : doing context engineering means deciding, for each call to the model, “what do I show it, in what order, and what do I leave out”. Taxonomies vary, but the best practices converge: filter, compress, retrieve just in time.

Several taxonomies are in circulation. Anthropic’s (September 2025) distinguishes five components of an agent’s context:

  1. System prompt — the baseline instructions. Anthropic recommends a “Goldilocks” level of abstraction: neither too rigid to block adaptation, nor too vague to lose the model.
  2. Tools — tool definitions should return token-efficient information. A verbose return pollutes the window.
  3. Examples — well-chosen examples are worth more than a long explanatory instruction.
  4. Message history — history accumulates. Purging and summarizing regularly is an engineering decision, not a detail.
  5. External data — “just-in-time” retrieval rather than loading everything upfront.

LangChain proposes four complementary strategies: Write (save outside the window, persistent memory), Select (retrieve the right information at the right time via RAG), Compress (summarize or truncate what can be), Isolate (multi-agent setups with separate contexts to avoid contamination).

Phil Schmid (HuggingFace, June 30, 2025) offers the most operational definition: “the discipline of designing and building dynamic systems that provide the right information, to the right tools, in the right format, at the right time.”


Context rot: the empirical evidence

In plain terms : the more you fill the context window, the less well the model uses it. Information placed in the middle of a long text? The model “sees” it less clearly than at the beginning or the end. Several independent studies measure this degradation.

The term “context rot” refers to the degradation of model performance as context grows. This isn’t an intuition — it’s documented by several independent studies.

Lost in the Middle (Liu et al., Stanford/UC Berkeley/Samaya AI, TACL 2024): models perform better when useful information is placed at the beginning or end of the context. Performance follows a U-shaped curve — a primacy bias and a recency bias combined. When information moves from the beginning or end toward the middle, degradation exceeds 30%. On GPT-3.5-Turbo in multi-document question answering, the drop exceeds 20%.

RULER (NVIDIA, COLM 2024): 17 models evaluated from 4K to 128K tokens across 13 representative tasks. Result: only half of the models claiming to handle 32K tokens or more reach the quality threshold at that length. Almost all models see their performance drop significantly with length, even those posting near-perfect scores on “needle in a haystack” (NIAH) tasks.

Chroma Research (July 2025): 18 models tested, consistent degradation with input length across all models. A counter-intuitive result: disordered contexts improve performance over coherent ones — perhaps because the implicit structure steers the model toward certain parts rather than others.

LongCodeBench (Rando et al., 2025): on code tasks, Claude 3.5 Sonnet drops from 29% to 3% accuracy when the context moves from 32K to 256K tokens.

These data invalidate the hypothesis that larger windows would solve the problem. Gemini 1.5 Pro, GPT-4.1, Llama 4 all announce windows of a million tokens or more. But effective performance diverges from the declared window: nominal size does not guarantee actual ability to exploit the whole of that window.


Toward automated context management

In plain terms : rather than writing by hand what you give to the model, the idea is to let another piece of code decide — add, summarize, prune based on feedback. Early results are promising, but the right balance between “compressing” and “not losing anything important” remains an open problem.

The ACE framework (Agentic Context Engineering, Zhang et al., arXiv 2510.04618, October 2025) proposes to treat context as an “evolving playbook” that updates based on performance feedback. Without labeled supervision, with a small open-source model, ACE gains +10.6% on agent benchmarks and +8.6% on financial tasks.

The 12-Factor Agents framework (humanlayer, GitHub), inspired by Heroku’s 12-Factor Apps, states a similar principle: “Own your context window” — explicitly manage which data is included and how it is summarized. Agents are designed as “stateless reducers”: each operation is independent and reproducible.

Documented limitations: Anthropic (September 2025) notes that overly aggressive compaction leads to the loss of subtle context whose importance only surfaces later. The compression/retention tradeoff is not algorithmically solved to date.


Emerging discipline or rebranding?

The critique exists. Several experienced developers note that the importance of context was already known before the term circulated. A more radical argument (Serge Liatko, LAWXER): prompt engineering and context engineering are both already outdated — the real work is the architecture of automated workflows where every piece of context is generated, managed, and delivered by code.

The opposing position is defended by LangChain and Anthropic: prompt engineering optimizes an isolated instruction; context engineering is a system discipline that manages the complete informational state at inference time. Phil Schmid puts it this way: agent failures are “context failures”, not “model failures”.

The tension is real. LangChain itself acknowledges that “even with all the context, how you assemble it in the prompt still matters”. Context engineering does not replace the care given to drafting instructions — it subsumes it in a broader scope.

What is currently missing: no standardized framework exists to measure “context quality”. Available benchmarks (RULER, NIAH, LoCoBench) are task-specific and do not transfer easily to general practice.


What to remember

  • Context engineering refers to managing everything the model receives at inference — not just the prompt, but history, retrieved data, tool returns, examples.
  • Peer-reviewed studies (Lost in the Middle, RULER) document performance degradations of 20 to 30% depending on information position in the context, and sharper drops as the window lengthens.
  • Larger context windows do not solve the problem: effective performance systematically diverges from the declared window.
  • Automating context management (ACE) shows measured gains on benchmarks, but the tradeoff between compaction and retention of critical context remains open.
  • The question “new discipline or rebranding?” is not settled in the literature — the absence of standardized metrics for context quality leaves the debate empirically unresolved.

Skill signals

No skill shortcoming identified on this deliverable. Figures from single unrecouped sources (300% RAG, 15× multi-agent, LangMem figures) were deliberately left out of the body of the article to respect the [NOT VERIFIED] rule — they would have required inline tagging that would have weighed down the text without providing real evidence. The editorial choice: only cite data from peer-reviewed papers or cross-recouped sources.