What Is Context Engineering? The Skill That Replaced Prompting
Context engineering is what prompt engineering grew into. The short version: instead of only polishing how you ask a question, you design everything the model sees when it answers — the instructions, the retrieved documents, the tool results, the conversation history, the notes from earlier sessions. The term took off in mid-2025, and by now Anthropic, LangChain, and most of the industry use it seriously. If you build anything with LLMs beyond a toy chatbot, this is the discipline that decides whether your thing works.
The name is new. The problem is not. Let me walk through where the term came from, what actually counts as context, why the biggest models made the problem worse, and what the practical techniques look like.
What context engineering is, and where the term came from
The timeline matters, because it tells you what people were frustrated with.
June 2025: Tobi Lütke, Shopify's CEO, used the phrase in an internal memo that later went public — context engineering as "the art of providing all the context for the task to be plausibly solvable by the LLM." Days later, Andrej Karpathy amplified it: he preferred context engineering over prompt engineering, defining it as "the delicate art and science of filling the context window with just the right information for the next step." Note what Karpathy lists as context-engineering moves: few-shot examples, RAG, tools, state, history, compacting. Stuff practitioners had been doing for years without a name that fit.
Simon Willison, who coined "prompt engineering" back in 2022, put his finger on why the old term stopped working: its inferred definition had slid into "a laughably pretentious term for typing things into a chatbot." The new term describes the actual job. Then in September 2025, Anthropic published an engineering guide calling context engineering the natural evolution of prompt engineering, and that was that — the vocabulary flipped.
One more definition worth keeping, from a GitHub repository that synthesized 1,400 papers on the subject: prompt engineering is "what you say"; context engineering is "everything else the model sees." I find that the fastest way to explain it to someone who's heard of prompting.
What actually counts as context
Here's the inventory. At any given moment, the context is everything in the model's context window — roughly seven kinds of things:
- The system prompt (stable instructions)
- Your current message
- Short-term state — the task in progress, the conversation so far
- Long-term memory — anything carried over from past sessions
- Retrieved knowledge — the R in RAG, documents pulled in at query time
- Tool definitions and tool results
- Required output formats
All seven compete for the same fixed capacity. That's the part people miss. The window is a shared budget — we have a full piece on the context window mechanism if you want how the capacity side works — and every component you add dilutes the attention available to the rest. Context engineering is budgeting that shared space.
Why the term showed up now: bigger windows, less room
This is the counterintuitive core, and honestly the reason the field needed a new word.
You'd think that with models accepting a million tokens, you could stop curating and just dump everything in. The research says the opposite. A study titled "Context Rot" measured what happens as input tokens increase: recall and reasoning both degrade. The effect has a mechanism. Attention is a budget — each token you add thins the attention available to every other token, and since attention models pairwise relationships between tokens, the interactions grow with the square of the length. More stuff, each thing seen more poorly.
But wait, you might say, didn't those long-context models ace the needle-in-a-haystack tests? They did, and that's the trap. The classic test hides a fact in a long document and asks a question whose wording nearly matches the hidden fact — it's a lexical-matching exercise, closer to string search than reasoning. When researchers rewrote the "needle" so the answer required connecting an indirect statement, or added near-miss distractors that look related but aren't the answer, performance fell with length, hard. Real work — "here's a pile of code, find why it breaks" — is the ambiguous, distractor-heavy kind, not the string-match kind.
The number I'd tattoo somewhere: in one benchmark comparison, answering over a full conversation history of ~120,000 tokens did worse than answering over a 300-token version containing just the relevant facts. Twelve thousand times fewer tokens, better answers. Curation isn't a nice-to-have after a certain point; it's the whole game. Half of context engineering is deciding what not to put in.
Context engineering vs prompt engineering
The comparison people actually search for, so let's do it properly. The honest framing, and the one LangChain's co-founder argues: prompt engineering didn't die, it became a subset. His line is that agent failures are, more often than not, context failures — not model failures. When your agent does something dumb, the model usually had everything it needed to be dumb: a bloated history, the wrong documents, three tools that overlap.
| Prompt engineering | Context engineering | |
|---|---|---|
| What you shape | The wording of one instruction | Everything entering the window |
| Form | A static template | A dynamic system assembled per task |
| Core question | How do I ask? | What goes in, what stays out, in what format? |
| Typical moves | Rephrase, add few-shot examples | Wire up RAG, trim history, design tools, keep notes |
A concrete example makes it land. You tell a model, "I want to return this item." With just the prompt, it answers mechanically: "please submit within 7 days." Now build the context: who the customer is, the order details, the store's rules, what was promised earlier in the conversation. Same model, same question — now it says the earbuds are still within the window and the return is already filed. The difference was never the wording of the question. For the full landscape of approaches — where fine-tuning and plain prompting sit relative to this — see our piece on fine-tuning vs RAG vs prompting.
The working principles: what goes in, what stays out
Anthropic's engineering guide is the most useful practical writeup I've read on this, and its organizing principle is beautifully compressed: aim for the smallest set of high-signal tokens. Every interaction, you're asking whether each piece of information earns its place in the budget.
A few principles worth stealing:
Fetch just-in-time. Don't preload everything the model might need; retrieve exactly when it's needed. The guide's own example is Claude Code, which reads the head or tail of a file, globs the directory, greps for a pattern — instead of loading the whole codebase into context. The information arrives on demand, already narrowed.
Write the system prompt at the right altitude. Too vague ("be a helpful assistant") and the model has no direction; exhaustively specific and your rules start colliding with each other. The sweet spot is stable behavioral principles one level up from case-by-case rules.
Design tools for a confused reader. Keep tool outputs compact, keep tool purposes from overlapping, and name them for intent. The guide puts it bluntly: if a human engineer can't tell which of two tools to use, don't expect the agent to. I've debugged exactly this failure — two nearly identical internal tools, agent picking the wrong one, everyone blaming the model.
And the gap that separates a demo from a product: an "AI schedules my meeting" demo runs on a system prompt alone. The working version feeds in available time slots, time zones, room capacity, company booking rules, last week's notes. The model didn't get smarter between the demo and the product. Its context did. This is also why context engineering and AI agents grew up together — an agent is the maximum-context scenario: long-running, multi-tool, state accumulating over hours.
Three techniques for long-running work
When a task runs long enough, the context problem stops being "what do I include" and becomes "what do I do with everything that's accumulated." Three techniques from the Anthropic guide, each answering a different failure:
Compaction. When history grows too long, actively summarize and clear it. The non-obvious part is what gets kept: recent files, architectural decisions, unresolved bugs — while processed tool output gets flushed. Compaction is triage, not blind truncation.
Structured note-taking. Have the agent write down what it needs to remember — a NOTES.md file, a step counter (the Claude-plays-Pokémon setup famously used a thousand-step counter as memory) — instead of trusting it to stay alive inside the window. Notes survive; window contents don't.
Sub-agent architectures. Spawn child agents with their own windows. A sub-agent might burn tens of thousands of tokens doing a messy search, then hand back a 1,000-2,000 token summary. The mess never touches the main context.
And the guide's closing advice, which I appreciate precisely because it's unglamorous: do the simplest thing that works. Most tasks don't need all three. RAG itself, when you think about it, is context engineering's most popular move — pulling in retrieved knowledge just when it's relevant. Start there, add machinery only when something actually breaks.
Quick answers
Is context engineering just a rebrand of prompt engineering? Partly, yes — there's fashion in it. But the substance grew: prompting was one input among many, and the rebrand reflects that the other inputs (retrieval, tools, memory, history management) now do more of the work.
Do I need context engineering if I just use ChatGPT? As a casual user, you're doing light context engineering already — attaching files, pasting the right snippet, starting a fresh chat when the old one gets confused. You just don't call it that.
Does RAG count as context engineering? It's the single most common context-engineering technique: deciding the model sees retrieved facts at the right moment, in a compact form.
Is prompt engineering dead then? No — it's a subset. Wording still matters. It just stopped being the whole job.
If you keep one idea from all this: the model at the other end is fixed, but what it sees is yours to design. The teams winning with LLMs aren't the ones with secret magic prompts. They're the ones treating context as an engineered system — measured, budgeted, and ruthlessly trimmed.