← All posts

Engineering · AI

Context Engineering: Why Your AI Agent Forgets What You Told It

Mihajlo Petrović5 min read

Agents do not get worse over a long session - their context does. What actually fills a context window, why quality degrades before you hit the limit, and five habits that make the model noticeably smarter for free.

You start a session, explain the architecture, agree on an approach, and get good work for twenty minutes. Then the agent quietly starts ignoring your conventions, reintroduces a bug you fixed together an hour ago, and writes a helper that already exists.

Nothing broke. This is the system working exactly as designed — and once you understand why, most "the model got worse" complaints turn out to be context problems with a fix.


The Model Has No Memory

This is the single most useful thing to internalise: a language model is stateless. It does not remember your last message. Every turn, the entire conversation is re-sent — system prompt, instruction files, all previous messages, every file the agent read, every tool result it received.

That means the "conversation" is really a document that grows with each turn. And a growing document has three consequences that bite in order:

  1. It gets expensive. You pay for the whole thing on every single turn.
  2. It gets slow. More input tokens, more time to first token.
  3. It gets worse. Long before you hit the limit, quality degrades — the signal you care about is buried under thousands of tokens of noise.

Current frontier models advertise context windows around a million tokens. That is a capacity, not a target. Filling it is a choice, and usually a bad one.


What Actually Eats Your Context

In a coding session, the conversation is rarely the problem. The bulk is almost always one of these:

Whole-file reads. The agent needs one function and reads all 900 lines. Do that six times and you've spent tens of thousands of tokens to look at maybe 200 lines of relevant code.

Tool result spam. A test run that prints 400 lines of passing output. A git log with no limit. An npm install transcript. All of it stays in context for the rest of the session.

Dead ends. The agent tries an approach, it fails, you redirect. The failed attempt — including all the code it wrote — is still sitting in the conversation, quietly competing with your correction for the model's attention.

Repeated near-copies. File read, file edited, file re-read to verify. Now three slightly different versions of the same file are in context, and the model has to work out which one is current.


Five Things That Actually Help

1. Put the durable rules in a file, not in chat

Anything you'd have to repeat in a new session belongs in a project instructions file (CLAUDE.md, .cursorrules, whatever your tool reads), checked into the repo. Folder structure, state-management approach, naming conventions, the commands to run tests and lint.

The highest-value half is usually the prohibitions: no new dependencies without asking, no any, no inline styles, never swallow an exception. Positive instructions are guessable; your specific prohibitions are not.

Keep it short. An instructions file that grows to 500 lines is itself context bloat, and it's loaded on every single turn.

2. Read structure before you read content

Tools that return a file's shape — the symbol list, the outline — cost a fraction of reading the file. Let the agent orient itself with structure, then read only the region that matters. This is exactly why wiring up a language server pays off, as I wrote about in the agents-in-practice post.

3. Start a fresh session more often than feels natural

If a session has gone through two dead ends and a scope change, its context is mostly archaeology. Write down the current state in two sentences, start clean, paste those two sentences in.

The reflex to preserve a long session is a false economy: you're paying for and reasoning over a transcript in which most of the content is now wrong.

4. Delegate the reading

When a sub-task requires reading a lot to produce a little — "find every place we validate an IBAN" — hand it to a sub-agent with its own context. It burns 50,000 tokens on the search and returns a 200-token answer. Your main session gets the answer without the archaeology.

5. Keep the stable part stable

If you're building an LLM feature rather than using an agent, prompt caching turns a large fixed prefix into a cheap one — but only if the prefix is byte-identical between calls. A timestamp in the system prompt, an unsorted JSON blob, a tool list assembled in a different order, and your cache hit rate silently drops to zero.

Order matters too: tools, then system prompt, then messages. Stable content first, volatile content last. Verify with the cache-read token counts in the API response rather than assuming — a broken cache doesn't raise an error, it just costs more.


The Reframe

The job used to be writing the right prompt. It's now curating the right context — which is a different skill, closer to editing than to writing.

A useful test before any nontrivial agent task: if a competent new colleague were handed exactly what's in this context window, could they do this task correctly? Not "is the prompt clever" — is the information actually there, and is the irrelevant information actually gone.

Most of what people experience as model quality is context quality. Fix the context and the model gets noticeably smarter for free.

  • #ai
  • #context engineering
  • #llm
  • #prompt caching
  • #developer productivity

Written by

Mihajlo Petrović

Software engineer in Belgrade. Builds his own products and the AI automations that keep them running.

Have a task that repeats every week?

Tell me about it. If it can be automated well, I will show you how. If it cannot, I will say that too.

Tell me what to automate