Dev

Context engineering starts with an inventory: you cannot budget a context window you cannot list

Every practice in the guides is a verb of choosing: trim, defer, retrieve, compact. Choosing needs a list. The one Claude Code ships itemizes your configuration by source and size, then reports the whole conversation as a single number, and nowhere says how long anything stays.

Context engineering has a canonical text. Anthropic published it in September 2025, and its definition is still the one people quote:

text
Context engineering refers to the set of strategies for curating and
maintaining the optimal set of tokens (information) during LLM inference,
including all the other information that may land there outside of the prompts.

Read the last clause again. Information that may land there. The definition says, in passing, that a good part of the window is not put there by you. The guide knows this well. It has whole sections on agents that fetch their own data and histories that grow on their own. What it teaches in response is how to choose.

Context engineering guides teach you to curate the window, not to see it

The practices are all verbs of choosing. Keep the system instructions to the minimal set of information that fully describes the behavior you want. Curate a minimal viable set of tools with little overlap. Pick a few canonical examples over a pile of edge cases. Retrieve just in time, holding paths and links in place of the documents they point to. For long work, compact the history, have the agent keep notes outside the window, and hand focused tasks to sub-agents that start clean. Two phrases from the guide carry the frame under all of it:

text
a finite resource with diminishing marginal returns
the smallest possible set of high-signal tokens

I agree with every one of those practices. My objection is to the order. Each one is an edit, and an edit has an object. Trim what, and compact a history made of whose text? The guide asks you to consider the whole state available to the model at a given moment, and it has no step for getting that state in front of you. It is a guide to curation with no inspection workflow, and in an agent harness the part of the window you would recognize as yours is small.

What a Claude Code context window holds before you type

Claude Code's documentation is the clearest account I know of how small. Its walkthrough of a session lists what loads before the first prompt: the harness's own system instructions, the auto memory index, environment info, the names of MCP tools, one-line descriptions of skills, your user-level CLAUDE.md and the project CLAUDE.md. Every one of those is marked as invisible in the terminal. The token counts in the walkthrough are labeled illustrative, but the proportion is the point: the startup blocks add up to roughly 7,850 tokens, and the prompt the user then types is 45. The walkthrough draws the conclusion itself:

text
Your prompt is tiny compared to what's already loaded.

It keeps going once work starts. A file read shows in the terminal as a one-line notice while the 2,400 tokens of file content go only to the model. A path-scoped rule loads because the agent touched a matching file, and you see that it loaded, not what it says. A formatting hook reports back through a field that enters context and never appears on screen. None of this is secret. It is all in the docs. You just did not choose any of it at the moment it arrived.

So the window has several kinds of author before the session is a minute old. You wrote the prompt and, some time ago, the instruction files. The vendor wrote the system instructions. Whoever published each MCP server and each skill wrote its description. Whoever last committed to the repository wrote the files being read, and I have argued before that a file a stranger can edit is a different trust level from a rule only you can write. An earlier run of the model wrote the memory index.

Claude Code's /context is already half of a context inventory

The obvious reply is that the listing exists. It does, and it is better than the guides let on. /context prints a breakdown of the window by category with token counts, and /context all expands it per item. On v2.1.296, run as claude -p '/context', the report has one table per kind of configuration. Schematically, with the section name on the left and its column headers on the right:

text
Estimated usage by category    Category   | Tokens | Percentage
MCP Tools                      Tool       | Server | Tokens
Custom Agents                  Agent Type | Source | Tokens
Memory Files                   Type       | Path   | Tokens
Skills                         Skill      | Source | Tokens

That is a real inventory of the configuration. In the same August essay I said that what keeps a layered setup legible is being able to answer one narrow question: which instruction reaches which agent, and who was allowed to put it there. This report answers the first half for the files on disk. It has a cost column. Its Source, Server and Path columns are provenance clues: they say where a block was loaded from, a plugin or your own directory, this file or that one. They do not say who can change it, which is the half of the question that matters for trust, but they are where you would start. It is also the tool for checking, in your own session, the figure I argued from when I called tool definitions a budget.

The listing is itself software with a history. Before v2.1.196 the Skills row counted the full text of every skill description and could show a figure several times larger than what the model received. Before v2.1.280, an AGENTS.md that Claude Code read directly was in the window and absent from the list, which is the gap I described when AGENTS.md support shipped. Both are fixed now. I bring them up because a listing is a claim about the window, made by code that can be wrong, and a span missing from the list was never a span missing from the context.

Claude Code's /context Messages row is where the inventory stops

The second thing about the report is structural. The conversation itself, in that same output, is one line of the category table:

text
Messages

It is one row and one number. Inside it are your own turns, the model's replies, every file it read, every command output, each tool result from each server, whatever a hook injected, the path-scoped rules that loaded along the way, and, after a compaction, a summary of all of the above written by the model. The itemized part of the report covers things that sit on disk, that you could have opened anyway. The summed part is where provenance is mixed: a fetched page and a line you typed are tokens in the same row, and they should not carry the same standing.

The raw material is not gone. The session transcript on disk keeps every message, tool call and tool result. The transcript viewer shows tool activity in detail. An InstructionsLoaded hook can log each instruction file with the reason it loaded. What none of these does is read back as an inventory of the current window: this span, this many tokens, written by this party, here until that event. You can reconstruct it. Nothing hands it to you. And one span cannot be reconstructed from the window at all. A compaction summary re-authors everything it touches, so the provenance that went in as a dozen labeled results comes out as one block in the model's voice. That is the case the author-of-the-span argument was worried about: a label recorded at admission that has to survive in the window, and does not.

A context window inventory needs a lifetime column, not only a token count

The other missing column is time. A token count says how much room a span takes now. It does not say what will end it, and spans in the same session end in very different ways.

Claude Code documents this more carefully than any guide to context engineering does, in a table of what happens to each mechanism at compaction. The project-root CLAUDE.md and auto memory are re-injected from disk. Path-scoped rules and nested CLAUDE.md files loaded into message history, so they are summarized with it and reload only when a matching file is touched again. Context a hook added earlier is summarized with the rest. Invoked skill bodies return capped at 5,000 tokens each and 25,000 in total, oldest dropped first, and the walkthrough adds that the startup listing of skill descriptions is not re-injected after a compaction. Up to five of the files the agent read or edited are re-read, and one over 5,000 tokens comes back as a path with no content. The troubleshooting note for a vanished instruction reads like a lifetime column written out in prose:

text
If an instruction disappeared after compaction, it was given only in
conversation, lives in a nested CLAUDE.md that hasn't reloaded yet, or is
a path-scoped rule that hasn't matched a file since.

Three rules of identical wording and identical token cost can sit in three of those places. The first is reloaded from disk every time. The second comes back only when a matching file is opened again. The third survives a compaction only if the summary happens to keep it, and the session carries on looking exactly as it did either way. Nothing in a table of tokens tells them apart.

Lifetime also runs sideways. A fresh sub-agent, one that is not a fork of the conversation, starts with its own window. The CLAUDE.md files load again there unless the agent is one of the built-in research agents or is configured to omit them, and the main session's auto memory does not load. A rule you count as always present is present per window. That is one more reason a written brief, where you chose every line, is easier to account for than an inherited transcript.

Why a context window budget set from token counts cuts the wrong span

Budget from a token table and you cut what is large. In a Claude Code session the two large rows you can touch are the tool definitions and your own instruction file. Cutting the first is right, and I have argued it. Cutting the second, by moving detail out of CLAUDE.md into path-scoped rules and skills, is also good advice for cost. It is a trade of lifetime for tokens, and the two are rarely stated in the same sentence: the rule you moved now loads on a trigger, and after a compaction what remains of it is whatever the summary kept, until the trigger fires again.

The opposite error is quieter. A span can be cheap and still be the one that decides the run. A few hundred tokens of tool result carrying a sentence phrased as an instruction cost almost nothing on any meter. A budget ranks by size, and neither standing nor persistence correlates with size.

So the order I would argue for is inventory first, with three columns per span, and budget second. Author: who could have written this text and who can change it without you, which is not always the party that emitted it. Cost: how many tokens it occupies, and on how many requests. Lifetime: what ends it, whether a compaction, a reset, the end of a sub-agent, or nothing short of you editing a file, and it is fine for that cell to say conditional. Budgeting answers one question about a span, whether it is worth its room. The inventory lets you ask the other two of every span, including the large ones you wrote yourself: on whose word is the agent acting here, and will this still be in force an hour from now?

What a context window inventory cannot tell you

An inventory is a snapshot of something that moves. It tells you what is in the window, not what the model did with it. A span can be listed, attributed and dated and still be ignored, or be weighed against a contradicting span and lose, and no listing shows that. It is also per window, so a system of several agents has several inventories and no place where they add up.

If you build on the API directly, you are the harness, and the inventory is your own request-assembly code. That is the easier position. You can record the source and the intended lifetime of each span at the moment you add it, which is the best moment there is: later, both have to be inferred.

Context engineering is usually described as deciding what the model should see. That decision is real. It also comes second. The first is finding out what the model already sees, who put it there, and how long it stays. Claude Code answers the first of those for the files on your disk. For the conversation, which is most of a long session, the answer on offer is still a single number and a transcript you can go and read.

Discussion

No comment section here — all discussions happen on X.

Max Nardit

Max Nardit

@mnardit

More articles