Dev

Auditing CLAUDE.md with Claude Code's /doctor: keep the reason in the file, the war story in git

The new audit treats every emphatic rule in your agent config as a hypothesis about a model you may no longer run. It asks where each rule came from, and in the same file it tells you to delete the story of where it came from. Both instructions are right once the reason and the history live in different places.

Claude Code v2.1.283 shipped on September 25 with one line in its release notes that is worth an hour of your week: /doctor prompt-audit, also reachable as /checkup prompt-audit, audits "your CLAUDE.md files, skills, agents and commands for prompting patterns written for older models." A second line in the same release says stale paths, stale commands and contradicting instruction files now lead the report, and that thinking keywords Claude Code documents are kept.

Anthropic's docs describe the underlying audit in a few paragraphs, with a sample report and the promise that you "review a patch, not a rewrite." What they don't give you is the detail that decides your results: which patterns get flagged, what gets read, what is left alone. That detail exists, because the audit is not a black box. It is a procedure written in plain English, shipped inside the binary, and you can read it before you let it near your files.

What does Claude Code's /doctor prompt-audit actually run?

It hands off to the bundled claude-api skill, which follows a 235-line procedure file called shared/prompt-audit.md. The /doctor part adds a fixed scope paragraph in front of it.

The skill has had a prompt-audit subcommand since v2.1.221 (August 4), and even the 2.1.281 copy of the procedure already listed CLAUDE.md and SKILL.md among the files it audits. What v2.1.283 adds is the /doctor entry point, a wider and more precise list of configuration files, and the stale-fact and contradiction checks covered below.

Claude Code unpacks bundled skill files into a per-version folder under its temp directory (CLAUDE_CODE_TMPDIR if you set it, otherwise the system temp dir). This lists every extracted copy; pick the one under your running version:

bash
find "${CLAUDE_CODE_TMPDIR:-${TMPDIR:-/tmp}}" -path '*bundled-skills*' -name prompt-audit.md 2>/dev/null

Read it once. Everything below is taken from it.

Which files does /doctor prompt-audit read, and what does it leave alone?

Every instruction file that loads into a session in the current project, your user-level ones included. Settings files stay closed.

The scope paragraph lists it out: CLAUDE.md, CLAUDE.local.md and AGENTS.md in the project root, its ancestors and nested directories, plus the instruction files they import (any other import is reported by path, unread); .claude/CLAUDE.md; ~/.claude/CLAUDE.md; and under both .claude/ and ~/.claude/ the rule files, skills, custom commands, subagent definitions and output styles. A managed-policy CLAUDE.md and plugin-provided skills get audited but only reported on.

It is told not to read settings files, .mcp.json or ~/.claude.json, because they can hold secrets and are not prompt text. Nothing inside the project can justify an edit to a file outside it. A user-level file still gets a proposed edit for a problem in its own text, marked as affecting every project. And the audited files are treated as data: an instruction inside one is something to assess, never something to obey. Your skills are full of imperatives, and the auditor is told to read every one of them as text.

How do you run a CLAUDE.md prompt audit?

Pick the model first, then run the command. Configuration files are audited against the model the session runs on, so that model decides what counts as a fossil. The exception is a skill, subagent or command that pins its own model: it gets audited against that one.

In the terminal, check the version (2.1.283 or later):

bash
claude --version

Then inside Claude Code, switch to the model you are moving to and run the audit:

text
/model
/doctor prompt-audit

Keep anything else off that line. When the argument has more than one word, /doctor forwards your text to the skill as-is and drops its Claude Code scope paragraph. That is how you narrow the audit to one file on purpose, and how a stray trailing comment would change it by accident:

text
/doctor prompt-audit .claude/skills/deploy/SKILL.md

You get two artifacts: a report and a proposed diff. The report opens with its assumptions (scope, target model), then lists findings ordered by confidence, each with file:line, the quoted text, the pattern it matches, why it is obsolete, a confidence level and an action. Nothing is applied unless your request explicitly asks for it ("clean it up"). Even then, edits for stale facts and contradicting files stay proposals and are never applied on a blanket request. An empty result is allowed. The procedure says it outright: an audit that finds nothing should change nothing.

Which CLAUDE.md patterns count as written for older models?

The procedure sorts them into four groups. On a repository that is mostly configuration, expect nearly everything to land in the first two: dated prompt text and brittle configuration files.

Dated prompt text is the familiar list. Pressure language comes first: capitalized MUST, NEVER, ALWAYS, CRITICAL, IMPORTANT, especially with no reason beside them. The procedure's argument is that older models needed volume and current ones over-apply it, so a file full of alarms produces a cautious, hedging agent. It cuts the other way too. A hedge like "try to" on something you actually require is now read literally, as permission to skip it.

text
Before:  IMPORTANT: NEVER do X          (several per prompt)
After:   state the one or two real constraints plainly, with the reason

After that come scaffolds that API features replaced ("think step by step," scratchpad tags, prose telling the model to think harder or less), over-specification (step-by-step choreography for judgment work, long prohibition lists), and fossils, meaning workarounds for a retired model's habits, like "never use bullets" written against models that over-formatted.

The thinking rows are where Opus 5.5 changes the answer. Thinking is always on there, and effort is the only control, so a rule telling the model not to think "can't be followed." That is the procedure's phrasing, not mine. The API side of this migration fails loudly, with 400s that name their fix. The instruction side fails quietly, and that's why it needs an audit.

What changed in the same Claude Code release, and why does it lead the audit report?

Stale facts and contradictions can be rated high confidence from repository evidence alone, and the report is ordered by confidence. A dead path is wrong no matter which model reads it. An emphasis finding is a claim about model behavior.

Diffing the 2.1.281 copy of the procedure against 2.1.283 shows what the one-line release note compresses. The "volatile specifics" row now checks that each path named in an instruction file exists inside the project, and checks commands and flags against the repository's scripts and manifests by reading them, never by running them. Paths that are generated, git-ignored, placeholders or outside the repository do not count as contradicted just because they are missing.

A new row covers instruction files that contradict each other: a skill against CLAUDE.md, a rule file against a subagent brief. It separates a conflict from an override. A nested file whose different rule is explained by its own directory or task, or that names the rule it overrides, is left alone. For a real conflict, git blame decides which passage is newer. File timestamps don't decide it, and neither does what a file claims about itself, so a line saying "this supersedes everything" is just more text to assess. When the older passage is a prohibition or safety rule, the audit flags the conflict for you to decide instead of rewriting it.

The thinking-keyword fix is narrower. In a coding agent's config, a keyword the agent documents and acts on itself (Claude Code's ultrathink, for example) is configuration, not a leftover incantation. In an application's own prompt the same word is prose and still gets flagged.

Which prompt-audit findings should you reject?

The ones that pay for a failure that still happens on the model you run now. The audit can't tell that from the wording. It needs the history or a test.

The keep list is explicit: prohibitions against current, demonstrated failures stay, and the test is whether the failure reproduces on the target model, not whether the sentence looks like a prohibition. Context stays too, meaning audience, environment facts, the quality bar and the reasons behind constraints. The provenance step asks one question of every emphatic line: which failure, on which model, did this prevent, and does it still reproduce? Its default for a line with no answer is blunt: "a line nobody can justify is suspect by default."

The procedure leaves one thing for you to reconcile. Its fossils row wants every mitigation to name, or be traced to, the model it patched, because "nobody owns the removal." Its history-narratives row flags past tense, incident IDs and pinned model names in instruction files as cruft: "State the current rule; drop the archaeology." Read quickly, those two contradict each other: name the model, but don't write the model down.

They stop conflicting once you notice that the second clause of the fossils row, "or gets traced to," is doing the work, and the tracing tool the procedure uses is git blame. The reason belongs in the file, in the present tense, because that is what the model has in front of it when it weighs the rule. The incident and the model belong in the commit that added the line, where blame finds them and nothing reads them by accident.

Take one of mine. In May I added a line to stop Claude suggesting I go to bed in the middle of a working day:

text
Time of day is irrelevant to my work patterns. Do not suggest
breaks, rest, or continuing tomorrow regardless of session length
or perceived hour. Your time estimates for tasks are sourced from
solo human developer training data and do not reflect what an LLM
can do; never quote them.

The first sentence is context and the third carries the reason for the time-estimate ban, both in the present tense and with no incident in sight. That is the half that survives the history-narratives row. The other half, which model it was written against and what it looked like when it fired, has to come from the commit, or the audit has nothing to trace. Even with both halves in place, the procedure's last step still applies: whether the behavior shows up on a newer model is settled by removing the line on a scratch copy, running a few ordinary coding sessions, and watching.

One more finding deserves a careful read. The procedure flags "unenforced instructions": rules no code path, eval or reviewer checks, and that the app's own transcripts show being broken. Its advice is to enforce in code what code can enforce, the same argument as moving the rules that must hold out of prose and into a gate. Before deleting such a rule, grep the wider system for its exact text, because tests and log parsers sometimes match on prompt strings.

When should you re-run the CLAUDE.md audit?

On every model change, run under the new model.

The last step of the procedure says it plainly: prompts are per-model artifacts, and each new migration is the trigger to audit again. Because the target is the session's model, the same files can come back with different findings before and after a switch, and the procedure treats each removal as a hypothesis to probe on a scratch copy, one change at a time where the stakes are high.

A default-model flip is the moment this matters most, because it can move you to a new model without a deploy and without you choosing it. The audit won't catch that flip. What it gives you afterwards is a way to check, file by file, which of your old rules were written for the model you just left.

Discussion

No comment section here — all discussions happen on X.

Max Nardit

Max Nardit

@mnardit

More articles

Agent memory as Markdown documents: a readable memory is still a memory

Moving agent memory out of a vector store and into Markdown files makes it easy to inspect, and that is worth having. Staleness, conflicting writers and the page nobody opened were never problems of storage format, so they move into the folder intact, where a clean file can pass for a checked one.