A benchmark has to prove it tracks the clock
Once an agent will push any number you give it, the model stops being the part worth arguing about. The part you own is the evidence that the number and the goal move together, and that evidence expires as soon as the optimizer leaves the region where you checked it.
Anthropic published an account this week of how a small team made claude.ai and the desktop app about three times faster in two weeks. The numbers are a bit absurd: time to a typeable page on a fresh load went from 3.1 seconds to 0.55 at the 75th percentile, and more than three thousand changes merged without a customer-facing incident or rollback. Twelve of the thirteen targets fell by day three, mostly to the planned projects, a static composer baked into the HTML and a precompiled V8 code cache among them. After that, Claude did most of the climbing, across more than a hundred and fifty parallel threads. The line the authors put at the center of the post is this one: "With Claude, measuring something makes it tractable."
The post is also unusually honest about the catch. Before any new benchmark was allowed to stand, someone posted this to the channel:
please prove that hill climbing against each of these can result in
measurable wall clock perf wins. we’ll unship the benches for any
candidates that cannot prove thatand the proof sits under a chart titled "Does the count track the clock?" So the idea that you validate a proxy before you let an agent climb it is theirs, stated plainly. What the post does not say is how long that answer stays true once the agent keeps climbing, and that is the part I want to argue about.
Why the lab needed a different number
The team wanted to iterate faster than they could deploy. Claude could work for hours, overnight, and waiting on field reads for every prototype would have throttled it. So they needed lab measurements, and wall-clock time in a lab is noisy; in the post's words, "milliseconds are too flaky to use as a CI gate." The targets themselves stayed wall-clock, real-user p75 per journey. The lab numbers were stand-ins for them.
Sam, one of the engineers, asked whether they could count JavaScript instructions instead. Claude answered with a method: run the benchmark under Valgrind with node --predictable, one run, no statistics. For browser paths, where Chromium offers no instruction counting, it listed other deterministic counts: React commits per interaction, function calls from V8's precise coverage, layout and style recalculations, DOM mutations. Eleven minutes later five threads were running, one per measurement. The rule the team set for all of them:
We treated every new benchmark with some skepticism. Each one had two jobs:
first, a metric Claude could move in the lab; second, a guardrail in CI with
a number that could only ratchet down. If a benchmark was flaky, or if it
didn’t actually correlate with user latency, we threw it out rather than let
Claude climb the wrong hill.What Goodhart actually said
Charles Goodhart's 1975 formulation was about monetary policy, and it is sharper than the paraphrase that circulates:
any observed statistical regularity will tend to collapse once pressure is
placed upon it for control purposes.The key word is "observed". The regularity was seen without pressure on it, and pressure changes the regime. Nobody has to cheat for this to happen. An optimizer only has to keep finding directions in which the number moves, and an agent that opens fifty or a hundred PRs against one benchmark is about as much pressure as a number will ever see.
Manheim and Garrabrant's taxonomy of Goodhart effects has a name for the nearest version of this problem: extremal Goodhart. A relationship between proxy and goal holds in the region where it was observed, and selection pushes the system into regions where it may not. They split it further, into a relationship that was only approximately right to begin with and one that is right locally but different elsewhere. From outside you cannot tell which one applies here, and neither can the count. The two measured paths show the relationship in the region the agent started from, and nowhere else.
I made a related argument about search ranking months ago: a ranker that cannot measure quality picks proxies that correlate with it on average, and the marginal case gets lost in the average. The agent case is the same mechanism with one difference that matters: here the target can be read.
Two points, measured two ways
Here is the validation the post shows. Claude drove the instruction count down on two hot paths, the routine that assembles a conversation's message tree and a scanner for status lines in Claude Code output. Instructions dropped 48% and 31%. Wall-clock time dropped 78% and 44%. Counts came from Valgrind with node --predictable, timings from the same benchmark under plain node with the JIT warm.
That is good evidence that the proxy points the right way for those changes. It is also visibly not proportional. Removing about half the instructions removed more than three quarters of the time, which means the instructions that went away were far more expensive than the average one. The post says why: a quarter of the first path's instructions were megamorphic dictionary lookups, resolving the same message ID three times. An instruction count weighs a slow lookup and a cheap add the same. It happened to move in the right direction here because the first fix removed expensive work. My read is that nothing guarantees the next fix on the same path has the same mix, and the count cannot tell you when it stops.
What the ratchet keeps enforcing
After the proof, the team checked in two ratchets: any PR that raised the instruction count on those paths failed CI, "and a daily job lowered each ceiling whenever the count went down."
The first half is a regression guard and a good one. The second half is where the relationship stops being checked. The ceiling moves on the count alone. No field read is attached to that daily job, and the paired timing that validated the proxy is not rerun when it fires. So every lowering quietly reasserts the original claim, fewer instructions on this path means less latency, at a point further from where anyone measured it.
Take a trade the ratchet cannot see. Memoizing a result adds instructions on the benchmarked cold call and saves them on every later one. Whether that is a win depends on how often the path runs warm in real use, which the count cannot know. The ceiling is a checked-in baseline, so a person could raise it in the same PR with a sentence of justification. My guess is that an agent scoped to one benchmark rarely tries. It was told to drive the number down, and to that agent the ceiling reads as a wall, so the trade never gets proposed.
The post shows the opposite correction happening too. One 900-line PR got a one-line reply: "going to gavel that 2ms per send is not worth the complexity of maintaining this build plugin." A person overruled a metric improvement on grounds the metric could not see. The ratchet guards the number against getting worse. Deciding whether an improvement is worth having stayed with people, and so did deciding whether a regression was worth it.
I have argued before that a load-bearing rule belongs in a check the harness runs and enforces rather than in a sentence the model weighs. The ratchet is that move done right. But an enforced check only holds as well as the claim it enforces, and "instruction count tracks latency on this path" is an empirical claim with a shelf life. That earlier argument ended on the same limit: some violations show up in the world, not in the action a gate inspects.
Every instrument in the chain is a proxy
The loop the post describes does check the lab against the field. After a change shipped, Claude watched the deploy and read the field data; if performance improved, it ratcheted the benchmark down, and if not, it turned the flag off and iterated. Anything that could cause a user-visible problem went behind a short-lived flag, high-risk changes went to employees first, then one percent of users, then everyone, and every PR needed at least one human approval. The loop has a branch for a lab win that does not show up in the field. The post does not say how often that branch fired.
The tempting conclusion is that the field read is the real guard, but the post's own best finds cut against it. The sidebar jank that started one thread was invisible to every monitor they had; Cumulative Layout Shift scored each shift around 0.008, well inside the "good" threshold of 0.1. A leftover location.reload() was causing half a million hidden reloads a day "that none of our load metrics could see." The layout shift after the static composer shipped internally was reported by a teammate with a screen recording, not an instrument. In each case the field metrics were blind too, and a person watching a screen was not.
So real-user p75 is also a proxy, one step closer to the target than an instruction count and still not the target. The chain runs instruction count, lab timing, field percentile, and ends at a person looking at the product. The team's response each time was to build another instrument, which is the post's thesis working as intended. Mine is narrower. Each link is validated against the next one up, and only in the region someone checked, so a validation is only as current as the last time someone checked that link against the one above it.
Before the first run
Write down the target the number stands for, and how you will read it, before you pick the number. If you cannot assess the target independently, you can still use the proxy, but you will not see it fail.
Record the validation as a claim with a domain: which paths, which kind of change, measured how. The first paired runs tell you the proxy tracks the target for the changes that produced them.
Attach the recheck to the ratchet itself. The paired timing already exists as a benchmark under plain node. Rerun it whenever the daily job lowers a ceiling, log the ratio of instruction change to time change, and hold the lowering when that ratio drifts from the one you validated. It costs one extra benchmark run per ceiling move. Cheap, next to a ceiling that has been enforcing a stale claim for months.
The Anthropic team did most of this by hand, with a named owner on every thread. Climbing got cheap. Checking that the hill still leads where you meant to go did not.