Three records of one agent incident
OpenAI's logs, a URL scanner's public archive and a statistics bureau each tell it differently. Whoever keeps the fullest account is also the one grading it.
Data & AI Systems Engineer working on visibility, measurement, and agentic systems.
AI now sits between people and what they’re looking for, and between work and the software that runs it. It answers, routes, remembers, forgets, and now acts on its own. Usually before anyone can check whether it got things right. I study that layer with original data and build the tools to work inside it. My background is the supply side of digital marketing: the crawl, the analytics, the automation and reporting, the plumbing under what looked like marketing. The marketing was never really the problem. Visibility, measurement, and control were. AI didn’t make that problem smaller. It moved it somewhere you can’t see.
Beetroot featured in 窓の杜 (Japan) · original-data research · shipped open-source tools
the systems that decide
what gets found, trusted,
and acted on
Thesis
AI is changing two things at once: how people find information, and how work moves through software.
That shift isn’t only a search problem. It touches tracking, attribution, context, memory, handoff, and trust: the systems that decide what a person sees before they make a choice.
I work on that layer. Some of it is research: measuring what changes and publishing what holds up. Some of it is engineering: building tools and workflows that keep context, expose failures, and make AI-assisted work inspectable.
Focus
Writing
OpenAI's logs, a URL scanner's public archive and a statistics bureau each tell it differently. Whoever keeps the fullest account is also the one grading it.
One panel gives you a snapshot. A report compares two periods, and that is where most AI visibility changes turn out to be artifacts of the panel, the sample size or an unlogged model swap. The design that makes a trend checkable.
Switch to a model or an API account that can't read the earlier thinking, and the request still succeeds. The reasoning those turns built up is dropped before the model reads it, with no error and nothing on the bill. You only find out if you asked the API to report it.
Once an agent will push any number you give it, the model stops being the part worth arguing about. The part you own is the evidence that the number and the goal move together, and that evidence expires as soon as the optimizer leaves the region where you checked it.
Mention, answer inclusion and citation come apart the moment you sample an assistant more than once. One prompt can carry a brand's entire presence, a parser can move its citation rate 2.5x, and a hallucinated license still counts as a mention. A workflow for technical teams, built on 359 answers.
Cache diagnostics names the first place a request stopped matching the one before it. The useful part is reading that verdict next to the cache read count, and covering the parameter changes it cannot see yourself, because those are the misses hardest to find by eye.
Artifacts
Tools I’ve built around the same layer: memory, clipboard, agent handoff, monitoring. Local-first where it makes sense, from problems I hit directly.
A local-first clipboard environment for preserving workflow context on Windows. AI transforms, OCR, smart search, unlimited history. Built around the idea that the clipboard is not temporary — it is part of workflow memory.
Learn more →
A clipboard bridge between humans and AI agents. Read, write, and watch clipboard state through MCP. A small interaction primitive for agent-native workflows.
Learn more →
A browser-to-agent handoff surface. Send pages into agent workflows through a webhook, preserving context at the moment of browsing.
Learn more →
Scoped persistent memory for long-running coding agents. Local-first, MCP-native, designed so context survives across sessions without piling up.
Learn more →
A small operational monitor for Claude.ai limits.
Learn more →