Dev

Claude Code skills as a supply chain: 2.2 million adoptions, no registry, and a scanner that reads the source while Python runs the cache

Reading a skill before you run it is the standard advice, and it has two holes. A copied skill has no update channel, so the copy you read can drift from its source without telling you, and for a skill that ships Python, the file you read is not always the file the interpreter loads.

Two papers reached arXiv within a day of each other this week, and they describe the two ends of the same pipe. Skill Constellations looks at how SKILL.md files travel between repositories on GitHub. PyCache Trap looks at what happens when one of those files arrives with Python next to it and a scanner is asked whether it is safe.

The standard advice is to read a skill before you run it. Between them, the two papers are the two holes in that advice. What you read is a copy, and nothing tells you when its source changes. And when the skill ships Python, the interpreter may load a different file from the one you opened.

Neither paper is about Claude Code alone. The skill format is open and Codex, Cursor and others load it. But Anthropic introduced the format, Claude Code is where I use it, and the questions below are the ones I asked about my own skills folder. Figures from the papers are reported, not reproduced. The Python behavior in the middle of this post I did check, on my own machine.

How do Claude Code skills spread on GitHub without a registry?

By being copied, mostly in bulk. Skill Constellations works from the GitSkills dataset: every file named exactly SKILL.md with valid front matter, first committed on or after 1 October 2025 in a non-fork repository, followed through its git history. It counts 2,193,119 adoptions, where an adoption is one repository acquiring one skill, the original author's first commit included. That is a count of repository and skill pairs. It says nothing about how many people run them.

The interesting unit is the copy event: a commit that adds ten or more skills at once, where a single earlier repository already held at least half of the skill lineages the committing repository has. So a skill typically arrives as part of a whole folder that traces to one earlier holder, and not as one developer picking one tool.

The distribution is what you would expect from that. Most adoptions pass the skill to nobody, and a small number of repositories account for nearly all the transmissions.

Do GitHub stars show which skill repositories are worth auditing?

No, and by a wide margin. Stars are reported as nearly uninformative about how many repositories copy from a source. Ranking by how often a repository has already been copied identifies 4.0 times as many later transmissions as ranking by stars.

Counting copies naively misleads too. Of the 20 repositories with the highest copy counts, only 7 stay in the top 20 once each is credited only for skills it held first. The rest are late adopters that look like sources.

The headline number comes from a replay. Fit a model of which earlier holder a copy event picks, on data before 1 April 2026, then audit the top 100 repositories by that model, where auditing means removing their flagged skills. That prevents 14.9% of later high-risk adoptions. Auditing the 100 most starred prevents 0.5%. The ceiling is low either way: at the same 100 audits, the paper says even a ranking chosen with hindsight prevents only about a quarter, because roughly half of the high-risk adoptions arrive as new skills that no audit of existing sources would have seen.

Do copied Claude Code skills pick up fixes from their source?

Mostly not. The paper reports that 20.9% of skill copies follow a later edit of their source. For comparison it cites earlier work on copied code files, which adopt upstream security fixes in 47% to 84% of cases. Inside groups of identical copies, 11.1% of changes are consistent across all members.

The reason is in the mechanism. A skill is a folder. Once it is copied into your repository or your home directory, nothing obliges it to remember where it came from. Some installers record a source, and the paper uses those records to validate its dates, but a plain copy carries no version, no lockfile entry and no origin URL. A fix upstream reaches you only if you go and look.

Both kinds sit on my own machine. Plugins from a marketplace carry a source and a version, and updates arrive through it. The folders copied into ~/.claude/skills by hand have neither, and nothing tells me whether their source has moved since.

What does "high-risk" mean for a skill in Skill Constellations?

Capability, not malice. The paper flags skills that bundle executables, pre-approve tools or instruct risky actions, and says outright that the flag marks what a skill can do, not what its author intends. On a labelled sample of 200 skills the strict flag has 98.3% precision and 80.9% recall, so the risk counts are lower bounds.

In Claude Code the second item has a name. The allowed-tools field in a skill's front matter lists, in the docs' words:

Tools Claude can use without asking permission during the turn that invokes this skill.

With allowed-tools: Bash(python3 *) in the front matter, Claude can run a script that sits next to the skill without asking each time, unless a managed policy switches the grant off. That is the designed behavior, and it is useful. It also means a bundled Python script can run with nobody looking at the command, which is the setting PyCache Trap studies.

How does PyCache Trap get a skill past a scanner?

It gives the scanner one program and the interpreter another. The package contains a harmless helper.py and, in __pycache__, a compiled file for the same module with different behavior. Scanners read the source. Python, when the module is imported, loads the cache if its header passes the loader's checks.

The paper tests seven skill scanners on 100 skills. These are adaptive results: the attacker gets up to five candidates per skill and sees the scanner's feedback. Under those conditions the attack gets through 94% to 100% of the time depending on the scanner, and none of the seven recognized what the cache did. For four of the scanners the authors also ran the admitted skill in a Claude Code harness driven by MiniMax-M3: in 76 to 83 of 100 cases the skill passed the scanner and the hidden behavior ran. The payloads were synthetic, the secrets were dummies, and nothing was published to a real platform.

The paper puts the reason checking the header is not enough in one sentence:

Header consistency alone cannot establish derivation because the publisher controls
both header and body.

Does a substituted Python cache survive git clone?

It depends on which kind of cache it is, and the paper does not test delivery at all. So I checked, on Python 3.12.3 and git 2.43 on Linux, with a module that returns one string from source and another from the planted cache. What follows is what I observed on that setup.

Python has three ways to validate a .pyc, set in its header (PEP 552). The default stores the source file's modification time and size. The two hash-based modes store a hash of the source instead: one checks it on every import, the other trusts it unless the interpreter is told to check.

Planted cachegit clonezip, tar--check-hash-based-pycs always
Timestamp (default)source runscache runsnot applicable
Hash, unchecked, wrong hashcache runscache runssource runs
Hash, either mode, real source hashcache runscache runscache runs

The timestamp kind did not survive the clone: git does not preserve modification times, so the header no longer matched (a clone that lands in the same second as the original would still match). It did survive zip and unzip, a tarball and cp -a, which restore the times. A tool that does not restore them breaks it the same way the clone does.

The hash kinds do not care about file times at all. Look at the last row. The publisher computes the true hash of the clean source, writes it into the header, and attaches any body. Python hashes the source, finds a match, and loads the body. The strictest interpreter flag does not help, which is the paper's point about header and body reproduced in a few lines. All of this assumes the same Python minor version, since a cache built for another one is ignored.

Two limits on this. A cache only arrives by git if the publisher commits it; most Python projects ignore __pycache__, but a skills repository with no .gitignore adds it on a plain git add -A, as mine did in the test. And only imported modules load from the cache. A script the skill runs directly is always compiled from its source.

Does python -B or deleting __pycache__ protect a Claude Code skill?

-B does not, deleting does. python -B and PYTHONDONTWRITEBYTECODE stop Python from writing caches. They do not stop it from reading one that is already there, and in my test the planted cache ran under -B in every case where it ran without it.

Three things worked in my test. Deleting the __pycache__ directories, after which Python compiles from the source you can read. Setting PYTHONPYCACHEPREFIX to a directory of your own, so Python does not look next to the source at all. And running python -m compileall -f over the folder, which is what the paper's repair step amounts to (it removed the substituted bytes in all 99 code-bearing packages they tried). The last one is the weakest of the three: it rewrites the cache only for the interpreter and optimization level you run it with, and leaves any other .pyc in the directory alone. Delete first.

What would versioned references change for Claude Code skills?

They would give a copy somewhere to look. Skill Constellations ends on that recommendation: platforms should distribute versioned references rather than copies. A reference has a source you can pin or revoke. A copied folder has nothing on the other end.

Plugins are the nearest thing Claude Code has today. A skill installed through a plugin has a source and a version, and an update mechanism: automatic for some marketplaces, a setting you turn on for others. That solves the frozen copy and opens the other question, the one I raised about the plugin directory: a listing is a statement about a review, and reviews happen at fixed moments. The directory scans every new bundle version; a plugin added from a marketplace URL gets no review at all. Either way an update channel delivers fixes and whatever else the publisher ships next.

The paper's own ceiling belongs here too. At the same budget of 100 repositories, even a ranking chosen with hindsight prevents about a quarter of later high-risk adoptions. Ranking and registries narrow the problem. What a skill may do once it is on the machine is still decided there, the same way it is for any tool you install.

How do you audit a Claude Code skills folder this week?

Four commands, and the first three are about caches. They assume personal skills in ~/.claude/skills and plugins in ~/.claude/plugins; add a project's .claude/skills if you have cloned one.

bash
# 1. Bytecode caches sitting next to skill code
find ~/.claude/skills ~/.claude/plugins -name '__pycache__' -type d
 
# 2. Hash-based caches: the kind that survives a clone
find ~/.claude/skills ~/.claude/plugins -name '*.pyc' \
  -exec sh -c 'xxd -s 4 -l 4 -p "$1" | grep -qE "^0[13]000000" && echo "$1"' _ {} \;
 
# 3. Remove them all; Python recompiles from the source you can read
find ~/.claude/skills ~/.claude/plugins -name '__pycache__' -type d -prune -exec rm -rf {} +
 
# 4. Skills that pre-approve tools
grep -rlE '^allowed-tools:' --include=SKILL.md ~/.claude/skills ~/.claude/plugins

On my machine the first command finds one directory, under a skill I wrote, with a timestamp header, so my own interpreter made it, and the second finds nothing. No plugin shipped a cache. That is the expected result. Hash-based caches have honest uses, reproducible builds among them, so finding one proves nothing by itself. But a cache of any kind shipped inside a skill you did not write is worth a look, and deleting it costs nothing: Python rebuilds it from the source.

Read the fourth list slowly. Nothing is wrong with a skill that pre-approves a shell pattern and ships a script. But it is where reading the SKILL.md tells you least, because the prose is not what runs. Same lesson as a cloned repo's settings file: the file you skim and the thing that executes are two objects, and the approval covers the second.

Discussion

No comment section here — all discussions happen on X.

Max Nardit

Max Nardit

@mnardit

More articles