Dev

Karpathy's tips for understanding LLM output, and the ASD-STE100 cheat sheet that gets the standard wrong

Andrej Karpathy's ladder for reading what models produce goes from controlled English to diagrams, HTML pages and explainer videos, each easier to take in than the last. The cheat sheet attached to his post is a clean, confident summary of ASD-STE100 that inverts one dictionary rule, approves a verb the standard rejects, and presents a recommendation as a dictionary entry. A clearer format can expose an error or make it easier to believe; the useful question is what each one lets you check.

On 2 October Andrej Karpathy posted a short guide to a problem most of us now have every day. It starts like this:

We'll be spending a lot more time trying to understand the outputs of language models.
A few thoughts, tips & tricks:

When I captured it on 4 October it showed 6.4 million views. What follows is a ladder of four output formats, each step up announced with "But even better:", and an image: a one-page, engineering-drawing summary of ASD-STE100, the controlled English that aircraft maintenance manuals are written in.

I read the post, then read the standard it recommends, then checked the image against the standard. The layout makes the rules easy to scan, most of it is right, and three of its dictionary rows would teach a reader the wrong thing. That turns out to be the most useful part of the whole thread, because it is the post's own subject, demonstrated by its own attachment.

What did Karpathy recommend for understanding LLM output?

Four formats, in increasing order of ambition. In his words, trimmed:

Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100,
it's a controlled language specification originally developed for aerospace maintenance
documentation.
Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot
easier to process, parse, and understand.
Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage.
Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer
videos generated on any arbitrary topic.

For the writing rung he adds a practical softening: he sometimes asks for "80% of the way to ASD-STE100" because "the spec is quite stringent." For video he gives a sample prompt, "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration", and notes that you can instead ask the model to find "decent free alternatives that use your local compute."

The summary is the actual thesis. As models do more of the work, "a lot more of our work will rise up the abstractions into oversight and understanding." Because intelligence and code are getting cheap, "you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before." He has been heading here for a while: his 2025 year in review already called code "discardable after single use" and said models should speak to us in our favored format.

I agree with the direction. The rest of this piece is about what each format lets you check.

What is ASD-STE100, and why ask an LLM to write in it?

ASD-STE100 is Simplified Technical English, a controlled language that began in 1979, when European airlines asked aircraft manufacturers for maintenance documentation that a mechanic reading in a second language could not misread. The first guide appeared in 1986. It is maintained by ASD, the European aerospace and defence association, and the current version, Issue 9 of 15 January 2025, turned it from a specification into an international standard.

It has two parts. The writing rules: 53 of them in nine sections, covering words, multi-word nouns, verbs, sentences, procedural and descriptive writing, safety instructions, punctuation and writing practices. And a dictionary of about 900 approved words, generally one meaning and one part of speech each, plus about 1,200 unapproved words with approved alternatives. Writers can also use technical nouns and technical verbs from defined categories, so STE is not confined to those 900 words. The limits most people quote are real: at most 20 words in a procedural sentence and 25 in a descriptive one, six sentences in a paragraph, three words in a multi-word noun, one instruction per sentence unless actions happen at the same time, active voice, no semicolons.

Why it suits a model: the rules remove exactly the habits that make LLM prose tiring. No "it is imperative that," no "prior to commencing," no stacked clauses, and a strong preference for saying who does what. Karpathy's point is that models know the style well and that the result is "a lot more readable." And for human readers, the evidence that it helps is real: a 1996 study of 175 aircraft maintenance technicians found comprehension improved with Simplified English, most for the hardest work cards and for non-native readers.

Before you use it, know what you are borrowing:

  • It is free to obtain and restricted to redistribute. You request a copy through the official downloads page; the document is ASD's copyright, with reproduction limited to the permissions in its notice.
  • Its own FAQ says it "is not intended for general-purpose writing," while allowing that its principles, such as short sentences and the active voice, travel well. Karpathy's "80% of the way" is that same compromise.
  • It is hard to write well: "STE was created for the maximum benefit of the reader. This does not necessarily mean that it is simple to write."

Karpathy did not start this. A wave of "make the model write in STE" skills hit GitHub and Hacker News in July 2026, two of them with more than 3,000 stars each. His post put a much larger audience on it.

Does the ASD-STE100 cheat sheet in Karpathy's post get the standard right?

Many details are right, and several entries would teach the wrong rule. Here is the image attached to the post:

The 'Simplified Technical English: overview' cheat sheet attached to Karpathy's post

Image from Karpathy's post on X, 2 October 2026, reproduced for review; its creator is not stated. It is not an ASD publication.

I checked it panel by panel against Issue 9, and where useful against Issue 8. A lot of it holds: every numeric limit (20, 25, 6, 3, one instruction per sentence), the six approved verb forms, the two-part, nine-section structure, the uppercase convention for approved words, and the example sentences with their word counts. Seven of the ten dictionary rows have the right word and status. A tired reader would trust it, and mostly should.

Now the dictionary panel, where it goes wrong:

Detail of the cheat sheet's dictionary panel, including the approximately and TEST rows

Detail of the same image.

  • approximately is approved, and the sheet says it isn't. The row marks "approximately (adv)" as not approved and offers ABOUT instead, with "Wait for ABOUT 10 minutes" as the correct form. In the standard it is the other way round. APPROXIMATELY is an approved adverb meaning "almost correct or accurate." ABOUT is approved only as a preposition meaning "concerned with," and the standard uses this exact pair as a teaching example: "Drain about 2 liters of fuel from the tank" is its example of what not to write. The sheet teaches the error the standard warns against.
  • TEST is not an approved verb. The row marks "TEST (v)" approved, with a definition, "To find if it operates correctly," that I could not find in the standard. TEST is approved only as a noun; the standard's own example is "DO A FUNCTIONAL TEST OF THE SOFTWARE," with "Functionally test the software" as the version to avoid. Someone quoted this exact rule on Hacker News in July.
  • "in order to" has no entry I could find. Shortening it to TO is sensible advice and in the spirit of the standard, but I found no such row in the dictionaries of Issue 9 or Issue 8. The sheet presents a recommendation in the format of a dictionary entry.

Outside the dictionary there are two more substantive problems. The verb panel says descriptive text "can use the passive only when it is necessary." The rule is narrower: in descriptive writing "you can use the passive voice only when the agent is unknown." "When necessary" is a judgement call; "when the agent is unknown" is a test. And the structure panel lists technical names and technical verbs as dictionary content, while the FAQ says plainly that the dictionary does not list them; they are categories defined in the writing rules.

There is also a quieter signal across the sheet: its terminology predates Issue 9 and matches Issue 8. It says "noun clusters" (now "multi-word nouns"), "technical name" (now "technical noun"), "Specification" in the title block (now a standard), and its history timeline stops before the 2025 change. One section title, "Procedures," matches neither edition.

The post doesn't say who made the image or how, and whatever its origin, the result is the same: useful summaries next to incorrect dictionary advice, in a layout that makes both look equally authoritative. It is already being used as a source. One repository created the day of the post takes its limits from "Karpathy's sheet" (the limits happen to be the correct part). And a separate STE skill written in July, before the post, has an open issue listing twelve rows that ban approved words, approximately among them, so machine-built STE word tables fail the same way independently. Readers did catch part of this: about 20 hours after the post, a reply flagged the approximately row by pasting a ChatGPT answer, and a GitHub reference file documented the same error and the older terms. The TEST row, the "in order to" row and the passive-voice paraphrase I found no one flagging.

The standard's own maintainers saw this coming. In June 2026 they published a white paper on STE and AI, and the downloads page summarizes it in two sentences I would put above every AI style guide:

AI-generated text can appear clear, authoritative, and consistent with STE, even when it
does not correctly apply the rules and vocabulary of the standard. Plausibility must not
be confused with verified compliance.

Do LLMs actually follow ASD-STE100 when you ask?

Loosely, and they tend to believe they follow it better than they do.

A technical writer who looked closely calls what models produce "STE-flavored English, not STE": without checking against the current dictionary, a model can imitate the style while breaking its vocabulary rules. In the replies under Karpathy's post, one person reports that three current models "simply ignore" the request; others saw little difference between the STE answer and the normal one. On Hacker News in July, a recurring view was that compliance drifts quickly and only a linter or a commit hook holds it.

The measurements are thin. The most-starred STE skill advertised that it "cuts slop 72.9%"; that number counted violations of the skill's own linter rules, and in a later eight-scenario audit of version 2.0.0 its author found that a short prompt produced shorter replies with fewer formatting defects. That was not a comprehension study, and the project has since rebuilt the skill. Lucian Ghinda ran a small, explicitly informal test on code explanations: with Claude, a loose "simple technical English" prompt mentioned 8.5% fewer of the scored facts than the baseline answer, and a strict ASD-STE100 prompt 46.8% fewer; with Codex the corresponding losses were 43.3% and 40.0%. Four code samples, two sessions. Treat it as a warning that simplification can drop facts, not as a ranking. Two 2026 preprints on constraint following find what you would expect: models overestimate their own compliance, and they can restate a rule while breaking it.

The practical conclusion is the one I made about prompts in general: a style rule in a prompt is a request, not a guarantee. If you actually need STE, check it outside the model. The prose linter Vale can encode the rules you choose to enforce, which also replaces "80%" with a list you decided on. A linter checks the rules you configured; it doesn't certify full STE compliance, or that the text is true.

Are diagrams better than text for understanding LLM output?

Larkin and Simon put the answer in the title of their 1987 paper: "Why a Diagram is (Sometimes) Worth Ten Thousand Words." A diagram holding the same information as text can be far cheaper to use, because location does the indexing that a reader would otherwise do in their head. A 2025 meta-analysis by Cromley and Chen, across 181 studies of multimedia learning, found an overall effect around 0.37, with large variation by design principle and outcome; text-and-diagram results were among the more consistent ones, animation much less so.

Two cautions specific to models. Generated diagrams are still unreliable as diagrams: benchmarks on SVG, Mermaid and similar formats show real gaps. And a diagram has its own way of overstating. One reply put it well: when you simplify, protect words like "if," "unless" and "not yet verified," because without them "a clearer sentence can become a stronger claim," and the same goes for a diagram, where a missing arrow quietly asserts that there is no dependency. The counter-case also came up: someone caught a real error because one arrow pointed the wrong way, which the prose had hidden. Both are true. A diagram makes structure visible, including wrong structure, and makes omissions invisible.

How I would use it: ask for a diagram in a text format you can read and diff (Mermaid, Graphviz, a simple SVG), look at the rendered result, and ask the model to list every relationship it drew as a sentence underneath. Then read the list.

Why ask an LLM for output in HTML?

Bret Victor described the goal in 2011, in "Explorable Explanations": "A reactive document allows the reader to play with the author's assumptions and analyses, and see the consequences." That is what an HTML answer offers that prose doesn't: tables and sliders, collapsible sections, small simulations. It was the rung the replies liked best, and Claude's artifacts and ChatGPT's canvas brought generated content into dedicated workspaces in 2024.

The best suggestion in the thread, made independently by several people, was a variation: don't ask for a page that shows the answer, ask for a page that lets you change one assumption and watch what breaks. That makes the explanation's assumptions inspectable. It is not yet a test of the real thing. If a model explains a caching bug with an interactive page, the page's toggle verifies the page's model of the bug, not your system; to check the diagnosis, reproduce it with the actual code and inputs.

The costs: an HTML artifact takes more tokens and time than an answer, and it can be polished and wrong in the same way the cheat sheet is. Ask for a self-contained file with its assumptions listed visibly on the page.

Can an LLM make a 3Blue1Brown-style explainer video?

Increasingly, and this is where Karpathy is most excited and the evidence is thinnest. One way to get the "3b1b style" is Manim, the animation engine Grant Sanderson wrote for 3Blue1Brown (MIT licensed, about 94,500 stars), or ManimCE, the better-documented community fork; their setup differs, so say which one you want. A model writes Manim code, the code renders the animation, a text-to-speech service narrates it. In the replies, people did exactly that within a day: on stochastic calculus, on how payments work, one entirely with local tools and no API keys, which is the alternative Karpathy himself mentions.

There is real research behind the idea. TheoremExplainAgent (ACL 2025), evaluated on a benchmark of 240 theorems, found something encouraging: the videos exposed flaws in the model's reasoning that its text explanations had hidden. Animation forces you to commit to what happens next. Whether that improves what viewers learn is a separate question it didn't test.

Three practical notes. ElevenLabs, which Karpathy's sample prompt names, logs the text you send by default, and turning that off is an enterprise feature, so keep private material out of hosted narration or use the local route. Put the key in the tool's credential settings rather than in a prompt. And video is harder to scan without a transcript or chapters: in a study of 6.9 million viewing sessions across 862 edX videos, median engagement time topped out at about six minutes. Keep the script, the captions and the source project, so the claims stay searchable. One early reply put the risk in a sentence: "a wrong claim narrated over a 3b1b animation is way harder to catch than a wrong sentence."

Does easier-to-read LLM output mean easier-to-verify output?

Not by itself. A clearer format can expose an error, as the arrow and the theorem videos did, and it can make an unsupported claim easier to accept. The useful question is what each format lets you inspect.

The research on the second effect is worth knowing. In a 2024 study (Si et al., NAACL), about 80 crowdworkers fact-checking claims with an LLM's explanation were right 87% of the time when the explanation was correct and 35% when it was wrong, below the 49% they managed with no evidence at all on those cases. A CHI 2021 study found explanations increased people's acceptance of AI advice whether or not the advice was right. Work on text simplification (Devaraj et al., ACL 2022) found factual errors common in simplified versions and argued that a readable but inaccurate version can be worse than no access at all. And the illusion of explanatory depth, our habit of overrating how well we understand mechanisms, makes me cautious about mistaking a visible mechanism for a complete explanation.

None of those studies compared Karpathy's four formats, so I won't claim each rung is harder to check than the last. What I'd say is narrower: each rung moves the claims somewhere new. Prose you can quote and search. A diagram puts them in arrows and absences. A page puts them in code and layout. A video puts them in timing and voice. The checking has to follow them there, and none of the four formats does it for you.

Karpathy's summary says our work will rise "into oversight and understanding." I think that is right, and I'd add that the two aren't the same activity. Understanding a translation is not the same as verifying it, and the same is true of any output you couldn't have produced yourself.

How should you use Karpathy's techniques without being fooled?

Use all four rungs; they are good. Keep one thing on each rung that you can check.

For writing in STE or anything like it:

  • Ask for the simplified version and keep the original next to it. Simplification should be a view, not a replacement.
  • Tell the model to keep identifiers, parameter names, error strings and technical terms exactly as written, and to keep every "if," "unless" and "not verified." Those are the words a simplifier drops first.
  • If the style matters, check it with a linter, not with the model's word.

For diagrams: request a text format, look at the render, then read a sentence per edge.

For HTML: ask for the page that lets you change an assumption, with its assumptions listed, and test the real system separately.

For video: keep the script as text, read it once before you watch, and keep private material out of hosted narration.

And for anything you will act on: pick one load-bearing claim and check it against the primary source. That is how the cheat sheet comes apart. It took one look at one dictionary entry.

What does Karpathy's post say about where work is going?

That the scarce skill is moving from producing output to judging it, and that models can help with the judging too, by producing artifacts that would never have been worth making by hand. The discardable explainer, the one-off interactive page, the video made for an audience of one: these are real and new.

What I would add is a distinction the post leaves implicit. Feeling that you understand something is not evidence that it is correct. The ladder is very good at the first. The second has to be built in on purpose, in whatever format you choose. The most widely shared summary of a standard for clear, unambiguous writing this week inverted one of that standard's own teaching examples, in the clearest possible layout. I would keep the primary source open next to the beautiful thing.

Discussion

No comment section here — all discussions happen on X.

Max Nardit

Max Nardit

@mnardit

More articles

Notion's ChatGPT token sharing: your AI bill now has two meters

Notion now lets a ChatGPT Plus or Pro subscription pay for some of its AI. It covers one agent and one model family, it still needs a Notion Business or Enterprise plan, and its eligibility changed by tweet within two days of launch. What token sharing is, how to set it up safely, what it covers, and why I still can't tell how much of either allowance a Notion task will use.