OpenAI Founding Member’s Insights: Enhancing LLMs Output with Controlled English and Visual Formats
OpenAI Founding Member Urges Users To Make LLMs Write Like Aircraft Manuals
Andrej Karpathy says asking language models to use aerospace-grade controlled English produces cleaner output — and he's already moved on to diagrams, HTML pages, and custom AI-generated explainer videos.
The advice comes from one of AI's most influential voices, and it points to a broader shift in how humans may soon interact with large language models. As AI systems become more capable, the question is no longer just what models can do — it's how their output is best presented to the people reviewing it.
Karpathy's Case for Controlled English
On Thursday, Karpathy posted on X that he had found success asking language models to explain topics using ASD-STE100, — a controlled writing standard originally developed to make aircraft maintenance manuals easier to read. Karpathy, a founding member of OpenAI who joined Anthropic's pretraining team in May 2026, described the standard's output as "a lot more readable" due to its "heavy constraints on clean writing style."
ASD-STE100 is maintained by ASD, the organization representing European aerospace, security, and defense companies. Known formally as Simplified Technical English (STE), the standard is built around a strict set of writing rules and a dictionary of approximately 900 approved words, most carrying only a single meaning. The standard's own documentation describes STE as an international benchmark for technical writing rather than a tool for general communication.
If you're exploring how AI handles language at this level of precision, it's worth understanding what artificial intelligence is and how it processes language — because the interaction between structured writing standards and LLM behavior is more nuanced than it first appears.
Karpathy acknowledged that the standard can be demanding. He noted that he has sometimes softened his prompt to ask for writing that is "80% of the way to ASD-STE100" when the full standard felt too restrictive. Nevertheless, he said language models are already familiar with STE and can apply it effectively.
A reply under his post pointed to an open-source skill already available on GitHub that applies STE rules to text used by AI agents. The tool's documentation makes clear it is not designed for marketing copy and does not include the standard's full dictionary.
Why STE Produces Cleaner AI Output
The practical value of STE as a prompting tool lies in what the standard eliminates. By restricting vocabulary to approximately 900 approved words — each assigned a single meaning — STE removes the ambiguity that causes AI-generated text to become verbose, inconsistent, or difficult to scan. When a model is instructed to write within those constraints, the output becomes more predictable and easier to verify. That predictability matters most in professional or technical contexts where misreading a result carries real cost.
For teams already using AI in documentation workflows, applying even a partial STE constraint to prompts may reduce the editing burden significantly. The key is treating the standard not as a rigid requirement, but as a calibration tool — one that shifts the model's defaults toward precision.
Moving Beyond Text: Diagrams, HTML, and Explainer Videos
Karpathy's post did not stop at controlled writing. He laid out three additional formats for receiving model output, describing each one as "even better" than the last.
Diagrams and Visual Output
His second suggestion was to ask a model for a diagram instead of written text. He described visual output as "a lot easier to process, parse, and understand" for certain types of information. For structured topics — processes, hierarchies, comparisons — a diagram can communicate relationships that would take several paragraphs of prose to describe with equivalent clarity.
HTML Pages Over Markdown
His third recommendation was to request output "in HTML," prompting the model to generate an interactive web page rather than a block of prose. This advice echoes a point Karpathy made in May 2026, roughly a week before he joined Anthropic. At that time he shared a post by Anthropic's Thariq Shihipar explaining why members of the Claude Code team prefer HTML over Markdown. A version of that piece was later published on the Claude Blog.
The HTML format offers a practical advantage beyond presentation: interactive pages can be structured with headers, collapsible sections, and embedded visuals, making AI-generated output easier to navigate and share across teams.
Custom AI-Generated Explainer Videos
The format Karpathy described as the one he is "most bullish on" is custom explainer videos generated by AI on any topic. His example prompt asks the model for a "3b1b style" explainer — a direct reference to the widely followed 3Blue1Brown mathematics video channel — and instructs the model to use an ElevenLabs API key for narration. For users without an API key he suggested asking the model to identify free alternatives that run locally on a personal computer.
"This is actually starting to work!" he wrote.
The implications for knowledge work are significant. A custom video explainer — generated on demand, tailored to a specific topic, and narrated automatically — represents a category of output that would have been impractical to commission at scale even two years ago. Karpathy's framing is deliberate: he is describing artifacts that "would have never made sense to create before."
Why This Signals a Larger Shift in AI Workflows
The Move Toward Oversight
Karpathy's four suggestions share a common thread. Each one is a different method of changing how model output is presented to the person responsible for reviewing it. Writing, diagrams, and HTML pages are all requests made directly to the model. The video format adds an external tool — ElevenLabs or a local alternative — to handle narration.
The framing around oversight is significant. Karpathy wrote that as models handle more of the underlying work, "a lot more of our work will rise up the abstractions into oversight and understanding." That statement reflects a view held by a growing number of researchers: that the human role in AI-assisted workflows is moving away from generation and toward verification.
Understanding this shift requires some familiarity with how these systems learn and improve over time. Understanding machine learning and how models are trained provides useful context for why output format choices have a measurable effect on the quality and consistency of results.
Practical Implications for Businesses
For businesses integrating AI into content, documentation, or internal knowledge systems, the implications are practical. Clearer model output reduces the time spent interpreting results and lowers the risk of misreading AI-generated information. Formats like HTML pages and visual diagrams can also make AI output easier to share and review across teams.
The connection between output format and review efficiency is not incidental. As deep learning and machine learning capabilities continue to advance, the models themselves will become more capable of generating complex, multi-format outputs — making the human skill of evaluating that output increasingly important relative to the skill of producing it.
What Professionals Should Take From This
Karpathy closed his post by encouraging users to ask models for "large, custom, discardable software artifacts" that "would have never made sense to create before." The ASD group distributes the full ASD-STE100 standard at no cost, though its documentation states that no software tool can replace the standard itself.
For professionals following developments in AI, Karpathy's post makes one thing clear: experimenting with output format — not just prompt content — can significantly improve the usability of model responses. As AI handles more generative work, the skills that matter most may increasingly involve knowing how to evaluate and interpret output, rather than simply how to produce it.