On This Page8 sections
Key Takeaways
- Answer me with HTML is an Agent Skill that converts complex AI answers into readable, self-contained HTML pages. Instead of making the model generate an entire webpage, it lets the model produce compact structured content while a local renderer handles layout, diagrams, CSS, themes, and SVG positioning.
- Its core idea is simple: let the LLM decide what to explain, and let deterministic code decide how to display it.
- In the project's published benchmark using Claude Sonnet 5.5, average output fell from 6,873 tokens to 923 tokens, while completion time fell from 46 seconds to 13 seconds.
- The same benchmark reported roughly $0.22 for direct HTML versus $0.26 with the skill, meaning fewer output tokens do not automatically translate into lower total API cost.
- The project supports Agent Skill-compatible environments including Claude Code, Codex, Cursor, and OpenCode.
- It can render flow diagrams, sequence diagrams, trees, timelines, comparison tables, annotations, callouts, and narrated explainer videos.
What Is Answer me with HTML?
Answer me with HTML is an open-source Agent Skill and local CLI designed to turn substantial AI explanations into visual HTML artifacts instead of long walls of Markdown or terminal text.
The project was created by QingYunA and follows a straightforward principle: complex answers are often easier to understand when the presentation format matches the structure of the information.
Instead of telling a coding model to generate hundreds of lines of HTML, CSS, SVG, and JavaScript, the skill asks the model to produce a much smaller structured draft. The bundled am renderer then converts that draft into the final page.
The workflow can be simplified as:
User question
↓
AI agent
↓
Compact structured Markdown
↓
am CLI
↓
Layout engine + components + SVG
↓
Single HTML fileThis makes Answer me with HTML more than an HTML prompt template. It acts as a reusable presentation layer for AI agents.
Why Is Answer me with HTML Getting Attention?
The project arrived during a broader discussion about improving how humans consume AI-generated information.
In early October 2026, Andrej Karpathy highlighted a progression for AI explanations that moves beyond ordinary prose toward controlled writing, diagrams, interactive HTML, and eventually generated explainer videos.
Answer me with HTML takes a similar direction but introduces an important optimization: the language model does not manually build the final webpage.
Instead, the AI handles semantic decisions while deterministic software performs repetitive visual work.
That distinction matters because generating HTML directly can cause models to spend thousands of output tokens on:
- wrapper elements;
- utility classes;
- CSS declarations;
- SVG coordinates;
- card layouts;
- responsive rules;
- dark-mode styles;
- diagram arrows and labels.
Those tokens make the result prettier, but they often contribute very little additional reasoning.
Answer me with HTML tries to move that deterministic work out of the model and into reusable software.
How Answer me with HTML Works
The model generates an extended Markdown-style representation instead of a complete webpage.
For example, a TCP explanation can be represented with something conceptually similar to:
---
title: TCP three-way handshake
---
## Three-way handshake
sequence:
Client -> Server: SYN
Server -> Client: SYN+ACK
Client -> Server: ACKThe renderer then decides how to draw the sequence, calculate spacing, style the page, and position visual elements.
This creates three separate responsibilities.
The LLM Handles Meaning
The model determines:
- which facts are important;
- how concepts should be grouped;
- whether a flow, sequence, table, tree, timeline, or callout is appropriate;
- what conclusions should be emphasized.
The Renderer Handles Layout
The CLI handles:
- coordinates;
- spacing;
- panel placement;
- diagram geometry;
- visual hierarchy;
- colors;
- themes;
- responsive presentation.
The Browser Handles Consumption
The finished artifact is a standalone HTML page that can be opened locally like a normal file.
This architecture follows an increasingly useful pattern for AI systems: use probabilistic models for ambiguous semantic decisions and deterministic software for repeatable mechanical work.
Why Not Just Ask Claude, Codex, or Cursor to Generate HTML?
Modern coding models can already generate impressive HTML, CSS, JavaScript, and React interfaces.
For a custom application, landing page, prototype, or branded interface, direct frontend generation may be preferable.
The efficiency problem appears when HTML is only being used to present an answer.
If an agent repeatedly regenerates the same card systems, CSS rules, SVG primitives, typography, and responsive layouts, a significant percentage of its output budget is being spent on presentation boilerplate.
Answer me with HTML trades some visual freedom for consistency and speed.
| Approach | Main Strength | Main Weakness |
|---|---|---|
| Plain Markdown | Fast and flexible | Complex answers can become difficult to scan |
| Direct AI-generated HTML | Maximum design freedom | High output-token usage and inconsistent layouts |
| Answer me with HTML | Fast and repeatable visual output | Limited by the renderer's supported components |
The skill is therefore best suited to explaining information, rather than building arbitrary web applications.
Benchmark: How Much Faster Is It?
The project includes a benchmark comparing direct HTML generation against its structured-rendering approach.
The test used Claude Sonnet 5.5 and covered three technical topics. Each configuration was run multiple times and median results were reported.
| Task | Direct HTML Tokens | Skill Tokens | Direct Time | Skill Time |
|---|---|---|---|---|
| TCP handshake and teardown | 8,968 | 972 | 56s | 11s |
| Redis vs Memcached | 6,092 | 935 | 43s | 16s |
| Git merge vs rebase | 5,560 | 862 | 40s | 12s |
| Average | 6,873 | 923 | 46s | 13s |
The reported averages imply roughly:
- 7.4× fewer output tokens;
- 3.6× faster completion.
However, the cost result is just as important.
The project reported approximately:
- Direct HTML: $0.22;
- Answer me with HTML: $0.26.
The reason is that an Agent Skill can introduce additional turns for loading instructions and invoking the renderer. Those turns may reread conversation context, which can offset output-token savings.
The correct interpretation is therefore:
Answer me with HTML can substantially reduce generated presentation tokens and latency, but it does not guarantee a lower total API bill.
The benchmark is also relatively small, so the figures should be treated as evidence for the tested workloads rather than a universal performance guarantee.
How to Install Answer me with HTML
The project requires Node.js 20 or newer.
A standard Agent Skills installation uses:
npx skills add QingYunA/answer-me-with-htmlFor a global non-interactive installation:
npx -y skills add QingYunA/answer-me-with-html -g -yTo target Claude Code specifically:
npx -y skills add QingYunA/answer-me-with-html -g -y -a claude-codeThe project documents compatibility with several major coding-agent environments, including:
- Claude Code;
- Codex;
- Cursor;
- OpenCode.
Installing Answer me with HTML in Claude Code
Claude Code also has a plugin-oriented installation path.
/plugin marketplace add QingYunA/answer-me-with-html
/plugin install answer-me-with-html@answer-me-with-htmlThis approach is useful when the renderer is intended to become part of a regular Claude Code workflow rather than an occasional tool.
What Questions Work Best?
The skill is most useful when the answer naturally benefits from structure or visualization.
Good examples include:
Explain the TCP three-way handshakeMap out how the modules in this repository fit togetherRedis or Memcached for this caching layer?Show the history of Kubernetes as a timelineExplain this authentication flowCompare merge and rebase visuallyShow the request path through this backend
A simple factual or command-line question usually does not need a visual artifact.
For example:
How do I show hidden files with ls?
A short direct answer is better than opening an entire webpage.
This selective behavior is important because visual output becomes counterproductive when it is generated for every trivial interaction.
Using the CLI Directly
The renderer can also be used independently of an AI agent.
A typical CLI path can be assigned like this:
AM=skills/answer-me-with-html/scripts/am.mjsRender an example:
node $AM render examples/tcp.en.mdChoose an output file without opening it automatically:
node $AM render notes.md -o out.html --no-openSelect the shadcn theme:
node $AM render notes.md --theme shadcnRun the writing checker:
node $AM lint notes.mdList available components:
node $AM listInspect configuration:
node $AM configThis direct CLI mode is useful for users who want the renderer without relying entirely on automatic Agent Skill invocation.
Supported Visual Components
Answer me with HTML uses predefined visual structures instead of arbitrary webpage generation.
Flow Diagrams
The flow component is suitable for:
- system architecture;
- data pipelines;
- dependency graphs;
- service relationships;
- decision branches;
- call chains.
Sequence Diagrams
The sequence component works well for:
- API requests;
- TCP exchanges;
- OAuth flows;
- authentication;
- distributed-system communication;
- client-server interactions.
Trees
The tree format works well for:
- repository structures;
- folders;
- module hierarchies;
- taxonomies.
Timelines
The timeline format is suitable for:
- product history;
- project phases;
- release histories;
- chronological events.
Limits
The limits component can visualize values relative to limits or thresholds.
Annotations
The annot component is intended for sentence-level or word-level explanations and corrections.
Key-Value Information
The kv component presents compact metadata and structured facts.
Callouts
The callout component highlights conclusions, warnings, recommendations, and important details.
Comparison Tables
Normal Markdown tables can become structured visual comparisons. Special values such as ok, no, and warn can also be converted into clearer status indicators.
The important concept is that the model describes what the information means, while the renderer determines how that information should look.
Themes and Page Layouts
Answer me with HTML includes two notable visual themes:
blueprint;shadcn.
The blueprint theme resembles an engineering diagram or technical drawing. The shadcn theme is closer to a modern SaaS or developer-tool interface.
The renderer also supports light and dark presentation.
Two common page structures are available:
sheetfor multi-panel visual layouts;docfor longer single-column explanations.
A configuration block can look like:
---
template: sheet
theme: shadcn
title: Cache comparison
subtitle: Redis and Memcached
cols: 3
---This is far more compact than forcing the language model to generate a complete layout system every time.
Why Single-File Offline HTML Matters
One of the project's practical strengths is that rendered output is designed to work as a standalone HTML file.
That provides several advantages:
- pages can be opened without deploying a web application;
- repository explanations can remain local;
- technical reviews can be archived with project documentation;
- generated artifacts can be shared as ordinary files;
- no frontend build pipeline is required;
- pages can remain useful even without an active internet connection.
This local-first model is particularly attractive for coding agents, where the output may contain information about private repositories or internal architecture.
Always-On Mode
By default, the skill can decide whether a page is useful.
An optional always-on integration can make visual artifacts appear more frequently for outputs such as:
- conclusions;
- summaries;
- plans;
- comparisons;
- reviews;
- technical explanations.
For Claude Code, the additional plugin can be installed with:
/plugin marketplace add QingYunA/answer-me-with-html
/plugin install answer-me-with-html-always@answer-me-with-htmlAlways-on mode should still be used selectively. Automatically generating a webpage after every command or short factual response would add friction rather than remove it.
Answer me with HTML Can Also Generate Explainer Videos
The project has expanded beyond static pages with an am video workflow.
The same structured explanation can drive an animated presentation where:
- diagram elements appear progressively;
- arrows animate;
- the camera focuses on relevant nodes;
- matching objects transition between scenes;
- narration remains synchronized with the visuals.
With an appropriate voice provider configured, narration can also be generated automatically.
The workflow can be thought of as:
Structured explanation
↓
Visual HTML
↓
Animated explainer
↓
Narrated videoThe important part is that each output format can reuse the same semantic representation.
That is much harder to achieve when the model directly produces a completely custom webpage from scratch.
ASD-STE100-Inspired Writing Checks
Answer me with HTML also includes writing-quality checks inspired by ASD-STE100 Simplified Technical English.
The purpose is to encourage explanations that are easier to interpret and less ambiguous.
Documented checks include concepts such as:
- shorter procedural sentences;
- shorter descriptive sentences;
- limited paragraph length;
- preference for common words;
- warnings about passive voice;
- corresponding length and style checks for Chinese;
- basic handling of Japanese text.
The default mode provides warnings rather than aggressively blocking content.
A stricter mode can be enabled with:
am config set style strictThe important idea is not formal STE certification. It is that clearer language and clearer visuals reinforce each other.
A visually impressive diagram cannot rescue an explanation that is semantically vague.
Technical Architecture
The project is intentionally lightweight.
Its Node.js package uses a small runtime dependency set, including:
markedfor Markdown processing;@dagrejs/dagrefor graph layout.
The CLI exposes short and long command names:
am;answer-me-with-html.
This compact architecture helps explain why the tool can create relatively rich output from a small amount of generated text.
The AI does not rebuild a rendering engine for every answer. It calls an existing one.
Where Answer me with HTML Works Best
Architecture Explanations
A repository can be represented through:
- service boundaries;
- request flows;
- dependencies;
- database interactions;
- queues;
- events.
A visual diagram can communicate these relationships faster than several paragraphs of prose.
Technology Comparisons
Examples include:
- Redis vs Memcached;
- Git merge vs rebase;
- PostgreSQL vs SQLite;
- REST vs GraphQL;
- queues vs streams.
These questions naturally fit comparison tables and callouts.
Repository Onboarding
A tree combined with a call graph can make an unfamiliar codebase easier to understand than a long folder-by-folder explanation.
Protocols and Distributed Systems
Sequence diagrams are useful for:
- TCP;
- OAuth;
- authentication;
- payments;
- retries;
- messaging;
- consensus protocols.
Plans and Technical Reviews
A multi-panel artifact can combine:
- current architecture;
- proposed changes;
- implementation stages;
- risks;
- trade-offs;
- final recommendations.
This can make an agent-generated plan significantly easier for humans to review.
Where It Is Not the Best Choice
Tiny Factual Questions
A one-line answer does not need an entire webpage.
Highly Custom Frontend Development
If the goal is to create a unique interactive product or branded website, direct React or HTML generation offers substantially more flexibility.
Long Narrative Content
Essays, policy analysis, and other argument-driven writing may work better as traditional documents.
Pixel-Perfect Design
The renderer gains consistency by using predefined visual components. That same constraint makes it less appropriate for highly specific brand or design requirements.
Unverified Information
Better presentation does not make incorrect information accurate. A polished visual artifact can actually make hallucinated information look more authoritative, so factual verification remains essential.
Common Pitfalls
Fewer Output Tokens Do Not Guarantee Lower Cost
The project's own benchmark is a good example. Output tokens fell dramatically, but reported total cost increased slightly because additional agent turns introduced extra context processing.
Do Not Enable Visual Output for Everything
A page is useful when information has meaningful structure. It is unnecessary for trivial responses.
Do Not Treat the Writing Checker as Formal STE Certification
The project applies rules inspired by ASD-STE100. It should not be confused with official standards compliance.
Do Not Expect Unlimited Design Freedom
The renderer is efficient precisely because it uses a constrained set of components. Direct frontend generation remains more appropriate when arbitrary design or interaction is required.
Review Third-Party Agent Skills Before Installing Them
Agent Skills can execute local tools and participate in coding workflows. Their instructions and scripts should be reviewed before being granted access to repositories containing sensitive code, credentials, or private data.
Privacy and Update Behavior
The project's documentation describes generated pages as local artifacts and says its update mechanism checks project version information rather than automatically updating the installed package.
The update check can be disabled with:
am config set update_check offThis local-first approach is useful, but Answer me with HTML should still be treated like any executable third-party developer tool: inspect the source, review updates, and avoid exposing secrets unnecessarily.
Answer me with HTML vs Direct HTML Generation
| Capability | Direct HTML Generation | Answer me with HTML |
|---|---|---|
| Maximum design freedom | Excellent | Limited |
| Fast technical explanations | Moderate | Excellent |
| Output-token efficiency | Variable | Strong |
| Consistent layouts | Variable | Strong |
| Custom JavaScript applications | Excellent | Not the main goal |
| Architecture diagrams | Possible | Built in |
| Offline visual explainers | Possible | Core use case |
| One-off landing pages | Better choice | Usually unnecessary |
A practical rule is:
If the HTML itself is the product, generate the HTML directly. If HTML is only the presentation layer for an explanation, a deterministic renderer can be much more efficient.
Why Answer me with HTML Matters Beyond This Project
The most important idea behind Answer me with HTML is larger than the repository itself.
Most AI interfaces still assume that the final output should be prose.
But different kinds of information have different ideal representations:
- comparisons want tables;
- architectures want graphs;
- protocols want sequence diagrams;
- histories want timelines;
- metrics want charts;
- instructions want steps;
- complex concepts may benefit from animation.
A more capable AI agent should therefore choose not only what to say, but also which information interface best communicates the answer.
Answer me with HTML demonstrates one possible architecture:
- let the LLM produce a semantic intermediate representation;
- keep that representation compact;
- use deterministic software to handle visual layout;
- preserve the underlying source;
- render the answer into the most useful human-facing format.
The same idea could eventually power:
- HTML explainers;
- dashboards;
- documentation;
- presentations;
- interactive simulations;
- architecture diagrams;
- narrated videos.
This is an early example of what can be described as artifact-native AI output.
FAQ
Is Answer me with HTML a Website?
No. It is an open-source Agent Skill and local command-line renderer that generates HTML files.
Does Answer me with HTML Work with Claude Code?
Yes. Claude Code is one of the environments explicitly targeted by the project, including a dedicated plugin installation flow.
Does It Work with Codex and Cursor?
Yes. The project documents compatibility with Codex and Cursor through the Agent Skills ecosystem.
Does It Save Tokens?
The published benchmark shows a substantial reduction in output tokens for the tested workloads, with average output dropping from 6,873 tokens to 923 tokens.
That does not mean total API cost will fall by the same percentage.
Can It Generate Diagrams?
Yes. Supported structures include flows, sequences, trees, timelines, comparisons, limits, annotations, and callouts.
Can It Generate Videos?
Yes. Its video workflow can reuse structured explanation data to produce animated explainers and, with the required local tooling, video output.
Does the HTML Work Offline?
The project's design emphasizes self-contained HTML output without depending on a normal frontend deployment process.
Is It Free?
The project is open source under the MIT license. AI model usage, voice generation, or other external services may still have their own costs.
Conclusion
Answer me with HTML is built around a strong architectural idea: LLMs should spend their output budget on reasoning and useful content rather than repeatedly regenerating presentation boilerplate.
For complex technical explanations, architecture reviews, comparisons, repository onboarding, and implementation plans, its structured-content-plus-deterministic-renderer approach can produce output that is faster to generate and easier to understand than either plain Markdown or fully model-generated HTML.
The early benchmark is promising, but its results should be interpreted carefully. It demonstrates significant reductions in output length and latency for a small set of tests, not guaranteed cost savings for every workload.
The larger trend is more important. As AI agents generate more information, presentation becomes part of the reasoning interface itself. Markdown will remain useful, but diagrams, dashboards, interactive pages, and generated explainers are likely to become increasingly normal forms of AI output.
For developers using Claude Code, Codex, Cursor, or OpenCode, Answer me with HTML is worth exploring as an early example of that shift from chat-native AI to artifact-native AI.
Continue Reading
More articles connected to the same themes, protocols, and tools.
Referenced Tools
Browse entries that are adjacent to the topics covered in this article.






