# Answer me with HTML: The Agent Skill Turning AI Answers Into Visual Pages

Answer me with HTML turns Claude Code, Codex and Cursor responses into visual offline pages while reducing output tokens and wait time.

Canonical URL: https://aiidelist.com/blog/answer-me-with-html

Language: en

Published: 2026-10-04

Updated: 2026-10-04

## Key Takeaways

- **Answer me with HTML is an Agent Skill that converts complex AI answers into readable, self-contained HTML pages.** Instead of making the model generate an entire webpage, it lets the model produce compact structured content while a local renderer handles layout, diagrams, CSS, themes, and SVG positioning.
- Its core idea is simple: **let the LLM decide what to explain, and let deterministic code decide how to display it.**
- In the project's published benchmark using Claude Sonnet 5.5, average output fell from **6,873 tokens to 923 tokens**, while completion time fell from **46 seconds to 13 seconds**.
- The same benchmark reported roughly **$0.22 for direct HTML versus $0.26 with the skill**, meaning fewer output tokens do not automatically translate into lower total API cost.
- The project supports Agent Skill-compatible environments including **Claude Code, Codex, Cursor, and OpenCode**.
- It can render flow diagrams, sequence diagrams, trees, timelines, comparison tables, annotations, callouts, and narrated explainer videos.

## What Is Answer me with HTML?

**Answer me with HTML** is an open-source Agent Skill and local CLI designed to turn substantial AI explanations into visual HTML artifacts instead of long walls of Markdown or terminal text.

The project was created by QingYunA and follows a straightforward principle: complex answers are often easier to understand when the presentation format matches the structure of the information.

Instead of telling a coding model to generate hundreds of lines of HTML, CSS, SVG, and JavaScript, the skill asks the model to produce a much smaller structured draft. The bundled `am` renderer then converts that draft into the final page.

The workflow can be simplified as:

```text
User question
    ↓
AI agent
    ↓
Compact structured Markdown
    ↓
am CLI
    ↓
Layout engine + components + SVG
    ↓
Single HTML file
```

This makes Answer me with HTML more than an HTML prompt template. It acts as a reusable **presentation layer for AI agents**.

## Why Is Answer me with HTML Getting Attention?

The project arrived during a broader discussion about improving how humans consume AI-generated information.

In early October 2026, Andrej Karpathy highlighted a progression for AI explanations that moves beyond ordinary prose toward controlled writing, diagrams, interactive HTML, and eventually generated explainer videos.

Answer me with HTML takes a similar direction but introduces an important optimization: **the language model does not manually build the final webpage**.

Instead, the AI handles semantic decisions while deterministic software performs repetitive visual work.

That distinction matters because generating HTML directly can cause models to spend thousands of output tokens on:

- wrapper elements;
- utility classes;
- CSS declarations;
- SVG coordinates;
- card layouts;
- responsive rules;
- dark-mode styles;
- diagram arrows and labels.

Those tokens make the result prettier, but they often contribute very little additional reasoning.

Answer me with HTML tries to move that deterministic work out of the model and into reusable software.

## How Answer me with HTML Works

The model generates an extended Markdown-style representation instead of a complete webpage.

For example, a TCP explanation can be represented with something conceptually similar to:

```text
---
title: TCP three-way handshake
---

## Three-way handshake

sequence:
Client -> Server: SYN
Server -> Client: SYN+ACK
Client -> Server: ACK
```

The renderer then decides how to draw the sequence, calculate spacing, style the page, and position visual elements.

This creates three separate responsibilities.

### The LLM Handles Meaning

The model determines:

- which facts are important;
- how concepts should be grouped;
- whether a flow, sequence, table, tree, timeline, or callout is appropriate;
- what conclusions should be emphasized.

### The Renderer Handles Layout

The CLI handles:

- coordinates;
- spacing;
- panel placement;
- diagram geometry;
- visual hierarchy;
- colors;
- themes;
- responsive presentation.

### The Browser Handles Consumption

The finished artifact is a standalone HTML page that can be opened locally like a normal file.

This architecture follows an increasingly useful pattern for AI systems: **use probabilistic models for ambiguous semantic decisions and deterministic software for repeatable mechanical work**.

## Why Not Just Ask Claude, Codex, or Cursor to Generate HTML?

Modern coding models can already generate impressive HTML, CSS, JavaScript, and React interfaces.

For a custom application, landing page, prototype, or branded interface, direct frontend generation may be preferable.

The efficiency problem appears when HTML is only being used to present an answer.

If an agent repeatedly regenerates the same card systems, CSS rules, SVG primitives, typography, and responsive layouts, a significant percentage of its output budget is being spent on presentation boilerplate.

Answer me with HTML trades some visual freedom for consistency and speed.

| Approach | Main Strength | Main Weakness |
|---|---|---|
| Plain Markdown | Fast and flexible | Complex answers can become difficult to scan |
| Direct AI-generated HTML | Maximum design freedom | High output-token usage and inconsistent layouts |
| Answer me with HTML | Fast and repeatable visual output | Limited by the renderer's supported components |

The skill is therefore best suited to **explaining information**, rather than building arbitrary web applications.

## Benchmark: How Much Faster Is It?

The project includes a benchmark comparing direct HTML generation against its structured-rendering approach.

The test used **Claude Sonnet 5.5** and covered three technical topics. Each configuration was run multiple times and median results were reported.

| Task | Direct HTML Tokens | Skill Tokens | Direct Time | Skill Time |
|---|---:|---:|---:|---:|
| TCP handshake and teardown | 8,968 | 972 | 56s | 11s |
| Redis vs Memcached | 6,092 | 935 | 43s | 16s |
| Git merge vs rebase | 5,560 | 862 | 40s | 12s |
| **Average** | **6,873** | **923** | **46s** | **13s** |

The reported averages imply roughly:

- **7.4× fewer output tokens**;
- **3.6× faster completion**.

However, the cost result is just as important.

The project reported approximately:

- Direct HTML: **$0.22**;
- Answer me with HTML: **$0.26**.

The reason is that an Agent Skill can introduce additional turns for loading instructions and invoking the renderer. Those turns may reread conversation context, which can offset output-token savings.

The correct interpretation is therefore:

> Answer me with HTML can substantially reduce generated presentation tokens and latency, but it does not guarantee a lower total API bill.

The benchmark is also relatively small, so the figures should be treated as evidence for the tested workloads rather than a universal performance guarantee.

## How to Install Answer me with HTML

The project requires **Node.js 20 or newer**.

A standard Agent Skills installation uses:

```bash
npx skills add QingYunA/answer-me-with-html
```

For a global non-interactive installation:

```bash
npx -y skills add QingYunA/answer-me-with-html -g -y
```

To target Claude Code specifically:

```bash
npx -y skills add QingYunA/answer-me-with-html -g -y -a claude-code
```

The project documents compatibility with several major coding-agent environments, including:

- Claude Code;
- Codex;
- Cursor;
- OpenCode.

## Installing Answer me with HTML in Claude Code

Claude Code also has a plugin-oriented installation path.

```text
/plugin marketplace add QingYunA/answer-me-with-html
/plugin install answer-me-with-html@answer-me-with-html
```

This approach is useful when the renderer is intended to become part of a regular Claude Code workflow rather than an occasional tool.

## What Questions Work Best?

The skill is most useful when the answer naturally benefits from structure or visualization.

Good examples include:

- `Explain the TCP three-way handshake`
- `Map out how the modules in this repository fit together`
- `Redis or Memcached for this caching layer?`
- `Show the history of Kubernetes as a timeline`
- `Explain this authentication flow`
- `Compare merge and rebase visually`
- `Show the request path through this backend`

A simple factual or command-line question usually does not need a visual artifact.

For example:

`How do I show hidden files with ls?`

A short direct answer is better than opening an entire webpage.

This selective behavior is important because visual output becomes counterproductive when it is generated for every trivial interaction.

## Using the CLI Directly

The renderer can also be used independently of an AI agent.

A typical CLI path can be assigned like this:

```bash
AM=skills/answer-me-with-html/scripts/am.mjs
```

Render an example:

```bash
node $AM render examples/tcp.en.md
```

Choose an output file without opening it automatically:

```bash
node $AM render notes.md -o out.html --no-open
```

Select the shadcn theme:

```bash
node $AM render notes.md --theme shadcn
```

Run the writing checker:

```bash
node $AM lint notes.md
```

List available components:

```bash
node $AM list
```

Inspect configuration:

```bash
node $AM config
```

This direct CLI mode is useful for users who want the renderer without relying entirely on automatic Agent Skill invocation.

## Supported Visual Components

Answer me with HTML uses predefined visual structures instead of arbitrary webpage generation.

### Flow Diagrams

The `flow` component is suitable for:

- system architecture;
- data pipelines;
- dependency graphs;
- service relationships;
- decision branches;
- call chains.

### Sequence Diagrams

The `sequence` component works well for:

- API requests;
- TCP exchanges;
- OAuth flows;
- authentication;
- distributed-system communication;
- client-server interactions.

### Trees

The `tree` format works well for:

- repository structures;
- folders;
- module hierarchies;
- taxonomies.

### Timelines

The `timeline` format is suitable for:

- product history;
- project phases;
- release histories;
- chronological events.

### Limits

The `limits` component can visualize values relative to limits or thresholds.

### Annotations

The `annot` component is intended for sentence-level or word-level explanations and corrections.

### Key-Value Information

The `kv` component presents compact metadata and structured facts.

### Callouts

The `callout` component highlights conclusions, warnings, recommendations, and important details.

### Comparison Tables

Normal Markdown tables can become structured visual comparisons. Special values such as `ok`, `no`, and `warn` can also be converted into clearer status indicators.

The important concept is that the model describes **what the information means**, while the renderer determines how that information should look.

## Themes and Page Layouts

Answer me with HTML includes two notable visual themes:

- `blueprint`;
- `shadcn`.

The blueprint theme resembles an engineering diagram or technical drawing. The shadcn theme is closer to a modern SaaS or developer-tool interface.

The renderer also supports light and dark presentation.

Two common page structures are available:

- `sheet` for multi-panel visual layouts;
- `doc` for longer single-column explanations.

A configuration block can look like:

```text
---
template: sheet
theme: shadcn
title: Cache comparison
subtitle: Redis and Memcached
cols: 3
---
```

This is far more compact than forcing the language model to generate a complete layout system every time.

## Why Single-File Offline HTML Matters

One of the project's practical strengths is that rendered output is designed to work as a standalone HTML file.

That provides several advantages:

- pages can be opened without deploying a web application;
- repository explanations can remain local;
- technical reviews can be archived with project documentation;
- generated artifacts can be shared as ordinary files;
- no frontend build pipeline is required;
- pages can remain useful even without an active internet connection.

This local-first model is particularly attractive for coding agents, where the output may contain information about private repositories or internal architecture.

## Always-On Mode

By default, the skill can decide whether a page is useful.

An optional always-on integration can make visual artifacts appear more frequently for outputs such as:

- conclusions;
- summaries;
- plans;
- comparisons;
- reviews;
- technical explanations.

For Claude Code, the additional plugin can be installed with:

```text
/plugin marketplace add QingYunA/answer-me-with-html
/plugin install answer-me-with-html-always@answer-me-with-html
```

Always-on mode should still be used selectively. Automatically generating a webpage after every command or short factual response would add friction rather than remove it.

## Answer me with HTML Can Also Generate Explainer Videos

The project has expanded beyond static pages with an `am video` workflow.

The same structured explanation can drive an animated presentation where:

- diagram elements appear progressively;
- arrows animate;
- the camera focuses on relevant nodes;
- matching objects transition between scenes;
- narration remains synchronized with the visuals.

With an appropriate voice provider configured, narration can also be generated automatically.

The workflow can be thought of as:

```text
Structured explanation
        ↓
Visual HTML
        ↓
Animated explainer
        ↓
Narrated video
```

The important part is that each output format can reuse the same semantic representation.

That is much harder to achieve when the model directly produces a completely custom webpage from scratch.

## ASD-STE100-Inspired Writing Checks

Answer me with HTML also includes writing-quality checks inspired by **ASD-STE100 Simplified Technical English**.

The purpose is to encourage explanations that are easier to interpret and less ambiguous.

Documented checks include concepts such as:

- shorter procedural sentences;
- shorter descriptive sentences;
- limited paragraph length;
- preference for common words;
- warnings about passive voice;
- corresponding length and style checks for Chinese;
- basic handling of Japanese text.

The default mode provides warnings rather than aggressively blocking content.

A stricter mode can be enabled with:

```bash
am config set style strict
```

The important idea is not formal STE certification. It is that **clearer language and clearer visuals reinforce each other**.

A visually impressive diagram cannot rescue an explanation that is semantically vague.

## Technical Architecture

The project is intentionally lightweight.

Its Node.js package uses a small runtime dependency set, including:

- `marked` for Markdown processing;
- `@dagrejs/dagre` for graph layout.

The CLI exposes short and long command names:

- `am`;
- `answer-me-with-html`.

This compact architecture helps explain why the tool can create relatively rich output from a small amount of generated text.

The AI does not rebuild a rendering engine for every answer. It calls an existing one.

## Where Answer me with HTML Works Best

### Architecture Explanations

A repository can be represented through:

- service boundaries;
- request flows;
- dependencies;
- database interactions;
- queues;
- events.

A visual diagram can communicate these relationships faster than several paragraphs of prose.

### Technology Comparisons

Examples include:

- Redis vs Memcached;
- Git merge vs rebase;
- PostgreSQL vs SQLite;
- REST vs GraphQL;
- queues vs streams.

These questions naturally fit comparison tables and callouts.

### Repository Onboarding

A tree combined with a call graph can make an unfamiliar codebase easier to understand than a long folder-by-folder explanation.

### Protocols and Distributed Systems

Sequence diagrams are useful for:

- TCP;
- OAuth;
- authentication;
- payments;
- retries;
- messaging;
- consensus protocols.

### Plans and Technical Reviews

A multi-panel artifact can combine:

- current architecture;
- proposed changes;
- implementation stages;
- risks;
- trade-offs;
- final recommendations.

This can make an agent-generated plan significantly easier for humans to review.

## Where It Is Not the Best Choice

### Tiny Factual Questions

A one-line answer does not need an entire webpage.

### Highly Custom Frontend Development

If the goal is to create a unique interactive product or branded website, direct React or HTML generation offers substantially more flexibility.

### Long Narrative Content

Essays, policy analysis, and other argument-driven writing may work better as traditional documents.

### Pixel-Perfect Design

The renderer gains consistency by using predefined visual components. That same constraint makes it less appropriate for highly specific brand or design requirements.

### Unverified Information

Better presentation does not make incorrect information accurate. A polished visual artifact can actually make hallucinated information look more authoritative, so factual verification remains essential.

## Common Pitfalls

### Fewer Output Tokens Do Not Guarantee Lower Cost

The project's own benchmark is a good example. Output tokens fell dramatically, but reported total cost increased slightly because additional agent turns introduced extra context processing.

### Do Not Enable Visual Output for Everything

A page is useful when information has meaningful structure. It is unnecessary for trivial responses.

### Do Not Treat the Writing Checker as Formal STE Certification

The project applies rules inspired by ASD-STE100. It should not be confused with official standards compliance.

### Do Not Expect Unlimited Design Freedom

The renderer is efficient precisely because it uses a constrained set of components. Direct frontend generation remains more appropriate when arbitrary design or interaction is required.

### Review Third-Party Agent Skills Before Installing Them

Agent Skills can execute local tools and participate in coding workflows. Their instructions and scripts should be reviewed before being granted access to repositories containing sensitive code, credentials, or private data.

## Privacy and Update Behavior

The project's documentation describes generated pages as local artifacts and says its update mechanism checks project version information rather than automatically updating the installed package.

The update check can be disabled with:

```bash
am config set update_check off
```

This local-first approach is useful, but Answer me with HTML should still be treated like any executable third-party developer tool: inspect the source, review updates, and avoid exposing secrets unnecessarily.

## Answer me with HTML vs Direct HTML Generation

| Capability | Direct HTML Generation | Answer me with HTML |
|---|---|---|
| Maximum design freedom | **Excellent** | Limited |
| Fast technical explanations | Moderate | **Excellent** |
| Output-token efficiency | Variable | **Strong** |
| Consistent layouts | Variable | **Strong** |
| Custom JavaScript applications | **Excellent** | Not the main goal |
| Architecture diagrams | Possible | **Built in** |
| Offline visual explainers | Possible | **Core use case** |
| One-off landing pages | **Better choice** | Usually unnecessary |

A practical rule is:

**If the HTML itself is the product, generate the HTML directly. If HTML is only the presentation layer for an explanation, a deterministic renderer can be much more efficient.**

## Why Answer me with HTML Matters Beyond This Project

The most important idea behind Answer me with HTML is larger than the repository itself.

Most AI interfaces still assume that the final output should be prose.

But different kinds of information have different ideal representations:

- comparisons want tables;
- architectures want graphs;
- protocols want sequence diagrams;
- histories want timelines;
- metrics want charts;
- instructions want steps;
- complex concepts may benefit from animation.

A more capable AI agent should therefore choose not only **what to say**, but also **which information interface best communicates the answer**.

Answer me with HTML demonstrates one possible architecture:

1. let the LLM produce a semantic intermediate representation;
2. keep that representation compact;
3. use deterministic software to handle visual layout;
4. preserve the underlying source;
5. render the answer into the most useful human-facing format.

The same idea could eventually power:

- HTML explainers;
- dashboards;
- documentation;
- presentations;
- interactive simulations;
- architecture diagrams;
- narrated videos.

This is an early example of what can be described as **artifact-native AI output**.

## FAQ

### Is Answer me with HTML a Website?

No. It is an open-source Agent Skill and local command-line renderer that generates HTML files.

### Does Answer me with HTML Work with Claude Code?

Yes. Claude Code is one of the environments explicitly targeted by the project, including a dedicated plugin installation flow.

### Does It Work with Codex and Cursor?

Yes. The project documents compatibility with Codex and Cursor through the Agent Skills ecosystem.

### Does It Save Tokens?

The published benchmark shows a substantial reduction in output tokens for the tested workloads, with average output dropping from 6,873 tokens to 923 tokens.

That does not mean total API cost will fall by the same percentage.

### Can It Generate Diagrams?

Yes. Supported structures include flows, sequences, trees, timelines, comparisons, limits, annotations, and callouts.

### Can It Generate Videos?

Yes. Its video workflow can reuse structured explanation data to produce animated explainers and, with the required local tooling, video output.

### Does the HTML Work Offline?

The project's design emphasizes self-contained HTML output without depending on a normal frontend deployment process.

### Is It Free?

The project is open source under the MIT license. AI model usage, voice generation, or other external services may still have their own costs.

## Conclusion

Answer me with HTML is built around a strong architectural idea: **LLMs should spend their output budget on reasoning and useful content rather than repeatedly regenerating presentation boilerplate**.

For complex technical explanations, architecture reviews, comparisons, repository onboarding, and implementation plans, its structured-content-plus-deterministic-renderer approach can produce output that is faster to generate and easier to understand than either plain Markdown or fully model-generated HTML.

The early benchmark is promising, but its results should be interpreted carefully. It demonstrates significant reductions in output length and latency for a small set of tests, not guaranteed cost savings for every workload.

The larger trend is more important. As AI agents generate more information, presentation becomes part of the reasoning interface itself. Markdown will remain useful, but diagrams, dashboards, interactive pages, and generated explainers are likely to become increasingly normal forms of AI output.

For developers using Claude Code, Codex, Cursor, or OpenCode, Answer me with HTML is worth exploring as an early example of that shift from **chat-native AI to artifact-native AI**.
