# Claude Sonnet 5.5 Examples: 14 Real Demos, Coding Tests, and Agent Use Cases

Explore 14 Claude Sonnet 5.5 examples covering coding, apps, agents, games, finance, visual tasks, slides, and reproducible prompts.

Canonical URL: https://aiidelist.com/blog/claude-sonnet-5-5-examples

Language: en

Published: 2026-09-29

Updated: 2026-09-29

## Key Takeaways

Claude Sonnet 5.5 is best understood through complete workflows rather than isolated benchmark scores. Its strongest public examples cover creative coding, long-horizon visual agents, app building, large-codebase work, finance, slides, document verification, and enterprise agents.

- Anthropic's official creative-coding demos include a **400-starling murmuration**, **wind-shaped sand dunes**, and a **clock made from 24 smaller clocks**, all requested as single-file HTML experiences.
- Anthropic reports that Sonnet 5.5 is the **first Sonnet model to beat Pokémon Red using only screenshots**, making it a notable vision-plus-agent example.
- Base44 evaluated **118 real app builds** and reported 3.6 iterations per build on average for Sonnet 5.5 versus 7.7 for Opus 5.
- Unity reported **90% completion** on its multi-step Unity Editor and coding benchmark with runtime verification.
- Anthropic's internal knowledge-work example turned earnings materials, an earnings-call transcript, and a slide template into a **10-slide operating review**.
- The most convincing Sonnet 5.5 examples include planning, execution, inspection, and verification—not just code or prose generation.

## What Is Claude Sonnet 5.5?

Claude Sonnet 5.5 was released on September 28, 2026. Anthropic positions it as a balanced model for coding, documents, slides, spreadsheets, design, and well-scoped professional work.

| Specification | Claude Sonnet 5.5 |
| --- | --- |
| Model ID | `claude-sonnet-5-5` |
| Context window | 1M tokens |
| Maximum standard output | 128K tokens |
| Input price | $2 / 1M tokens |
| Output price | $10 / 1M tokens |
| Cache read | $0.20 / 1M tokens |
| Inputs | Text and images |
| Reliable knowledge cutoff | June 2026 |

## Example 1: 400-Starling Murmuration

One of Anthropic's headline demonstrations starts with a very short prompt:

```text
A murmuration of 400 starlings in one HTML file
```

A convincing result requires more than drawing 400 dots. A flocking simulation typically needs local rules for separation, alignment, cohesion, steering limits, boundary handling, and efficient frame updates.

This makes the example a useful test of:

- JavaScript and Canvas;
- simulation logic;
- animation performance;
- visual polish;
- translating a vague visual concept into runnable code.

## Example 2: Wind-Shaped Sand Dunes

Another official creative-coding prompt is:

```text
Wind shaping sand dunes in one HTML file
```

This is a useful procedural-graphics test because the model must communicate wind direction, accumulation, erosion, and dune structure instead of simply displaying random noise.

A strong implementation may combine Canvas or WebGL with procedural noise, a particle or height field, wind vectors, deposition rules, elevation shading, and real-time animation.

## Example 3: A Clock Made of 24 Smaller Clocks

Anthropic also demonstrates:

```text
A clock made of 24 small clocks in one HTML file
```

The challenge is coordination. Each small clock has independent hands, but the full set must form a larger visual composition.

This tests geometric mapping, coordinated rotations, smooth interpolation, time state, responsive layout, and visual hierarchy.

## Example 4: Pokémon Red Using Only Screenshots

Anthropic reports that Sonnet 5.5 is the first Sonnet model to beat Pokémon Red using only screenshots.

The workflow can be represented as:

```text
Screenshot
→ identify game state
→ recall recent progress
→ infer the next objective
→ choose an action
→ observe the new screenshot
→ update the plan
→ repeat
```

The task combines visual recognition, navigation, state tracking, long-horizon planning, action selection, and recovery from mistakes.

Its significance extends beyond games. The same pattern appears in browser automation, QA agents, desktop agents, and systems that must operate through changing graphical interfaces.

## Example 5: 118 Real App Builds at Base44

Base44 evaluated Sonnet 5.5 across **118 real app builds**. According to the customer result published by Anthropic, Sonnet 5.5 averaged **3.6 iterations per build**, compared with 7.7 for Opus 5, while also showing strong tool-call reliability.

A reproducible app-building prompt could look like this:

```text
Build a production-quality SaaS analytics dashboard.

Requirements:
- React and TypeScript
- responsive navigation
- MRR, revenue, churn and active-user charts
- customer table with search and filters
- date-range control
- empty, loading and error states
- dark and light themes
- realistic sample data
- mobile layout

Run the application, inspect the result, fix runtime and visual problems,
and do not report completion until a real build or test succeeds.
```

The important point is verification. A coding example should not count as complete merely because the model generated plausible code.

## Example 6: Unity Game Development With Runtime Verification

Unity reported **90% completion** on a multi-step Unity Editor and coding benchmark where projects were reopened and checked at runtime.

A Unity agent may need to coordinate C# scripts, GameObjects, serialized properties, scene state, physics, animation, UI, and runtime behavior.

Example prompt:

```text
Add a stamina-based sprint system to this Unity project.

Requirements:
- hold Shift to sprint
- stamina drains while sprinting
- stamina regenerates after a short delay
- sprint stops automatically at zero stamina
- update the existing HUD with a stamina bar
- preserve controller support
- do not change unrelated movement behavior

Implement the feature, run the relevant scene, verify it at runtime,
and report any check that could not be completed.
```

## Example 7: Large-Codebase Work at Epic Games

Epic Games' early evaluation focused on codebase-scale work involving gameplay-system architecture, system-design audits, data-flow reviews, and tasks spanning tens of thousands of lines of code.

A stronger repository test than a generic summary would be:

```text
Trace the gameplay ability flow from player input to the final replicated effect.

Identify:
1. where input is interpreted,
2. where abilities are validated,
3. how server authority is enforced,
4. how state is replicated,
5. where animation and UI feedback are triggered,
6. duplicated state,
7. race-condition risks,
8. the smallest safe architectural improvements.

Reference the relevant files and functions.
Do not propose a rewrite unless the current architecture makes the requested behavior unsafe.
```

## Example 8: Earnings Materials to a 10-Slide Operating Review

Anthropic reports an internal test in which Sonnet 5.5 received quarterly earnings materials, an earnings-call transcript, and an existing slide template, then produced a **10-slide operating review**.

This kind of task combines document retrieval, numerical accuracy, financial interpretation, narrative prioritization, chart selection, slide structure, and template adherence.

A serious reproduction should verify revenue, margins, guidance, period comparisons, chart units, and material numerical claims against the underlying documents.

## Example 9: 2,441 Finance Tasks at Balyasny

Balyasny Asset Management evaluated Sonnet 5.5 on a private suite of **2,441 finance tasks** covering Q&A, extraction, analysis, forecasting, analyst search, and retrieval.

The broader lesson is that strong agent performance is not only about producing more reasoning. It can also come from **doing less unnecessary work while preserving accuracy**.

A realistic finance test can ask the model to reconcile a 10-Q, earnings transcript, investor deck, and historical model, then flag disagreements and produce a scenario table.

## Example 10: Source-Document Rechecking at Box

Box's evaluation highlights a useful behavior: returning to the source document instead of trusting an intermediate summary.

A document-verification prompt can look like this:

```text
Verify every revenue, margin, growth and guidance figure in the draft memo
against the original source documents.

For every discrepancy, return:
- figure in the draft
- verified figure
- document
- page or section
- explanation of the mismatch

Do not use the draft memo itself as evidence.
```

This tests evidence-grounded correction, not generic summarization.

## Example 11: Slackbot With Fewer Steps

Slack reported that Sonnet 5.5 performed better than Sonnet 5 on almost all of its offline Slackbot evaluations while taking fewer steps and using fewer output tokens.

A practical enterprise-knowledge example is:

```text
Find the important Project Atlas discussions from the last 30 days.

Return:
- confirmed decisions
- unresolved blockers
- owners
- deadlines
- open disagreements
- the most relevant conversation references

Do not treat proposals as decisions.
```

The difficult part is distinguishing decision state, chronology, ownership, and unresolved work.

## Example 12: Zendesk Support Decisions

Zendesk evaluated Sonnet 5.5 on real support use cases involving replies and escalation decisions.

A structured example can separate policy reasoning from writing quality:

```text
Customer:
I cancelled yesterday but was charged for another month.

Use the supplied billing policy, subscription history, refund rules,
and current account state.

Return one decision:
REFUND
CREDIT
ESCALATE
NO_ACTION

Then explain the evidence and draft a concise customer-facing reply.
```

## Example 13: Lovable and Efficient Coding Loops

Lovable reported fewer tool calls and shell runs in its Sonnet 5.5 coding evaluations.

That matters because every unnecessary tool call adds latency, cost, and another potential failure point.

For coding agents, the goal should be the **smallest reliable sequence of actions that produces a verified result**.

## Example 14: Atlassian Rovo Agents

Atlassian reported faster Rovo Agent execution with Sonnet 5.5 than with Sonnet 5.

Across Base44, Unity, Slack, Zendesk, Lovable, and Atlassian, a consistent pattern appears: model quality should be evaluated together with **workflow efficiency**.

Track:

1. total tool calls;
2. failed tool calls;
3. wall-clock task time;
4. tokens per successful task;
5. human interventions.

## Best Sonnet 5.5 Example Categories

| Category | Example | What It Tests |
| --- | --- | --- |
| Creative coding | 400-starling murmuration | Simulation and browser graphics |
| Procedural graphics | Wind-shaped dunes | Physical modeling and animation |
| Generative UI | 24-clock composition | Geometry and coordinated UI |
| Vision agent | Pokémon Red | Screenshot understanding and long-horizon action |
| App building | Base44 app builds | Full workflow and tool reliability |
| Game development | Unity benchmark | Editor actions and runtime verification |
| Large codebases | Epic Games | Architecture and multi-hour work |
| Slides | Operating review | Documents, analysis and design |
| Finance | Balyasny evaluation | Extraction, forecasting and efficiency |
| Document verification | Box | Source rechecking and correction |
| Enterprise knowledge | Slackbot | Retrieval and decision synthesis |
| Customer support | Zendesk | Policy reasoning and escalation |
| Coding agents | Lovable | Tool efficiency |
| Enterprise agents | Atlassian Rovo | Workflow speed |

## How to Prompt Sonnet 5.5 for Better Results

### Choose Effort Deliberately

For agentic coding, start with `medium` for well-specified work and move to `high` for harder or longer tasks. Reserve `xhigh` and `max` for workloads where evaluation shows a measurable gain.

### Require Real Verification

A stronger completion instruction is:

```text
After changing runnable code, execute a real test, build, type-check,
or runtime check that exercises the change before reporting completion.

If no meaningful check can run, state exactly what could not be verified.
```

### Control Scope at Very High Effort

At very high effort, additional review or hardening can become counterproductive for tightly scoped work.

Use an explicit stopping rule:

```text
When the requested work is complete and its checks pass, stop.
Do not start additional refactors, hardening, documentation,
or reviewer-agent passes unless they were explicitly requested.
```

## Seven Original Sonnet 5.5 Examples Worth Reproducing

### 1. Interactive Three.js Miniature World

Build a spherical world with terrain, roads, buildings, vehicles, clouds, day-night lighting, orbit controls, and responsive performance.

### 2. Physics Playground

Create gravity, draggable bodies, collisions, joints, reset controls, and live parameter editing.

### 3. Procedural City Generator

Generate roads, blocks, parks, buildings, traffic, camera controls, and deterministic seeds.

### 4. Production SaaS Dashboard

Use an existing React repository and require charts, tables, filters, responsive design, accessibility, loading states, tests, and build verification.

### 5. Complete Browser Game

Require a title screen, controls, scoring, win and lose states, restart flow, sound toggle, mobile controls, and saved high score.

### 6. Financial Filing to Investment Memo

Provide several filings and transcripts and require every numerical claim to be traceable to the documents.

### 7. Repository Feature Implementation

Give the model a feature request touching frontend, backend, tests, and database state.

The result should only pass when the implementation works, unrelated files remain untouched, tests pass, the build succeeds, and the completion report accurately describes what was verified.

## How to Evaluate a Sonnet 5.5 Example Fairly

For every run, record:

- **Prompt:** exact task text.
- **Model:** `claude-sonnet-5-5`.
- **Effort:** low, medium, high, xhigh, or max.
- **Environment:** Claude Code, API harness, browser agent, or custom runtime.
- **Input tokens:** including documents and repository context.
- **Output tokens:** task output and reasoning usage where available.
- **Tool calls:** successful and failed.
- **Wall-clock time:** start to verified completion.
- **Verification:** build, tests, runtime check, visual inspection, or expert review.
- **Human edits:** changes required after the model stopped.
- **Artifact:** code, screenshot, video, deck, spreadsheet, or report.

The most useful production metric is **cost per accepted result**, not raw token count.

## Common Testing Mistakes

### Showing Only the Final Screenshot

A polished interface does not prove that forms, navigation, responsive behavior, state, or builds actually work.

### Hiding the Effort Level

Effort changes latency, token use, and behavior. Without it, two model runs are difficult to compare fairly.

### Comparing Different Harnesses

Shell access, browser tools, subagents, retry policies, file search, and system prompts can materially change outcomes.

### Treating Customer Testimonials as Independent Benchmarks

Customer evaluations published by a model vendor are useful real-world signals, but private datasets and harnesses are usually not independently reproducible.

### Using Max Effort by Default

More reasoning is not automatically better. Extra review and unnecessary work can increase cost and scope.

### Accepting the Model's Claim That Work Is Done

For code, completion should mean a real build, test, type-check, or runtime result—not merely a sentence saying the task is finished.

## FAQ

### What are the best Claude Sonnet 5.5 examples?

The strongest public examples include the 400-starling murmuration, wind-shaped sand dunes, the 24-clock composition, the screenshot-only Pokémon Red run, Base44 app builds, Unity development tasks, Epic Games codebase analysis, the earnings operating-review deck, and enterprise evaluations from Balyasny, Box, Slack, Zendesk, Lovable, and Atlassian.

### Is Claude Sonnet 5.5 good for coding?

Yes. Coding is one of its main use cases, especially when tasks require repository understanding, tool use, implementation, and verification.

### Can Sonnet 5.5 build complete apps?

Yes, but a serious evaluation should verify the application at runtime instead of judging only the generated code or screenshot.

### Can Sonnet 5.5 create animated or 3D websites?

Its official creative-coding demos demonstrate sophisticated browser animation. Three.js scenes, procedural graphics, and interactive 3D experiences are natural workloads to evaluate independently.

### Can Sonnet 5.5 work with images?

Yes. The API accepts image inputs, and Anthropic reports a long-horizon Pokémon Red agent demonstration driven by screenshots.

### How large is the Sonnet 5.5 context window?

Sonnet 5.5 supports a 1M-token context window, making it suitable for large repositories, document sets, and long-running knowledge workflows.

### Which effort setting should be used?

Start with `medium` for well-specified agentic coding tasks and move to `high` for harder work. Use `xhigh` or `max` only when evaluation demonstrates enough additional quality to justify the extra cost and latency.

## Conclusion

The best Claude Sonnet 5.5 examples are **end-to-end tasks with observable success conditions**, not generic chatbot prompts.

Creative-coding demos show that short instructions can become sophisticated interactive programs. Pokémon Red illustrates long-horizon visual action. Base44, Unity, and Epic Games demonstrate coding workflows, while finance, slides, document verification, support, and enterprise-agent examples show that the model's value extends well beyond software generation.

For developers evaluating Sonnet 5.5, the most useful next step is to build a small internal task suite covering **frontend design, a repository-level code change, a visual agent task, and a document-heavy knowledge task**.

Record effort, tokens, tool calls, latency, verification, and human edits for every run.

That creates something more useful than another benchmark table: **evidence of how Claude Sonnet 5.5 performs on the work that actually matters to the product.**
