Google · Image generation & editing

Nano Banana 2.1

Google's high-efficiency image generation and editing model

Generate and edit 1K–4K images, combine up to 14 reference images, keep characters and products consistent, render legible text and choose a thinking level for complex visual tasks.

  • Google
  • Image generation
  • Image editing
  • Gemini 3.6 Flash
  • 1K · 2K · 4K
  • Up to 14 references
  • Released October 6, 2026

Last verified October 7, 2026 against Google, Google DeepMind and Arena.Sources

Three framed interpretations of a banana: a photograph, amber glass and folded yellow paper in a dark studio.

Editorial concept

One subject, three visual treatments

Photography, materials and illustration offer different ways to develop the same creative idea.

Editorial concept illustrations created with OpenAI image generation.

Nano Banana 2.1 Quick Facts

Developer
Google
Model ID
Type
Image generation & editing
Base model
Gemini 3.6 Flash
Release
October 6, 2026
Inputs
Text, Image, Video, PDF
Outputs
Image, Text
Resolution
1K (default) · 2K · 4K
Aspect ratios
14 ratios, 1:4, 4:1, 1:8, 8:1 included
Reference images
Up to 14 (10 objects + 4 characters)
Thinking
Minimal · Medium (default) · High
Token limits (API)
131,072 input · 32,768 output
Search grounding
Supported
Batch API
Supported
Image output
From $0.0336 per 1K image

Why Nano Banana 2.1 Matters

01

Better editing

In Google's editing evaluations the thinking variant scores 1106 for multi-character consistency (Nano Banana 2: 978), 1049 for mask-based editing (965) and 1066 for multi-reference editing (988).

02

Half-price 1K and 2K images

Image output drops to $0.0336 at 1K and $0.0504 at 2K, exactly half of Nano Banana 2. At 4K the cut is smaller: $0.113 versus $0.151.

03

Visual reasoning

Configurable thinking and Google Search grounding target infographics and fact-dependent visuals. Google's infographic factuality score rises from 0.179 (Nano Banana 2) to 0.521.

What Nano Banana 2.1 Can Do

A yellow camera with a black lens and red button shown in a neutral studio, a sunlit setting and a dark studio.

Editorial concept

Product identity across scenes

A visual concept for keeping a product recognizable while changing its setting, lighting and mood.

A cream poster reading Make Something Bright combines bold black typography, a yellow circle and amber glass, with a matching square print.

Editorial concept

Typography meets composition

A poster concept combining a short headline, graphic shapes and photographic material textures.

Text-to-image

Generate 1K, 2K or 4K images from a prompt in any of 14 aspect ratios.

Conversational editing

Edit an image over several turns, for example translating an infographic while keeping every other element unchanged.

Multi-reference composition

Combine up to 14 reference images in one request: up to 10 objects and up to 4 characters.

Character and product consistency

Keep up to 4 characters and up to 10 products or objects consistent across new scenes.

Typography and infographics

Render legible, stylized text for infographics, menus, diagrams and marketing assets. Small and long text remain weak spots.

Search-grounded generation

Use Google Search as a tool so images such as weather charts or event graphics reflect current information.

Panoramic formats

Use 1:4, 4:1, 1:8 and 8:1 ratios; Google fixed tiling artifacts at 2K and 4K in this release.

  • Image generationsupported
  • Image editing (multi-turn)supported
  • Thinkingsupported
  • Google Search groundingsupported
  • Batch APIsupported
  • Context cachingnot supported
  • Structured outputsnot supported
  • Function callingnot supported
  • Code executionnot supported

Nano Banana 2.1 Pricing

Image output is billed at $30.00 per million image tokens; a 1K image uses 1,120 tokens. Batch requests cost half. There is no free tier for this model in the Gemini API.

Nano Banana 2.1 price per image
ResolutionImage tokensStandardBatch
1K1,120$0.0336$0.0168
2K1,680$0.0504$0.0252
4K3,780$0.113$0.0567

Cost per 1,000 images

Nano Banana 2.1 cost per 1,000 images
ResolutionStandardBatch
1K$33.60$16.80
2K$50.40$25.20
4K$113.40$56.70

Calculated from image tokens × the per-token rate, so 4K shows the unrounded $0.1134 per image behind Google's $0.113.

What actually costs money?

Image output

$30.00 per 1M image tokens ($15.00 batch). Usually the largest line item.

Input tokens

$1.50 per 1M tokens for text, images and video ($0.75 batch). Reference images count as input.

Text and thinking output

$7.50 per 1M tokens ($3.75 batch). Higher thinking levels produce more of it; interim thought images are not charged.

Search grounding

5,000 free search requests per month (shared across Gemini 3.x models), then $14 per 1,000 requests.

$0.0336 is the image output alone. Prompts, reference images, thinking tokens, multi-turn edits and grounding requests are billed on top, so a real request usually costs more.

Nano Banana 2.1 / Nano Banana 2 / Nano Banana Pro pricing comparison

Standard price per image
Model1K2K4K
Nano Banana 2.1$0.0336$0.0504$0.113
Nano Banana 2$0.067$0.101$0.151
Nano Banana Pro$0.134$0.134$0.24

At 1K, a Nano Banana 2.1 image costs about a quarter of Nano Banana Pro ($0.0336 vs $0.134). At 4K the gap narrows to $0.113 vs $0.24. Batch prices for all three models are about half of these.

Nano Banana 2.1 Benchmarks

Vendor results from the Google DeepMind model card. Elo from side-by-side human evaluations; infographic factuality from a single-sided autorater.

Google benchmarks: Text-to-image

Text-to-image benchmarks
TestNano Banana 2.1 (thinking)Nano Banana 2.1 (no thinking)Nano Banana 2 (thinking)Nano Banana Pro
Overall preference1050 ± 141015 ± 13990 ± 7935 ± 8
Infographic design1048 ± 171001 ± 17961 ± 12912 ± 12
Infographic factuality0.5210.3280.1790.265

Google benchmarks: Image editing

Image editing benchmarks
TestNano Banana 2.1 (thinking)Nano Banana 2.1 (no thinking)Nano Banana 2 (thinking)Nano Banana Pro
General editing1026 ± 12980 ± 15938 ± 11939 ± 10
Single-character consistency1028 ± 141021 ± 14981 ± 10991 ± 9
Multi-character consistency1106 ± 141068 ± 14978 ± 101011 ± 10
Mask / ink-based editing1049 ± 151042 ± 16965 ± 12927 ± 12
Product consistency1024 ± 18981 ± 18955 ± 22965 ± 14
Stylization1062 ± 201036 ± 17991 ± 12990 ± 12
Multi-reference editing1066 ± 221041 ± 20988 ± 13989 ± 12
Source: Google DeepMind model card (vendor benchmark)

Independent benchmarks

Blind preference leaderboards and vendor evaluations measure different things, so the numbers are not interchangeable. Early Arena scores are marked preliminary and will move as votes come in.

Arena leaderboard positions
LeaderboardRankScoreVotesStatus
Arena text-to-image#51328 ± 95,312Preliminary, October 6, 2026
Arena image edit#61428 ± 612,985Preliminary, October 6, 2026

Nano Banana 2.1 Thinking Levels

Set the level with generation_config.thinking_level. Guidance below is ours, not Google's.

Fastest

minimal

Best for

  • Simple generations
  • Background replacement
  • Thumbnails
  • Bulk workflows

Default

medium

Best for

  • Everyday production
  • Ads and marketing images
  • Multi-element scenes
  • Routine edits

Most reasoning

high

Best for

  • Typography and infographics
  • Many reference images
  • Complex or multi-step edits
  • Search-grounded visuals

Nano Banana 2.1 Resolution and Aspect Ratios

Resolution values must use an uppercase K. Pixel sizes below are for 1:1; other ratios keep the same resolution tier.

Resolutions
Resolution1:1 sizeImage tokens
1K (default)1024 × 10241,120
2K2048 × 20481,680
4K4096 × 40963,780

14 supported aspect ratios

  • 1:1
  • 2:3
  • 3:2
  • 3:4
  • 4:3
  • 4:5
  • 5:4
  • 9:16
  • 16:9
  • 21:9
  • 1:4
  • 4:1
  • 1:8
  • 8:1

Highlighted panoramic ratios (1:4, 4:1, 1:8, 8:1) had tiling artifacts fixed at 2K and 4K in this release.

Nano Banana 2.1 Reference Image Limits

Up to 14 reference images per request.

14 references does not mean 14 people: character consistency covers up to 4 characters, and up to 10 images can be objects or products.

Nano Banana 2.1 Prompt Structure

A practical anatomy for production prompts: say what must not change before describing what should.

  1. PreservationKeep the product geometry, logo and packaging unchanged.
  2. ScenePlace the product on a polished black stone surface in a minimalist luxury studio.
  3. LightingUse soft directional lighting with subtle reflections.
  4. CompositionCreate a centered commercial composition.
  5. TypographyAdd the headline “DESIGNED FOR TOMORROW” in the upper third.
  6. Output16:9 cinematic advertising image.
Keep the product geometry, logo and packaging unchanged.

Place the product on a polished black stone surface in a minimalist luxury studio.

Use soft directional lighting with subtle reflections.

Create a centered commercial composition.

Add the headline “DESIGNED FOR TOMORROW” in the upper third.

16:9 cinematic advertising image.

Nano Banana 2.1 API

Model ID Examples are Google's, unchanged, using the Interactions API.

Generate an image

python
from google import genai
from PIL import Image
import base64

client = genai.Client()

interaction = client.interactions.create(
    model="gemini-nano-banana-2.1",
    input="Create a picture of a nano banana dish in a fancy restaurant with a Gemini theme",
)

with open("generated_image.png", "wb") as f:
    f.write(base64.b64decode(interaction.output_image.data))

Edit an image

python
from google import genai
import base64

client = genai.Client()

with open("/path/to/cat_image.png", "rb") as f:
    image_bytes = f.read()

interaction = client.interactions.create(
    model="gemini-nano-banana-2.1",
    input=[
        {
          "type": "text",
          "text": "Create a picture of a nano banana dish in a fancy restaurant with a Gemini theme"
        },
        {
            "type": "image",
            "data": base64.b64encode(image_bytes).decode('utf-8'),
            "mime_type": "image/png"
        }
    ],
)

with open("generated_image.png", "wb") as f:
    f.write(base64.b64decode(interaction.output_image.data))

Aspect ratio and resolution

python
from google import genai
import base64

client = genai.Client()

prompt = "Da Vinci style anatomical sketch of a dissected Monarch butterfly. Detailed drawings of the head, wings, and legs on textured parchment with notes in English."

interaction = client.interactions.create(
    model="gemini-nano-banana-2.1",
    input=prompt,
    response_format={
        "type": "image",
        "mime_type": "image/jpeg",
        "aspect_ratio": "1:1",
        "image_size": "1K"
    },
)

print(interaction.output_text)

with open("butterfly.png", "wb") as f:
    f.write(base64.b64decode(interaction.output_image.data))

Set the thinking level

python
from google import genai
from PIL import Image
import base64
import io

client = genai.Client()

interaction = client.interactions.create(
    model="gemini-nano-banana-2.1",
    input="A futuristic city built inside a giant glass bottle floating in space",
    generation_config={"thinking_level": "high"},
)

print(interaction.output_text)

image = Image.open(io.BytesIO(base64.b64decode(interaction.output_image.data)))

image.show()

Multiple reference images

python
from google import genai
from google.genai import types
from PIL import Image
import base64

prompt = "An office group photo of these people, they are making funny faces."

client = genai.Client()

interaction = client.interactions.create(
    model="gemini-nano-banana-2.1",
    input=[
        {
            "type": "text",
            "text": prompt,
        },
        {
            "type": "image",
            "data": base64.b64encode(image_bytes).decode('utf-8'),
            "mime_type": "image/png"
        },
        {
            "type": "image",
            "data": base64.b64encode(image_bytes).decode('utf-8'),
            "mime_type": "image/png"
        },
        {
            "type": "image",
            "data": base64.b64encode(image_bytes).decode('utf-8'),
            "mime_type": "image/png"
        },
        {
            "type": "image",
            "data": base64.b64encode(image_bytes).decode('utf-8'),
            "mime_type": "image/png"
        },
        {
            "type": "image",
            "data": base64.b64encode(image_bytes).decode('utf-8'),
            "mime_type": "image/png"
        },
    ],
    response_format={
        "type": "image",
        "aspect_ratio": "5:4",
        "image_size": "2K"
    },
)

with open("office.png", "wb") as f:
    f.write(base64.b64decode(interaction.output_image.data))

Google Search grounding

python
from google import genai
import base64

client = genai.Client()

prompt = "Visualize the current weather forecast for the next 5 days in San Francisco as a clean, modern weather chart. Add a visual on what I should wear each day"

interaction = client.interactions.create(
    model="gemini-nano-banana-2.1",
    input=prompt,
    tools=[{"type": "google_search"}],
    response_format={
        "type": "image",
        "mime_type": "image/jpeg",
        "aspect_ratio": "16:9"
    },
)

with open("weather.png", "wb") as f:
    f.write(base64.b64decode(interaction.output_image.data))
View the official API documentation

Developer notes

Use the Interactions API

Google's examples use interactions.create in the google-genai (Python) and @google/genai (JavaScript) SDKs; the generateContent image guide is now marked legacy.

Follow the endpoint token limits

The API lists 131,072 input and 32,768 output tokens. The model card describes the Gemini 3.6 Flash base model with up to 1M input tokens; plan around the endpoint limits.

No caching, structured output or function calling

Context caching, structured outputs, function calling and code execution are not supported, so this model cannot be swapped into every Gemini text workflow.

Uppercase resolution values

image_size must be 1K, 2K or 4K with an uppercase K; lowercase values are rejected. 1K is the default.

Thought images are free, thinking text is not

Interim thought images generated while thinking are not charged, but thinking output tokens are billed at the text/thinking rate.

Shared grounding quota

The 5,000 free monthly Search grounding requests are shared across all Gemini 3.x models in a project.

Nano Banana 2.1 Limitations

  • Small text can be blurry, especially at 1K.
  • Long paragraphs and page-length text remain unreliable.
  • Character consistency between input and output images is not always perfect.
  • Mask and doodle-based edits can follow instructions only partially or leave ink behind.
  • Spatial instructions such as left and right can be confused.
  • Infographic facts still need checking: Google's own infographic factuality score is 0.521.
  • Every generated image carries a SynthID watermark.

Known issues come from Google's model card and API documentation.

Nano Banana 2.1 vs Nano Banana Pro

A quick comparison using Google's prices and evaluations. Full head-to-head guides will follow.

Model comparison
Nano Banana 2.1Nano Banana 2Nano Banana Pro
1K image$0.0336$0.067$0.134
4K image$0.113$0.151$0.24
Google overall preference (Elo)1050 ± 14990 ± 7935 ± 8
Model IDgemini-nano-banana-2.1gemini-3.1-flash-imagegemini-3-pro-image

Against Nano Banana 2, Nano Banana 2.1 is cheaper at every resolution and ahead in every Google text-to-image and editing evaluation above.

Against Nano Banana Pro, Nano Banana 2.1 is cheaper at every resolution and ahead in every Google text-to-image and editing evaluation above.

Nano Banana 2.1 FAQ

What is Nano Banana 2.1?

Nano Banana 2.1 is Google's high-efficiency image generation and editing model. It is based on Gemini 3.6 Flash and accepts text, image, video, pdf input, returning images and text.

When was Nano Banana 2.1 released?

Google released Nano Banana 2.1 on October 6, 2026.

What is the Nano Banana 2.1 model ID?

Use gemini-nano-banana-2.1 in the Gemini API.

How much does Nano Banana 2.1 cost?

Image output costs $0.0336 per 1K image, $0.0504 per 2K image and $0.113 per 4K image. Batch halves this to $0.0168, $0.0252 and $0.0567. Input tokens cost $1.50 and text/thinking output $7.50 per million tokens.

Is Nano Banana 2.1 cheaper than Nano Banana Pro?

Yes. A 1K image costs $0.0336 versus $0.134 for Nano Banana Pro, and a 4K image $0.113 versus $0.24.

How does Nano Banana 2.1 pricing compare with Nano Banana 2?

1K and 2K images cost about half as much as Nano Banana 2 ($0.067 and $0.101). At 4K the price falls from $0.151 to $0.113.

Does Nano Banana 2.1 support 4K?

Yes. It outputs 1K / 2K / 4K; 1K is the default. Values must use an uppercase K.

How many reference images does Nano Banana 2.1 support?

Up to 14: up to 10 object images and up to 4 character images.

What are the Nano Banana 2.1 thinking levels?

minimal / medium / high. The default is medium; set it with generation_config.thinking_level.

Does Nano Banana 2.1 support image editing?

Yes, including multi-turn conversational editing, mask-based editing and multi-reference editing.

Is there a free tier for the Nano Banana 2.1 API?

No. Google's pricing page lists no free tier for this model.

What does Nano Banana 2.1 not support?

The Gemini API lists context caching, structured outputs, function calling, code execution as not supported.

Do Nano Banana 2.1 images include a watermark?

Yes. Every generated image includes a SynthID watermark.

Sources and verification

Last verified October 7, 2026. Prices, limits and leaderboard positions can change; check the sources before committing budgets.

Browse all models