01
编辑能力提升
在 Google 的编辑评测中,开启思考的版本在多人物一致性上得分 1106(Nano Banana 2 为 978),蒙版编辑为 1049(前代为 965),多参考图编辑为 1066(前代为 988)。
Google · 图像生成与编辑
Google 的高效图像生成与编辑模型
支持生成和编辑 1K–4K 图像,最多可结合 14 张参考图,保持人物与产品的一致性,并生成清晰可读的文字。面对复杂的视觉任务,还可以调整思考级别。
最近核实:2026年10月7日,依据 Google、Google DeepMind 和 Arena 的资料。来源
01
在 Google 的编辑评测中,开启思考的版本在多人物一致性上得分 1106(Nano Banana 2 为 978),蒙版编辑为 1049(前代为 965),多参考图编辑为 1066(前代为 988)。
02
1K 图像输出降至每张 0.0336 美元,2K 为 0.0504 美元,约为 Nano Banana 2 的一半。4K 的降幅较小,从 0.151 美元降至 0.113 美元。
03
可调节的思考级别与 Google 搜索信息引用,适合制作信息图和依赖事实的视觉内容。Google 的信息图事实准确性得分从 Nano Banana 2 的 0.179 提升至 0.521。
通过提示词生成 1K、2K 或 4K 图像,支持 14 种宽高比。
通过多轮对话逐步编辑图像,例如在保留其他元素的同时,翻译信息图中的文字。
单次请求最多可结合 14 张参考图,其中物体参考图最多 10 张,人物参考图最多 4 张。
更换场景时,最多可保持 4 个人物与 10 个产品或物体的一致性。
为信息图、菜单、图表和营销素材生成清晰且有风格的文字,但小字号和长文本仍是薄弱环节。
调用 Google 搜索,让天气图表、活动配图等内容反映最新信息。
支持 1:4、4:1、1:8 和 8:1 画幅;Google 在本次更新中修复了这些比例在 2K、4K 下出现的拼接伪影。
图像输出按每百万图像 Token $30.00 计费;一张 1K 图像消耗 1,120 个 Token。批量请求的费率减半。 该模型在 Gemini API 中没有免费额度。
| 分辨率 | 图像 Token 数 | 标准 | 批量 |
|---|---|---|---|
| 1K | 1,120 | $0.0336 | $0.0168 |
| 2K | 1,680 | $0.0504 | $0.0252 |
| 4K | 3,780 | $0.113 | $0.0567 |
| 分辨率 | 标准 | 批量 |
|---|---|---|
| 1K | $33.60 | $16.80 |
| 2K | $50.40 | $25.20 |
| 4K | $113.40 | $56.70 |
按图像 Token 数乘以 Token 单价计算,因此 4K 显示的是每张 $0.1134 的未舍入费用;Google 公布的舍入价格为 $0.113。
每百万图像 Token $30.00,批量请求为 $15.00。这通常是费用中的主要部分。
文本、图像和视频输入按每百万 Token $1.50 计费,批量请求为 $0.75。参考图也计入输入。
每百万 Token $7.50,批量请求为 $3.75。思考级别越高,通常输出的思考 Token 越多;思考过程中的中间图像不收费。
每月 5,000 次免费搜索请求,由 Gemini 3.x 系列模型共享;超出后每 1,000 次收费 14 美元。
$0.0336 仅为图像输出费用。提示词、参考图、思考 Token、多轮编辑和搜索信息引用请求会另外计费,因此实际请求的总费用通常更高。
| 模型 | 1K | 2K | 4K |
|---|---|---|---|
| Nano Banana 2.1 | $0.0336 | $0.0504 | $0.113 |
| Nano Banana 2 | $0.067 | $0.101 | $0.151 |
| Nano Banana Pro | $0.134 | $0.134 | $0.24 |
在 1K 下,Nano Banana 2.1 每张图像的价格约为 Nano Banana Pro 的四分之一($0.0336 对比 $0.134)。4K 的价差缩小至 $0.113 对比 $0.24。这三款模型的批量价格均约为标准价格的一半。
以下为厂商评测结果,来源:Google DeepMind 模型卡。Elo 来自人工并排对比评测;信息图的事实准确性由自动评估器对单个输出评分。
| 评测项目 | Nano Banana 2.1(开启思考) | Nano Banana 2.1(关闭思考) | Nano Banana 2(开启思考) | Nano Banana Pro |
|---|---|---|---|---|
| 整体偏好 | 1050 ± 14 | 1015 ± 13 | 990 ± 7 | 935 ± 8 |
| 信息图设计 | 1048 ± 17 | 1001 ± 17 | 961 ± 12 | 912 ± 12 |
| 信息图事实准确性 | 0.521 | 0.328 | 0.179 | 0.265 |
| 评测项目 | Nano Banana 2.1(开启思考) | Nano Banana 2.1(关闭思考) | Nano Banana 2(开启思考) | Nano Banana Pro |
|---|---|---|---|---|
| 常规编辑 | 1026 ± 12 | 980 ± 15 | 938 ± 11 | 939 ± 10 |
| 单人物一致性 | 1028 ± 14 | 1021 ± 14 | 981 ± 10 | 991 ± 9 |
| 多人物一致性 | 1106 ± 14 | 1068 ± 14 | 978 ± 10 | 1011 ± 10 |
| 蒙版/涂鸦编辑 | 1049 ± 15 | 1042 ± 16 | 965 ± 12 | 927 ± 12 |
| 产品一致性 | 1024 ± 18 | 981 ± 18 | 955 ± 22 | 965 ± 14 |
| 风格化 | 1062 ± 20 | 1036 ± 17 | 991 ± 12 | 990 ± 12 |
| 多参考图编辑 | 1066 ± 22 | 1041 ± 20 | 988 ± 13 | 989 ± 12 |
盲测偏好榜单与厂商评测衡量的维度不同,分数不能直接对照。Arena 的早期得分标记为初步结果,会随投票增加而变化。
| 榜单 | 排名 | 得分 | 投票数 | 状态 |
|---|---|---|---|---|
| Arena 文生图榜单 | #5 | 1328 ± 9 | 5,312 | 初步结果, 2026年10月6日 |
| Arena 图像编辑榜单 | #6 | 1428 ± 6 | 12,985 | 初步结果, 2026年10月6日 |
通过 generation_config.thinking_level 设置思考级别。以下为本站的使用建议,并非 Google 官方建议。
速度最快
适合场景
默认
适合场景
更充分的推理
适合场景
分辨率参数中的 K 必须大写。下表列出的是 1:1 画幅的像素尺寸;其他宽高比沿用同一分辨率档位。
| 分辨率 | 1:1 像素尺寸 | 图像 Token 数 |
|---|---|---|
| 1K(默认) | 1024 × 1024 | 1,120 |
| 2K | 2048 × 2048 | 1,680 |
| 4K | 4096 × 4096 | 3,780 |
高亮显示的超宽/超长画幅(1:4, 4:1, 1:8, 8:1)在 2K 和 4K 下的拼接伪影已于本次更新中修复。
单次请求最多使用 14 张参考图。
14 张参考图并不代表支持 14 个人物:人物一致性最多覆盖 4 人,物体或产品参考图最多为 10 张。
实用的提示词写法:先说明哪些内容必须保留,再描述希望做出的变化。
保持产品的形状、标志和包装不变。
将产品放在简约高端影棚中的抛光黑色石面上。
使用柔和的定向光,并带有细腻的反射。
采用主体居中的商业广告构图。
在画面上方三分之一区域添加标题“DESIGNED FOR TOMORROW”。
输出 16:9、具有电影质感的广告图。模型 ID 以下保留 Google 原始示例,使用 Interactions API。
from google import genai
from PIL import Image
import base64
client = genai.Client()
interaction = client.interactions.create(
model="gemini-nano-banana-2.1",
input="Create a picture of a nano banana dish in a fancy restaurant with a Gemini theme",
)
with open("generated_image.png", "wb") as f:
f.write(base64.b64decode(interaction.output_image.data))from google import genai
import base64
client = genai.Client()
with open("/path/to/cat_image.png", "rb") as f:
image_bytes = f.read()
interaction = client.interactions.create(
model="gemini-nano-banana-2.1",
input=[
{
"type": "text",
"text": "Create a picture of a nano banana dish in a fancy restaurant with a Gemini theme"
},
{
"type": "image",
"data": base64.b64encode(image_bytes).decode('utf-8'),
"mime_type": "image/png"
}
],
)
with open("generated_image.png", "wb") as f:
f.write(base64.b64decode(interaction.output_image.data))from google import genai
import base64
client = genai.Client()
prompt = "Da Vinci style anatomical sketch of a dissected Monarch butterfly. Detailed drawings of the head, wings, and legs on textured parchment with notes in English."
interaction = client.interactions.create(
model="gemini-nano-banana-2.1",
input=prompt,
response_format={
"type": "image",
"mime_type": "image/jpeg",
"aspect_ratio": "1:1",
"image_size": "1K"
},
)
print(interaction.output_text)
with open("butterfly.png", "wb") as f:
f.write(base64.b64decode(interaction.output_image.data))from google import genai
from PIL import Image
import base64
import io
client = genai.Client()
interaction = client.interactions.create(
model="gemini-nano-banana-2.1",
input="A futuristic city built inside a giant glass bottle floating in space",
generation_config={"thinking_level": "high"},
)
print(interaction.output_text)
image = Image.open(io.BytesIO(base64.b64decode(interaction.output_image.data)))
image.show()from google import genai
from google.genai import types
from PIL import Image
import base64
prompt = "An office group photo of these people, they are making funny faces."
client = genai.Client()
interaction = client.interactions.create(
model="gemini-nano-banana-2.1",
input=[
{
"type": "text",
"text": prompt,
},
{
"type": "image",
"data": base64.b64encode(image_bytes).decode('utf-8'),
"mime_type": "image/png"
},
{
"type": "image",
"data": base64.b64encode(image_bytes).decode('utf-8'),
"mime_type": "image/png"
},
{
"type": "image",
"data": base64.b64encode(image_bytes).decode('utf-8'),
"mime_type": "image/png"
},
{
"type": "image",
"data": base64.b64encode(image_bytes).decode('utf-8'),
"mime_type": "image/png"
},
{
"type": "image",
"data": base64.b64encode(image_bytes).decode('utf-8'),
"mime_type": "image/png"
},
],
response_format={
"type": "image",
"aspect_ratio": "5:4",
"image_size": "2K"
},
)
with open("office.png", "wb") as f:
f.write(base64.b64decode(interaction.output_image.data))from google import genai
import base64
client = genai.Client()
prompt = "Visualize the current weather forecast for the next 5 days in San Francisco as a clean, modern weather chart. Add a visual on what I should wear each day"
interaction = client.interactions.create(
model="gemini-nano-banana-2.1",
input=prompt,
tools=[{"type": "google_search"}],
response_format={
"type": "image",
"mime_type": "image/jpeg",
"aspect_ratio": "16:9"
},
)
with open("weather.png", "wb") as f:
f.write(base64.b64decode(interaction.output_image.data))Google 示例使用 google-genai(Python)和 @google/genai(JavaScript)SDK 中的 interactions.create;generateContent 图像指南目前已标记为旧版。
API 标注的上限为 131,072 个输入 Token 和 32,768 个输出 Token。模型卡中 Gemini 3.6 Flash 基础模型的输入上限虽为 100 万 Token,实际接入仍应遵循端点限制。
不支持上下文缓存、结构化输出、函数调用和代码执行,因此不能直接替换所有 Gemini 文本工作流中的模型。
image_size 必须填写 1K、2K 或 4K,K 需大写;小写值会被拒绝。默认值为 1K。
思考过程中生成的中间图像不收费,但思考输出的 Token 会按文本/思考输出费率计费。
每月 5,000 次免费搜索信息引用请求,由同一项目中的所有 Gemini 3.x 模型共享。
已知问题依据 Google 模型卡与 API 文档整理。
依据 Google 的价格与评测结果做简要对比,后续将补充详细对比指南。
| Nano Banana 2.1 | Nano Banana 2 | Nano Banana Pro | |
|---|---|---|---|
| 1K 图像 | $0.0336 | $0.067 | $0.134 |
| 4K 图像 | $0.113 | $0.151 | $0.24 |
| Google 整体偏好(Elo) | 1050 ± 14 | 990 ± 7 | 935 ± 8 |
| 模型 ID | gemini-nano-banana-2.1 | gemini-3.1-flash-image | gemini-3-pro-image |
与 Nano Banana 2 相比,Nano Banana 2.1 在所有分辨率下的价格都更低,并在上方 Google 的每项文生图与编辑评测中领先。
与 Nano Banana Pro 相比,Nano Banana 2.1 在所有分辨率下的价格都更低,并在上方 Google 的每项文生图与编辑评测中领先。
Nano Banana 2.1 是 Google 的高效图像生成与编辑模型,基于 Gemini 3.6 Flash,支持文本、图像、视频、pdf输入,可输出图像和文本。
Google 于 2026年10月6日 发布了 Nano Banana 2.1。
在 Gemini API 中使用 gemini-nano-banana-2.1。
图像输出按张计费:1K 为 $0.0336,2K 为 $0.0504,4K 为 $0.113。批量价格约减半,分别为 $0.0168、$0.0252 和 $0.0567。每百万输入 Token 为 $1.50,每百万文本/思考输出 Token 为 $7.50。
是的。1K 图像每张为 $0.0336,Nano Banana Pro 为 $0.134;4K 图像每张为 $0.113,对方为 $0.24。
1K 和 2K 图像的价格约为 Nano Banana 2(分别为 $0.067 和 $0.101)的一半;4K 则从 $0.151 降至 $0.113。
支持。可输出 1K / 2K / 4K 图像,默认为 1K。参数中的 K 必须大写。
最多 14 张,其中物体参考图最多 10 张,人物参考图最多 4 张。
可选 minimal / medium / high,默认为 medium,通过 generation_config.thinking_level 设置。
支持,包括多轮对话式编辑、蒙版编辑和多参考图编辑。
没有,Google 价格页未提供该模型的免费额度。
Gemini API 列明不支持以下功能:上下文缓存、结构化输出、函数调用、代码执行。
有,每张生成图像都带有 SynthID 水印。
最近核实于 2026年10月7日。价格、使用限制和榜单排名可能变化,请在确定预算前查看原始来源。