AI Image Generator Text Rendering Benchmark 2026: 11 Models
AI image generator text rendering benchmark: 11 models, 5 e-commerce posters, 110 images. Exact-text pass rates for English, Chinese, SKU and prices, plus cost.

In our benchmark, one prompt asked every model for a Chinese-language rice cooker poster: a “50% off” headline, a “smart rice cooker” subtitle and a ¥199 price. FLUX.2 Max, FLUX.2 Pro and Ideogram 3 got the headline and the price right in all six of their images. In every one of those six, the last character of the subtitle (the one meaning “pot”) came out wrong: a glyph that isn’t a real character on the FLUX images, a different character on the Ideogram ones. A Chinese shopper would read a typo. Our first OCR judge didn’t notice it on three of the six.
That is the pattern across the AI image generator text rendering benchmark. We gave 11 text-to-image models on SandBase the same five e-commerce poster prompts, two images each, on 2026-10-03 (UTC), and checked every required string. English headlines and prices were close to solved. Chinese glyphs, small print and price layout weren’t, and a vision-LLM grader needed a human check on exactly those cases.
Key takeaway
- Seven of 11 models put every required string on all 10 posters: GPT Image 2, GPT Image 2.5 Flare, Nano Banana Pro, Nano Banana 2, Seedream 5.0 Pro, Qwen Image 3 and Wan 2.7 Pro.
- FLUX.2 Max and Pro scored 6/10 and Ideogram 3 scored 3/10. Every one of their Chinese posters failed on a malformed glyph. Midjourney v8.1 scored 5/10, losing the mixed list, the SKU hyphens and one price label.
- The cheapest 10/10 models were Nano Banana 2 ($0.039 billed per image, 18.7 s median) and Qwen Image 3 ($0.04, 52.7 s). GPT Image 2 also went 10/10, but at $0.15 and a 107 s median.
- The GPT-6.1 Sol OCR judge agreed with our eye on 75 of 83 reviewed images. 13 of the 14 strings it wrongly passed were Chinese text it silently “fixed”. Have a person check CJK text.
- Scope: 5 prompts, 2 images per model per prompt, default settings, one day. Not a general image-quality ranking.
Benchmark setup: 11 models
The scenario: a marketing agent writes the copy and needs a poster carrying it exactly, with nobody retouching it. Every string had to appear character for character.
All 11 models were enabled: true in GET https://api.sandbase.ai/v1/models/<id> on 2026-10-03 (we rechecked Qwen Image 3 at 13:54 UTC: still enabled: true). That endpoint needs the same Authorization: Bearer key as the calls; the public model pages show the same prices without one. We submitted every image through the same Unified Run endpoint, POST https://api.sandbase.ai/v1/run. Ten models answered 202 with a pending run, which we polled with GET /v1/run/{id}; GPT Image 2.5 Flare answered 200 with the finished image, so there was nothing to poll. Everything was left at its default except a 1:1 aspect ratio. Qwen Image 3 and Midjourney v8.1 have no aspect parameter, and their defaults were already square. Output was 1024×1024 for most models, 1328×1328 for Wan 2.7 Pro and 2048×2048 for Seedream 5.0 Pro. Generation ran from 10:34 to about 11:00 UTC.
| Model (SandBase id) | List price / image | Price shown on model page, 2026-10-03 |
|---|---|---|
openai/gpt-image-2 | $0.20 (quality high, the default) | $0.15 (25% off) |
openai/gpt-image-2.5-flare | $0.20 (quality high) | $0.15 (25% off) |
google/nano-banana-pro | $0.12 (1K) | $0.08 (35% off) |
google/nano-banana-2 | $0.06 (1K) | $0.04 (35% off) |
bytedance/seedream/5.0/pro | $0.0675 (with aspect_ratio set) | not captured |
bfl/flux-2/max | $0.07 | not captured |
bfl/flux-2/pro | $0.03 | not captured |
alibaba/qwen-image-3/text-to-image | $0.04 (1K) | $0.04 |
ideogram-ai/ideogram-v3 | $0.06 (BALANCED, the default) | not captured |
alibaba/wan/2.7/pro | $0.075 | not captured |
midjourney/midjourney-v8.1 | $0.10 per request (4 images) | not captured |
List prices come from each model card’s price_formula with our parameters. None of the 11 was above the $0.30/image cutoff we set, so none was dropped. Midjourney returns four images per request at $0.10. We scored only the first image of each request, so its $0.10 counts as one scored image.

Caption: The Qwen Image 3 model page lists $0.04 per run with async execution, which matches the $0.04 billed on each of its 10 test images (captured 2026-10-03).
Five poster prompts and the grading rule
Each prompt names a product, a background and the exact strings, then ends with the same sentence. 16 required strings in total:
| Prompt | Type | Required strings |
|---|---|---|
| P1 | English headline + price | STAY COLD 24 HOURS, $39.99 |
| P2 | Chinese headline + price | 限时五折 (“50% off”), 智能电饭煲 (“smart rice cooker”), ¥199 |
| P3 | Mixed CN/EN, 4-line list | AirFlow Pro 降噪耳机 (noise-cancelling earbuds), 40dB 主动降噪 (active noise cancelling), 续航 36 小时 (36-hour battery), Bluetooth 6.0, IPX5 防水 (waterproof) |
| P4 | Small print | HYDRA SERUM, SKU: BX-2026-07 |
| P5 | Two prices + date | FLASH SALE, Was $129.00, Now $89.00, Ends 2026.10.03 |
All five prompts, verbatim. Each one ends with the same sentence, “Render the text exactly as written, with no other text anywhere in the image.” The prompts are in code blocks because they contain the literal Chinese strings the models had to render.
P1 (English headline + price):
Square e-commerce product poster for a matte black stainless steel insulated water bottle on a light gray studio background. At the top, a large bold headline that reads exactly: "STAY COLD 24 HOURS". At the bottom right, a price tag that reads exactly: "$39.99". Render the text exactly as written, with no other text anywhere in the image.
P2 (Chinese headline + price):
Square e-commerce promotional poster for a white smart rice cooker on a red festive background. A large Chinese headline that reads exactly: "限时五折". Below it, a smaller Chinese subtitle that reads exactly: "智能电饭煲". A price badge that reads exactly: "¥199". Render the text exactly as written, with no other text anywhere in the image.
P3 (Mixed CN/EN 4-line list):
Square product poster for wireless noise-cancelling earbuds on a dark blue gradient background. A headline that reads exactly: "AirFlow Pro 降噪耳机". Below it, a bulleted list of exactly four lines: "40dB 主动降噪", "续航 36 小时", "Bluetooth 6.0", "IPX5 防水". Render the text exactly as written, with no other text anywhere in the image.
P4 (Small print SKU):
Square product poster for a glass dropper bottle of face serum on a soft beige background. A large headline that reads exactly: "HYDRA SERUM". In small print at the bottom left corner, a product code that reads exactly: "SKU: BX-2026-07". Render the text exactly as written, with no other text anywhere in the image.
P5 (Two prices + date):
Square flash sale poster for a pair of white running shoes on a bright yellow background. A headline that reads exactly: "FLASH SALE". Show the old price "Was $129.00" with a strikethrough line, and the new price "Now $89.00" larger next to it. At the bottom, a line that reads exactly: "Ends 2026.10.03". Render the text exactly as written, with no other text anywhere in the image.
Every request sent only model, prompt and, where the model has one, aspect_ratio: "1:1". The per-image labels, raw transcriptions and scoring records are kept internally; the tables in this article carry every figure we cite.
Grading had two stages. First, a vision LLM, openai/gpt-6.1-sol through POST /v1/chat/completions, transcribed each image. We told it to copy text exactly as rendered, keep malformed characters, write ? for unreadable ones and not correct anything. Then Python scored the transcription, with these normalization rules:
- Unicode NFKC, so full-width ¥ becomes ¥.
- Casefold, except the SKU, which must match case.
- Whitespace and leading bullet glyphs dropped, so a phrase may wrap across lines.
$ ¥ . : -kept, and a digit boundary enforced, so¥199fails inside¥1999.
A poster passes only if all of its strings pass. We ran a second judge, google/gemini-3.8-flash, on all 110 images for comparison.
Then we looked. Two reviewers checked images by eye from downscaled copies and zoomed crops, and their three overlapping images agreed. Where an image was reviewed, the human label is final. The others use the Sol judge, and both judges passed all of them.
| Prompt | Images | Reviewed by eye | Judge only |
|---|---|---|---|
| P1 | 22 | 6 | 16 |
| P2 | 22 | 22 | 0 |
| P3 | 22 | 22 | 0 |
| P4 | 22 | 11 | 11 |
| P5 | 22 | 22 | 0 |
| Total | 110 | 83 | 27 |
Results: 7 of 11 models passed every poster
| Model | Posters passing (10) | Strings (32) | P1 EN | P2 ZH | P3 Mixed | P4 SKU | P5 Prices | Billed / image | Median latency | Cost per passing poster |
|---|---|---|---|---|---|---|---|---|---|---|
| GPT Image 2 | 10 | 32 | 2 | 2 | 2 | 2 | 2 | $0.15 | 106.9 s | $0.15 |
| GPT Image 2.5 Flare | 10 | 32 | 2 | 2 | 2 | 2 | 2 | $0.15 | 20.0 s | $0.15 |
| Nano Banana Pro | 10 | 32 | 2 | 2 | 2 | 2 | 2 | $0.078 | 24.0 s | $0.078 |
| Nano Banana 2 | 10 | 32 | 2 | 2 | 2 | 2 | 2 | $0.039 | 18.7 s | $0.039 |
| Seedream 5.0 Pro | 10 | 32 | 2 | 2 | 2 | 2 | 2 | $0.0675 | 30.3 s | $0.0675 |
| Qwen Image 3 | 10 | 32 | 2 | 2 | 2 | 2 | 2 | $0.04 | 52.7 s | $0.04 |
| Wan 2.7 Pro | 10 | 32 | 2 | 2 | 2 | 2 | 2 | $0.075 | 16.2 s | $0.075 |
| FLUX.2 Max | 6 | 24 | 2 | 0 | 0 | 2 | 2 | $0.07 | 21.4 s | $0.117 |
| FLUX.2 Pro | 6 | 24 | 2 | 0 | 0 | 2 | 2 | $0.03 | 13.5 s | $0.05 |
| Midjourney v8.1 | 5 | 25 | 2 | 2 | 0 | 0 | 1 | $0.10 / request | 44.8 s | $0.20 |
| Ideogram 3 | 3 | 19 | 2 | 0 | 0 | 1 | 0 | $0.06 | 16.5 s | $0.20 |
Billed per image is the median cost from GET /v1/tasks/<id>/cost across all of a model’s tasks. Latency is submit to finished, polling every 2 s from one machine. Run ids were passed to GET /v1/tasks/<id>/cost as returned; the documented source of a task id is the x-task-id response header, which carried the same value as the run id in our calls.
By prompt type, all 22 P1 posters passed. P2 passed 16/22, P3 14/22, P4 19/22 and P5 19/22. Seven strings passed on every image from every model: both P1 strings, the P2 headline and price, the P3 waterproof line (IPX5 plus two Chinese characters), HYDRA SERUM and Ends 2026.10.03.

Caption: Prompt 2 outputs generated on 2026-10-03 by bfl/flux-2/max, ideogram-ai/ideogram-v3, alibaba/qwen-image-3/text-to-image and bytedance/seedream/5.0/pro (first image of each). Only the bottom two spell the “smart rice cooker” subtitle correctly.
Where the text broke
Uncommon Chinese characters. The FLUX.2 pair and Ideogram 3 rendered common characters, like the “50% off” headline and “waterproof”, correctly and broke on rarer ones. The “pot” character failed on all 6 of their P2 images. On P3, the two-character word for “noise cancelling” failed on all 6 of their images, in both the headline and the first bullet, usually with its second character malformed. The “36-hour battery” line failed on 5, with one wrong character or a missing word. Midjourney passed P2 but lost “noise cancelling” on both P3 images, and on one it also misspelled “earbuds”. FLUX.2 Pro dropped the Bluetooth 6.0 line once.

Caption: Prompt 3 outputs generated on 2026-10-03 by bfl/flux-2/pro, midjourney/midjourney-v8.1, openai/gpt-image-2 and alibaba/wan/2.7/pro (first image of each). The Latin parts survive everywhere; the Chinese word for “noise cancelling” doesn’t.
Small print. Midjourney wrote SKU: BX 2026 07 on both P4 images, with no hyphens. Ideogram wrote Bx-2026-07 once, which fails a case-sensitive product code. All other models got the SKU exactly. We confirmed 8 of those 19 passing SKUs by eye in cropped close-ups.
Price layout. Ideogram’s second P5 image hid the “a” of SALE behind the shoe and printed WAS $1299.00 and $89.0. On its first, NOW sat in the top corner, far from $89.00. We failed that one as a detached label; it’s a judgment call. Midjourney rendered Now as “Novr Low” once. Some models also ignored the strikethrough and underlined the old price instead: Midjourney twice and Seedream once among the posters we checked. We didn’t score styling, only text.
Text nobody asked for. Every prompt ended with “no other text anywhere in the image”, and several models added some anyway. In ad copy, these are the risky ones:
- Wan 2.7 Pro added “free shipping”, “genuine product” and “7-day returns” (in Chinese) to one rice cooker poster. Those are service promises the merchant never made.
- One Nano Banana Pro poster carried a faint watermark-style logo, which the Sol judge read as the name of a Chinese stock-image site.
- Ideogram invented brand marks: “Ricce” on a rice cooker, a garbled “D?CL” on the bottle. Wan printed “HYDRA STRUM” on a serum label.
- Swoosh-like logos appeared on shoes from Nano Banana 2, Nano Banana Pro and Ideogram.
Qwen Image 3 was the only model whose 10 transcriptions contained no unrequested text at all. Micro-text on appliance control panels, which most models drew, we treated as cosmetic.
Can a vision LLM grade the posters?
Mostly, if a person checks the CJK. Against the 83 images we reviewed by eye:
| Judge | Image verdicts agreeing | Strings agreeing | Judge pass, human fail | Judge fail, human pass |
|---|---|---|---|---|
| GPT-6.1 Sol | 75/83 | 280/298 | 14 | 4 |
| Gemini 3.8 Flash | 77/83 | 278/298 | 10 | 10 |
The two judges failed differently. 13 of Sol’s 14 lenient errors were Chinese strings. Asked for a verbatim transcription, it still “corrected” the malformed “pot” character on three FLUX.2 images and repaired the malformed “noise cancelling” word in 8 strings on FLUX.2 and Ideogram images. Its strict errors came from reading order: when a model stacked “Now” above “$89.00” or put “Was” and “Now” in a column above both prices, Sol transcribed the labels first and the string match failed.
Gemini 3.8 Flash had fewer lenient errors but more strict ones. On two P3 images it returned its own reasoning about stroke components instead of a transcription.
On Latin text the judges were reliable. Both agreed with us on all 17 P1 and P4 images we reviewed, including the failed Midjourney and Ideogram SKUs, and both passed the 27 we didn’t review. Sol’s Latin errors were the four P5 reading-order cases plus the detached Ideogram NOW. If you automate this, auto-accept Latin-only passes and route CJK passes to a person. The sample program below does exactly that.
Cost: billed prices were below list for four models
The study generated 110 scored images plus 4 billed duplicates. Those were retries after our client lost the connection while the job still finished upstream and was billed. One Midjourney task sat in running for 900 s, later ended failed and wasn’t charged.
| Item | Basis | Cost |
|---|---|---|
| Image generation, 114 billed tasks | Sum of GET /v1/tasks/<id>/cost | $8.96 |
| OCR pass, GPT-6.1 Sol | 192,820 input and 18,384 output tokens at $2 / $10 per million | $0.57 |
| OCR pass, Gemini 3.8 Flash | Usage at $1.50 / $7.50 per million | $0.38 |
| Probes (three models, list price) | One Qwen Image 3, one FLUX.2 Pro, one Midjourney request | $0.17 |
| First sample-program run | Two Qwen images plus two judge calls | $0.09 |
| Screenshot captures | 10 capture calls | $0.05 |
| Total, first day | about $10.22 |
The rerun of the sample program after the polling fix, described below, added $0.09.
For four models the billed cost was below the list price in the model card. The model pages explain why: they showed 25% off for both GPT Image models and 35% off for both Nano Banana models on 2026-10-03. The other seven billed exactly their list price. Discounts change, so treat the billed column as a 2026-10-03 snapshot.

Caption: The Nano Banana 2 page showed 35% off, $0.04 instead of $0.06, on 2026-10-03. Its 10 test images were billed $0.039 each (captured 2026-10-03).

Caption: GPT Image 2.5 Flare showed 25% off, $0.15 instead of $0.20, matching the $0.15 billed per image in this test (captured 2026-10-03).
Cost per passing poster is where the failures show up. FLUX.2 Pro is the cheapest model on the list at $0.03, but at 6/10 it costs $0.05 per usable poster across all five prompts. On English-only work it went 6/6, so there it really is $0.03. Ideogram and Midjourney both land at $0.20 per usable poster.
How to choose
| If you need | Start with | But |
|---|---|---|
| Lowest cost with exact Chinese and English text | Nano Banana 2: 10/10 at $0.039, 18.7 s | It drew both P2 posters as framed mockups on a gray background, not flat artwork; ask for “flat poster, edge to edge” if you need that |
| A low price with no invented copy | Qwen Image 3: 10/10 at $0.04, zero unrequested text | 52.7 s median, the second slowest |
| Fastest 10/10 | Wan 2.7 Pro: 16.2 s at $0.075 | Watch for added claims like “free shipping” or “genuine” |
| A second 10/10 vendor at moderate cost | Seedream 5.0 Pro ($0.0675, 2048 px) or Nano Banana Pro ($0.078) | Nano Banana Pro carried a watermark-like logo once |
| OpenAI stack | GPT Image 2.5 Flare: 10/10 at $0.15, 20 s | GPT Image 2 scored the same at a 107 s median |
| English-only posters, cheapest | FLUX.2 Pro: P1/P4/P5 6/6 at $0.03 | 0/4 on Chinese; don’t use it for CJK copy |
For an agent, the useful pattern is a default model, a fallback, and a check that decides between them. Generate with the default, transcribe, and score the strings in code. If a string fails, regenerate or switch to the fallback, and send CJK passes to a person before anything goes live.
Run the same check on SandBase
One SandBase API key covers the image model and the OCR judge. The program below does one model’s loop:
- Generate a poster with
POST /v1/run. - If the submit returns an unfinished run (
202), pollGET /v1/run/{id}until it’s done, then downloadoutputs[0].url. - Send the image to GPT-6.1 Sol through the OpenAI-compatible Chat Completions endpoint for a verbatim transcription.
- Check the required strings in Python.
It raises when a task doesn’t complete or the judge stops for any reason other than stop, and it marks CJK passes for a human look. It polls only when the submit returns an unfinished run, since some models (GPT Image 2.5 Flare in our test) answer synchronously. We kept GPT-6.1 Sol as the judge because it made fewer strict errors than Gemini 3.8 Flash (4 vs 10) and always returned a transcription. It is also the model our Sonnet 5.5 vs GPT-6.1 Sol agent test found 30/30 on short tool tasks.
"""Generate a product poster on SandBase, transcribe its text with a vision LLM, and check the required strings.
Usage: SANDBASE_API_KEY=... python3 poster_text_check.py
"""
import base64
import os
import re
import time
import unicodedata
import requests
BASE = "https://api.sandbase.ai/v1"
HEADERS = {"Authorization": f"Bearer {os.environ['SANDBASE_API_KEY']}"}
IMAGE_MODEL = "alibaba/qwen-image-3/text-to-image"
JUDGE_MODEL = "openai/gpt-6.1-sol"
PROMPT = ('Square e-commerce promotional poster for a white smart rice cooker on a red festive background. '
'A large Chinese headline that reads exactly: "限时五折". Below it, a smaller Chinese subtitle '
'that reads exactly: "智能电饭煲". A price badge that reads exactly: "¥199". '
'Render the text exactly as written, with no other text anywhere in the image.')
REQUIRED = ["限时五折", "智能电饭煲", "¥199"]
OCR_PROMPT = ("Transcribe every piece of visible text in this image exactly as it is rendered, character by "
"character. Do not correct spelling, do not fix garbled or malformed characters, do not guess what "
"the text was meant to say, and do not translate. Put each separate text line on its own line. "
"If a character is unreadable, write ? for it. If there is no text, reply with NO_TEXT.")
def generate(prompt: str, timeout_s: int = 600) -> tuple[str, bytes]:
"""Submit to POST /v1/run, poll GET /v1/run/{id} while the run is not finished, return (run id, image bytes)."""
resp = requests.post(f"{BASE}/run", headers=HEADERS, json={"model": IMAGE_MODEL, "prompt": prompt}, timeout=120)
resp.raise_for_status()
task = resp.json()
task_id, start = task["id"], time.time()
while task.get("status") not in ("completed", "failed", "timeout", "cancelled"):
if time.time() - start > timeout_s:
raise TimeoutError(f"{task_id} still {task.get('status')} after {timeout_s}s")
time.sleep(3)
task = requests.get(f"{BASE}/run/{task_id}", headers=HEADERS, timeout=60).json()
outputs = task.get("outputs") or []
if task["status"] != "completed" or not outputs:
raise RuntimeError(f"{task_id}: {task['status']} {task.get('error')}")
return task_id, requests.get(outputs[0]["url"], timeout=120).content
def transcribe(image: bytes) -> tuple[str, dict]:
"""Ask the vision model for a verbatim transcription (no correction)."""
data_url = "data:image/png;base64," + base64.b64encode(image).decode()
body = {"model": JUDGE_MODEL, "max_tokens": 2000, "messages": [{"role": "user", "content": [
{"type": "text", "text": OCR_PROMPT},
{"type": "image_url", "image_url": {"url": data_url}}]}]}
resp = requests.post(f"{BASE}/chat/completions", headers=HEADERS, json=body, timeout=300)
resp.raise_for_status()
out = resp.json()
choice = out["choices"][0]
if choice.get("finish_reason") != "stop":
raise RuntimeError(f"judge stopped with {choice.get('finish_reason')}")
return choice["message"].get("content") or "", out.get("usage") or {}
def norm(s: str) -> str:
"""NFKC (full-width ¥ -> ¥), casefold, drop whitespace and leading bullets."""
s = unicodedata.normalize("NFKC", s).casefold()
return re.sub(r"\s+", "", re.sub(r"^[•·●\-*]+", "", s, flags=re.M))
def check(text: str, required: list[str]) -> dict[str, bool]:
"""A string passes when it appears in the transcription, not glued to another digit (¥199 vs ¥1999)."""
blob = norm(text)
hits = {}
for req in required:
pat = re.escape(norm(req))
if req[-1].isdigit():
pat += r"(?![0-9])"
hits[req] = re.search(pat, blob) is not None
return hits
def needs_human_review(required: list[str]) -> bool:
"""In our test the judge silently 'fixed' near-miss Chinese glyphs, so CJK passes get a human look."""
return any("\u4e00" <= ch <= "\u9fff" for req in required for ch in req)
if __name__ == "__main__":
rows = []
for i in range(2):
t0 = time.time()
task_id, image = generate(PROMPT)
seconds = round(time.time() - t0, 1)
path = f"poster_{i + 1}.png"
open(path, "wb").write(image)
text, usage = transcribe(image)
hits = check(text, REQUIRED)
rows.append((path, task_id, seconds, hits, usage))
print(f"--- {path} ({task_id}, {seconds}s)\n{text}\n")
print(f"{'file':14} {'secs':>5} " + " ".join(f"{r:>6}" for r in REQUIRED) + " verdict")
for path, task_id, seconds, hits, usage in rows:
verdict = "PASS" if all(hits.values()) else "FAIL"
if verdict == "PASS" and needs_human_review(REQUIRED):
verdict = "PASS (check CJK by eye)"
print(f"{path:14} {seconds:5} " + " ".join(f"{'ok' if hits[r] else 'MISS':>6}" for r in REQUIRED)
+ f" {verdict}")
print(f"{'':14} judge tokens in/out: {usage.get('prompt_tokens')}/{usage.get('completion_tokens')}")
Tested on 2026-10-03 (UTC). We ran the program above verbatim once, from 13:48:31 to 13:51:02 UTC. Input was the fictional prompt 2 above, with no third-party content. Request: POST https://api.sandbase.ai/v1/run with {"model": "alibaba/qwen-image-3/text-to-image", "prompt": "<prompt 2>"}, which returned 202 with "status": "pending". The program then polled GET https://api.sandbase.ai/v1/run/{id}; the final poll for the first poster returned this complete object:
{"id": "e863b328-8e22-48ff-bf93-e89c7a614bb3", "status": "completed",
"model": "alibaba/qwen-image-3/text-to-image",
"outputs": [{"content_type": "image/png",
"url": "https://media.sandbase.ai/files/e863b328-8e22-48ff-bf93-e89c7a614bb3/0.png"}]}
Program output, with the two transcriptions left out. The second transcription had five extra lines of ? marks and a “10”, which are icons and a display on the cooker’s control panel.
file secs 限时五折 智能电饭煲 ¥199 verdict
poster_1.png 55.4 ok ok ok PASS (check CJK by eye)
judge tokens in/out: 1318/128
poster_2.png 55.7 ok ok ok PASS (check CJK by eye)
judge tokens in/out: 1318/222
We checked both posters by eye: the headline, the subtitle and the price were all correct on both. GET /v1/tasks/<id>/cost returned:
| Task | Id | Billed |
|---|---|---|
| Poster 1, Qwen Image 3 | e863b328-8e22-48ff-bf93-e89c7a614bb3 | $0.040000 |
| Poster 1, GPT-6.1 Sol judge | 84164611-91de-4bfa-a402-6958a7558ffd | $0.004574 |
| Poster 2, Qwen Image 3 | a4b2e276-3e8b-4a6e-813c-ccf0cc9324c8 | $0.040000 |
| Poster 2, GPT-6.1 Sol judge | c0838a38-8147-4f8a-a715-4ffe8b362d62 | $0.005514 |
The fields used (id, status, outputs[0].url, usage) match the documented asynchronous run response. The request fields are on the Qwen Image 3 model page, and other image models are in the SandBase model catalog. Get a SandBase API key to run the same check on your own catalog copy.
Scope of the data: all prompts describe fictional products, and every image is our own generated output. The prompts, required strings, grading rule and every aggregate are in this article; the per-image labels, raw transcriptions and scoring records are kept internally. For a feature-level view of these vendors, see best AI image generation APIs. For a deeper look at three of the 10/10 models, see Seedream vs Qwen Image vs Nano Banana.
FAQ
Which AI image generator renders text most accurately in 2026?
On our five e-commerce posters, seven models tied with every required string correct on all 10 images: GPT Image 2, GPT Image 2.5 Flare, Nano Banana Pro, Nano Banana 2, Seedream 5.0 Pro, Qwen Image 3 and Wan 2.7 Pro. That’s 10 images per model, so treat it as a shortlist, not a ranking among the seven.
Which image model is best for Chinese text on posters?
Nano Banana 2 and Qwen Image 3 were the cheapest of the seven models that got all four of their Chinese-bearing posters (P2 and P3) right in our test ($0.039 and $0.04 per image). FLUX.2 Max, FLUX.2 Pro and Ideogram 3 failed every Chinese poster. Midjourney v8.1 passed the simple Chinese poster and failed the mixed list.
Can FLUX.2 render text?
Yes, in English: FLUX.2 Pro and Max each passed all 6 of their English-only posters (P1, P4, P5), including the small SKU. Each failed all 4 of its posters with uncommon Chinese characters, so don’t use them for CJK copy.
Can I use an LLM to check text in generated images automatically?
For Latin text, mostly yes. Our GPT-6.1 Sol judge agreed with human review on 75 of 83 images, and both judges agreed with us on all 17 English-only P1 and P4 images we reviewed. For Chinese, no: the judge “corrected” malformed glyphs on its own, so route CJK passes to a person.
How much does a text-accurate AI poster cost?
Among models that passed all 10 posters, billed cost per image on 2026-10-03 ran from $0.039 (Nano Banana 2) and $0.04 (Qwen Image 3) to $0.15 (both GPT Image models). Add about half a cent per image for the OCR check ($0.0046 to $0.0055 billed in our sample run). Four models showed 25 to 35% discounts that day, so check the model page.
Limitations
This test used five prompts and two images per model per prompt, 110 images in total, on one day with default settings, square output and English-language prompts. Two images can’t separate models with a 90% and a 100% pass rate, so read 10/10 as “no failure seen”, not a guarantee. We scored text, not design quality, product fidelity or layout beyond the strings. The strikethrough and “no other text” instructions were noted but not scored. Human review covered 83 of 110 images. The remaining 27 rely on the OCR judge, which was reliable on Latin text in our sample. Latency was measured from one machine through the SandBase gateway, and prices and discounts are a 2026-10-03 snapshot. Rerun the prompts that look like your own catalog before you pick a default.