~8-step distill — iterate prompts aggressively
Decoupled-DMD / DMDR few-step distillation is the official pitch: composition, expression, wardrobe can be tested fast; waiting half a minute per tweak kills short-video cover schedules.
Alibaba Tongyi MAI · Z-Image-Turbo · 2025-11
~8-step distilled inference, photoreal skin and light, readable Chinese/English in-image text — use free credits on iMini for avatars, headshots, and bilingual poster drafts first.
Z Image Turbo is Tongyi MAI (Tongyi Lab)’s distilled text-to-image model open-sourced in November 2025 — a ~6B single-stream diffusion Transformer (S3-DiT) that ships at ~8 NFE steps on paper, sub-second latency on enterprise H800, and runs on ~16GB consumer GPUs. Official benchmarks put it against closed flagships on photoreal detail, light, materials, and bilingual in-image text (small poster type still readable while faces hold). GitHub / Hugging Face also cite #1 open model on Artificial Analysis text-to-image and strong Elo on Alibaba AI Arena vs closed models. Family split: Turbo for fast generation, Edit for instruction edits. On iMini it is the default for free-credit realism / portrait / bilingual drafts — switch to Qwen Image Plus for long Chinese poster layout, Nano Banana Pro / GPT Image 2 for native 4K polish. Blurry skin, fake light, or broken shop-sign Chinese kills headshots and bilingual KV — Turbo’s job is to cheap-test those failures before final art.
Decoupled-DMD / DMDR few-step distillation is the official pitch: composition, expression, wardrobe can be tested fast; waiting half a minute per tweak kills short-video cover schedules.
Official chapter on detail, light, texture, mood — avatars, headshots, makeup close-ups pass the “looks human” gate first; one fake face voids the whole portrait set.
Official emphasis on bilingual rendering rivaling top closed models; small poster type still holds. Wrong shop copy or blurred titles kill bilingual KV / campaign visuals.
Open weights (Apache 2.0), enter workspace from this page with free credits — no vendor key setup first.
Car-window rim light, indoor window light, gown side light, soft beauty light, Peking opera makeup, creative mask — official photoreal samples where skin and light hold up.






Z Image Turbo vs poster-oriented Qwen Image Plus, Google’s Nano Banana 2, and OpenAI’s GPT Image 2 — step count, photoreal portraits, bilingual text, and default picks side by side.
| Dimension | Z Image Turbo | Qwen Image Plus | Nano Banana 2 | GPT Image 2 |
|---|---|---|---|---|
| Speed profile | ~8-step distill · sub-second (H800) | Instant fast draft + quality tiers | ~8-step distill · sub-second (H800) | Instant fast draft + quality tiers |
| Portraits | Photoreal realism (official) | Good | Photoreal realism (official) | Photoreal realism (official) |
| In-image text | Good | Industry-best text precision | Photoreal realism (official) | Photoreal realism (official) |
| Resolution | Up to 2K | Up to 2K | Up to 4K | Native 4K (3840×2160) |
| Chinese support | Official strength in CN/EN in-image text | Official strength in CN/EN in-image text | Multi-reference support | Multi-reference support |
| Draft speed | ~8-step distill · sub-second (H800) | Instant fast draft + quality tiers | ~8-step distill · sub-second (H800) | Instant fast draft + quality tiers |
| Editing | Basic repaint / variations | Basic repaint / variations | Basic repaint / variations | High-fidelity edit + identity hold |
| Default pick | Default for photoreal portraits / bilingual drafts | Chinese posters and long copy | Speed + 4K daily workhorse | OpenAI default for new projects |
Few-step realism, portrait texture, bilingual text — cheap iteration on this lane before final art.
~8 steps per image — swap prompts, light, makeup without queue pain; slow iteration breaks sample schedules.
Skin, hair, rim light, depth of field — avatars and headshots pass “looks human” first; one fake face voids the set.
Shop signs, poster lines, bilingual titles readable; blurred or wrong text kills campaign assets.
Fabric, beads, dappled daylight — same portrait spec as official samples.
Burn free credits while direction is unsettled; upgrade to Nano Banana Pro or GPT Image 2 for 4K / multi-ref polish.
Official samples map to these deliverables: lifestyle candid, daily headshot, editorial, beauty, makeup, creative portrait.

Car-window rim light, travel candids — highlights and ambient light must land; fake light fails lifestyle review.

Indoor window-light half-body for social and profile drafts; fake skin or muddy hair kills profile output.

Gown, side light, depth of field — magazine half-body; collapsed light or fake face voids the editorial set.

Soft beauty light, lace and hair detail; muddy materials kill beauty / lookbook samples.

Peking opera makeup and beaded headpieces — features and materials must hold; muddy detail makes close-ups useless.

Masks, embroidery, unconventional styling still human; one drifted look kills creative pitches.
Teams that need cheap, fast photoreal portraits and bilingual drafts that pass first review.
Covers, thumbnails, portrait samples — iterate hard on free credits first.
Campaign KV and bilingual short-copy poster drafts, compared with flagships in-browser.
Model shots and headshot direction unsettled — lock light and expression first.
Pitch boards and look direction without the most expensive lane before sign-off.
Early brand visuals and persona art — cheap trials until the story is clear.
Many client sample versions, compared with Qwen / Banana / GPT Image on one screen.
In the iMini image workspace, select Z Image Turbo.
Subject, style, aspect ratio; quote in-image text.
Download what works, or compare with sibling models on iMini.
Starters aligned to official photoreal samples — copy, swap subject, generate free on iMini.

Young East Asian woman in passenger seat smiling at camera, highway guardrail and blue sky outside, strong rim light, hair highlights, real skin texture, 85mm half-body, vertical frame.

Long black hair young woman half-body, white printed tee, bookshelf and window light behind, natural daylight from right, real skin and hair detail, shallow depth of field, vertical frame.

Middle-aged East Asian man three-quarter profile, black tux and bow tie, golden hour light, lake and distant city skyline behind, cinematic shallow depth of field, magazine editorial portrait.

Young woman soft-light half-body, light pink lace top, long curls over shoulders, clear hair and skin detail, dappled natural light with slight prism flare, shallow depth of field, vertical frame.

Peking opera dan role half-body close-up, traditional white base red cheek makeup, ornate blue-gold beaded headdress, red embroidered costume, warm side light, blurred background, vertical photoreal portrait.

Young East Asian woman half-body, lower face in embroidered cat mask with gold bead edge and fine stitching, long bangs, dappled tree shade daylight, neutral wall background, vertical photoreal portrait.
Speed claims, realism, how it splits work with Qwen / flagships, and how credits work.
6B · ~8 steps · photoreal realism and CN/EN text — try on iMini before polish.
Generate free on iMiniFree credits to start · no separate API key