MODEL FIELD NOTES · Image experiments
Qwen Image Storyboards: Four Panels, Twelve Panels and Pro Trade-offs
A twelve-panel storyboard did not inevitably collapse in our tests. The harder question was how much work each panel had to do: people, actions, composition and text all compete for space. Two small experiment records show why panel count alone is a weak success measure.
在黄果制片阅读中文版Two case studies, not a model leaderboard
The first record is a four-image A/B exercise dated August 24, 2026, compared with earlier standard-version outputs. The second contains three 3×4, twelve-panel images. These samples do not establish a general success rate.
The names “3.0” and “3 Pro” were Qwen Image labels used by a third-party service at the time. We have not independently established their mapping to official weights, so this is not an official model benchmark.
Original creative assets are not published with this article. Visual findings below are summaries of the contemporary human review, not side-by-side evidence readers can independently inspect here.
Improvements and regressions occurred on different axes
Better compliance with a character or clothing instruction did not guarantee better margins, panel numbering or page structure. Reducing those observations to a single “better model” verdict would hide the trade-off.
| Check | Observation in the four-image A/B | Workflow implication |
|---|---|---|
| Four-panel line art | Pro sometimes added caption strips or a footer | Check whether the page template is still usable |
| Numbering | Some outputs substituted letters or incorrect numbers | Add final numbering in an editor |
| Realistic people and clothing | Pro followed some instructions more closely | Consider it for individual keyframes |
| Single-keyframe layout | A better subject could come with extra panels | Review the full page as well as the subject |
Correct panel count was only the first gate
The twelve-panel exercise used a 1536×2048 canvas arranged in three columns and four rows. It produced standard line art, Pro line art and Pro realism. The report recorded the requested twelve panels in all three outputs.
Panel-level review still found incomplete actions. Twelve correctly placed boxes do not prove twelve usable shots, and a storyboard sheet should not automatically become a reliable identity reference for video generation.
These samples justify testing a lighter allocation of action: alternate single-person, two-person and environment close-ups instead of filling every panel with complex interaction. They do not prove this pattern works across stories or seeds.
Separate the overview sheet from production keyframes
An overview establishes shot order, spatial relationships and rhythm. A production keyframe establishes identity, composition and the starting state of one shot. Asking one image to do both makes failures harder to isolate.
- Begin with four panels and inspect cast size, shot scale and readable action.
- When expanding to twelve panels, keep the story stable and record seed and settings. Track new layout failures separately.
- Score count, numbering, cast, clothing, action and background for each panel rather than awarding one global score.
- Regenerate a failed panel as an independent keyframe instead of repeatedly redrawing the whole sheet.
- Add important lettering and numbering afterward and keep editable source files.
A fairer comparison to run next
The historical records were not a controlled four-versus-twelve-panel benchmark. A stronger follow-up needs the same story, multiple paired seeds and a fixed rubric, with model changes separated from prompt changes.
Higher price should not be treated as higher usable-output rate. We have not rechecked live pricing and do not present old charges as current plan advice. Decide whether the task is a layout overview or a realistic keyframe, then compare the total cost of obtaining a usable result.