AI News

ChatGPT Images 2.5 Sketch Test: 23 First Outputs, One Retained Failure

Method note: this is a single-session, zero-cost ChatGPT test. I used eight deliberately rough local diagrams, not hand-drawn originals. The completed matrix has 24 planned conditions: 23 generated first outputs and one retained no-output attempt. No paid API calls, reruns, or image edits were used.

ChatGPT Images 2.5 makes a bad drawing useful. A rough sketch can supply the part that prose often loses: where things go, what is large, and which empty space matters. I tested that idea with a movie poster, living room, bottle ad, sneaker, thumbnail, creature, book cover and landing-page hero.

What the completed test found

Every scenario had three frozen first-output conditions: A, a direct text-only prompt; B, the rough sketch plus a minimal request to polish it; and C, the same sketch plus a full transformation brief. I scored layout match (40 points), idea recognition (25), detail compliance (20) and shareability (15).

Condition Generated outputs Mean score
A: text only 8 94.4 / 100
B: sketch only 7 64.6 / 100
C: sketch + full brief 8 97.8 / 100

The sketch-plus-brief condition won this particular first-output set. That does not make it a general benchmark: this is eight curated concepts, one output per condition, on one ChatGPT session. The service did not identify an underlying image model in the web UI, so the results do not establish model-level performance.

The scorecard

Scenario A: text only B: sketch only C: sketch + full brief
Movie poster 89 91 94
Living room 95 34 96
Beverage ad 95 72 99
Sneaker concept 94 FAIL 99
Creator thumbnail 99 85 96
Creature character 95 19 99
Cabin book cover 93 67 99
Landing-page hero 95 84 100

The FAIL is not a zero masquerading as a score. The sneaker’s sketch-only request returned no assistant output after 90 seconds, and the experiment kept that service-side failure rather than replacing it. The low creature sketch-only result is also part of the record: the first image was an unrelated face-and-cube composition.

Text alone can make a good image. A sketch can make it your image.

The text-only controls were usually attractive. They also show why a polished image is not the same as a controlled one. The full brief supplied style, material, palette and exclusions. The sketch supplied the hierarchy: a face left and cube right, a product region on the right of a landing page, or a blank title strip beneath a cabin.

Sketch-only was the unstable middle ground. It preserved the poster’s broad hierarchy, the bottle’s copy space and the landing-page wireframe. But one living-room request became an exterior, the cabin became a daylight photograph instead of a snowy cover, and the creature missed entirely. A drawing without a brief is a map without a destination.

Five ways to make Sketch useful

  1. Draw the order of attention. Make the main object larger than everything else.
  2. Mark negative space. Empty areas communicate room for a headline or logo.
  3. Use arrows and boxes. They carry movement, placement and hierarchy quickly.
  4. Write down what the drawing cannot show. Add medium, palette, lighting, materials and exclusions.
  5. Protect exact copy in the prompt. Use explicit words for visible text and verify the first output.

Try this workflow

Turn this rough sketch into a polished finished image. Preserve the composition and every object. Add no text.

Use that minimal prompt when you want to see what the layout carries by itself. For a usable asset, follow it with a transformation brief: name the subject, preserve the important spatial relationships, choose a medium and palette, and list what must not appear. OpenAI’s release notes describe image generation from the ChatGPT message box, alongside templates, image comments and prompt sharing for further steering.

Verdict

Sketch is useful when you think in layouts before you think in prompts. The evidence here supports a simple workflow: draw the composition, then add the brief that tells the model what the drawing should become. The combination delivered the highest average in this frozen first-output run. It is still one small, quota-limited experiment with system-drawn diagrams, so treat it as a practical example rather than proof of universal image quality.

Sources