AI Router · CLI · MCPCheapest eligible quotes before you create
evaluation · consideration

Evaluate typography in AI-generated graphics

Test exact text separately from composition and document the downstream replacement workflow.

Start with 20 free credits
Failure trace

Evaluate the words separately from the graphic

A polished composition is not proof that its typography is usable. Review every generated graphic on two tracks: first verify the literal string, then judge the visual treatment. The copy track checks spelling, capitalization, punctuation, word order, missing characters, substitutions, and invented lettering. The design track checks hierarchy, font fit, spacing, alignment, contrast, placement, and whether the type still reads at the intended display size. A candidate fails if either track fails, even when the overall image looks convincing.

For an OfflineCreator-focused test, Recraft V3 is the relevant starting point because the current Studio catalog positions it for brand graphics at 5 credits. The same page displayed 12 credits when accessed on August 9, 2026; that value is a historical snapshot, not the current price. Catalog positioning is not evidence that Recraft will reproduce exact copy. Recraft's own documentation says V3 can generate text but does not handle very small text reliably; distorted text and typos are documented limitations. Treat each output as a test result rather than a typography guarantee.

Error code index

Build a matched Studio sample set

Use one fixed brief for every candidate in the set. Keep the model, prompt, aspect ratio, requested text, language, capitalization, and generation count unchanged. A useful test string combines a short headline, a number, punctuation, and a less common proper noun so the review can expose substitutions without becoming an unrealistic paragraph-generation challenge. Save every result, not only the strongest image, and label it with the model, prompt, run order, and review date.

Do not compare a carefully selected winner with another model's first attempt. Review the same number of outputs, at the same target size, against the same checklist. TextInVision combines text-recognition measures with human review of whether generated text is accurate and clear, while TypoBench evaluates typography beyond spelling, including font, size, weight, color, alignment, spacing, curvature, placement, and preservation of surrounding design. Those studies support a split rubric instead of one vague aesthetic score.

Prompt anatomy

Score exact text before style

Transcribe the visible lettering without looking back at the requested copy, then compare the two strings character by character. Record a pass only when every required character is present in the correct order. Keep separate defect labels for omission, insertion, substitution, transposition, duplicated text, and letter-like noise. This prevents an attractive type treatment from hiding a misspelling and makes repeated runs comparable.

Automated OCR can help locate likely mismatches, but it is not the final judge. TextInVision used OCR and edit distance as quantitative measures, then added human evaluation for prompt following and whether text was accurate and clear. For a production review, inspect ambiguous glyphs at full resolution and at delivery size. A string that OCR reads correctly can still contain malformed strokes, unstable spacing, or low-contrast details that a person cannot comfortably read.

Output contact sheet

Review typographic fidelity and layout

After the string passes, assess the typography against the intended role. Check whether headline and supporting text have a clear hierarchy, whether letter and line spacing remain even, whether alignment is intentional, and whether the text stays inside its safe region. Inspect repeated letters for inconsistent shapes and look for collisions with faces, products, logos, or important background details. Review both the full image and the smallest planned placement because defects often disappear on a large artboard and reappear at thumbnail size.

Keep these observations as separate fields rather than collapsing them into one pass/fail judgment. TypoBench is grounded in 989 real design templates and reports that typography involves multiple properties, including font, size, weight, color, alignment, spacing, curvature, and inline style. Its generation tasks also measure whether inserted text remains in the target region and whether surrounding content is preserved. That is a useful boundary for a creator rubric, but its benchmark results are not OfflineCreator performance results.

Provider disclosure

Apply the accessibility gate to the final placement

For web delivery, ask whether the required words need to remain pixels at all. WCAG 2.2 Success Criterion 1.4.5 calls for text instead of images of text when the same presentation can be achieved with the available technology, with exceptions for customizable images and essential presentation such as logotypes. A practical pass is therefore often a strong generated composition with the exact headline rebuilt as live text, not approval of embedded lettering merely because it is spelled correctly.

If readable text remains in the image, measure contrast against the actual background behind each character. WCAG's minimum-contrast guidance sets at least 4.5:1 for ordinary text and 3:1 for large-scale text, subject to its stated exceptions. Also provide an appropriate text alternative: W3C's image tutorial says images of text should contain the same words in the alternative, informative images need a short description of essential information, and decorative images use an empty alternative. Choose based on the image's role and avoid repeating nearby live copy unnecessarily.

Related circuit

Use the Recraft V3 guide when the next step is another model-specific composition pass. Return to the model directory when the task or budget is still undecided. If exact words are the only failing element, stop regenerating the entire visual and hand the approved composition to a normal design or web-production workflow for controlled type placement.

Canonical plate

Page boundary and research limits

This page owns the evaluation question: how to test typography in AI-generated graphics without confusing composition quality with exact-copy fidelity. It does not promise that a model will spell every string correctly, reproduce a particular font, satisfy accessibility requirements automatically, or preserve text behavior across future versions. Model selection details belong in the model guides; the downstream replacement procedure belongs in the workflow guidance.

Recent community evidence is limited and anecdotal. One r/graphic_design poster described typography and typesetting as an easy way to spot AI-generated design, while one TikTok critique used a single flyer to call out too many fonts, too much text, and weak hierarchy. Neither item is a controlled comparison, a Studio test, or evidence of model-level performance. Treat Studio-specific typography performance as unmeasured because this draft includes no matched Studio sample set and reports no exact-copy, layout, contrast, or accessibility result for a Studio output. Before stronger editorial treatment, run an authorized matched test, preserve every candidate and its settings, and record the accessibility checks. Until that evidence exists, this page remains an unreviewed, unpublished research draft rather than a completed model evaluation.