Create a proof set before scaling AI creative
Validate a small, representative sample before spending on more variants.
A creative proof set is a deliberately small batch made before a larger generation run. Its job is to expose prompt ambiguity, recurring visual defects, unsafe source material, unsuitable aspect ratios, and an uneconomic model choice while the cost of changing direction is still low. It is not a miniature ad-performance study and it cannot establish which creative will win after publication.
Write the decision first: for example, “Does this prompt-and-model combination reliably leave usable copy space while preserving the product silhouette?” Freeze the brief, model, ratio, and review rubric for the set. Google Ads similarly advises experimenters to state a clear hypothesis, test one variable at a time, choose success metrics before the test, and keep records. Those principles make a proof set interpretable even though this small qualitative review is not a statistically powered campaign experiment.
Build, label, and review the set in six steps
1. Name one decision and one variable. 2. Choose three representative cases: a normal brief, a known edge case, and a failure-prone case. 3. Generate two outputs for each case, for six outputs total. 4. Give every output an ID and record its case, prompt version, model, ratio, and credits. 5. Score every output against the same prewritten criteria. 6. Stop, revise, or approve the recipe before producing more variants.
The six-output pattern is a disclosed operational starting point, not a universal sample-size claim. It borrows the structure of a balanced one-factor design: NIST describes one primary factor in terms of its levels and replications, and notes that equal replication at each level improves sensitivity. For a production performance claim, calculate sample size for the actual metric and risk tolerance instead of treating six generated assets as statistically conclusive.
- Sample disclosure
- 3 cases × 2 outputs = 6 reviewed outputsState the model, prompt version, ratio, case definitions, repetitions, exclusions, and review date beside the result.
- Control
- Change one planned factorKeep the brief and rubric fixed; if you change the model and prompt together, label the set exploratory rather than attributing the difference.
- Decision
- Approve, revise, or stopApprove only the tested recipe and scope. A pass does not predict audience response or guarantee that a larger stochastic run will contain no failures.
Clear source rights and the cloud boundary first
Do not put confidential or unlicensed source material into a proof set merely because the batch is small. Record who approved each source image, what downstream use is intended, and whether the prompt or upload may leave the device. If a case needs material that cannot be disclosed to an external processor, remove that case or choose an appropriate local workflow before generation.
OfflineCreator’s provider disclosure says Studio sends generation inputs to the provider serving the selected model, routes the current launch catalog through fal, and is not an offline service. That boundary belongs in the proof record whenever inputs include a product still, unreleased campaign details, people, or client material. A technically successful output should still fail review when its source or intended use lacks approval.
Log credits before scaling the recipe
Create one ledger row per attempt, including failed and rejected outputs: output ID, timestamp, case, model, prompt version, ratio, displayed unit cost, job result, credits charged or returned, and review outcome. Add the rows rather than multiplying only the selected examples. This exposes the real proof cost and prevents a contact sheet from hiding unsuccessful attempts.
The current Studio launch catalog lists FLUX Schnell at 1 credit, FLUX 1.1 Pro Ultra at 8, Recraft V3 at 5, Kling 2.6 Pro at 42, Kling 2.6 Pro Motion at 42, and Veo 3.1 Fast at 96. Studio’s image generator shows the selected model’s credit cost before submission. At those current rates, six FLUX Schnell attempts have a displayed base cost of 6 credits, while six FLUX 1.1 Pro Ultra attempts have a displayed base cost of 48 credits. Recheck the live display before running because this page’s arithmetic is a dated example, not a quote for future work.
Use a review rubric that can reject attractive failures
Score each output before looking for a favorite. A practical rubric can mark brief fit, composition, source fidelity, text safety, obvious artifacts, rights or disclosure concerns, and edit effort as pass, revise, or fail. Add a hard-stop column for defects that make an asset unusable, such as a changed product shape, invented label text, missing copy-safe space, or content the team is not authorized to send to a provider.
Google’s evaluation guidance recommends defining measurable criteria from ideal behavior, business goals, and known failure cases before building a benchmark. It also recommends covering happy paths, edge cases, and adversarial examples, and manually rating a small sample when aligning an automated evaluator. Applied here, the human-reviewed proof set becomes a reusable challenge set; it does not become proof that subjective creative quality has been automated.
Turn the proof into the next narrow action
Build a contact sheet when reviewers need to compare all six outputs without cherry-picking. Move to the vertical-video guide only after a still or shot recipe passes the proof criteria. Return to the workflow directory if the proof shows that the chosen medium, model class, or input path is wrong rather than merely under-prompted.
Keep this page scoped to pre-scale proof
This guide owns the operational proof-set decision: disclose the sample, account for every generation credit, apply fixed review criteria, and record why the team approved, revised, or stopped. The parent workflow directory owns broad model routing, while the contact-sheet guide owns comparative presentation. Campaign experimentation owns audience-level performance claims; this small generation review must never be presented as evidence of lift, conversion, or statistical significance.