AI Router · CLI · MCPCheapest eligible quotes before you create
use-case · consideration

Creative proof-set operations for generation teams

Define sample size, controlled variables, review rubric, and a stop-or-scale decision.

Start with 20 free credits
Tool rack

Operate proof sets as a queue of bounded decisions

Creative proof-set operations begin after a team has more candidate requests than it should generate at once. Give every request a decision owner, one question, a frozen input packet, one controlled variable, a maximum attempt count, a cost ceiling, and a due date. The proof set is an internal production gate: it can show whether a recipe is ready for a larger batch, needs revision, or should stop. It is not evidence that an asset will improve clicks, conversions, revenue, or audience response.

A current OpenAI customer story names Nobuhisa Hirata, Assistant Chief General Manager of MIXI's FamilyAlbum Business Department, as a practitioner using GPTs in advertising creative planning to enhance brand understanding and design A/B tests. He reports an approximate reduction of 28 work hours per month, and the same story describes teams that collaboratively test and refine prompts. This is useful named workflow evidence, but it is a vendor-published customer account, not an independent benchmark and not proof that another team will save the same amount.

Intake
One decision, owner, variable, ceiling, and deadlineReject requests that ask one small set to answer several unrelated creative or performance questions.
Exit
Stop, revise, or scale the exact recipeA scale decision authorizes a bounded next batch, not automatic publication or media spend.
Credit gauge

Choose sample size from coverage, repetition, and consequence

Do not hide sample size behind the word batch. Start with a coverage table listing the ordinary case, the most important edge case, and the failure-prone case. Assign at least one output to every case, then add repetitions only where stochastic variation could change the decision. Record planned and actual attempts separately, including rejected and failed attempts. A team evaluating three cases with two repetitions should disclose six planned outputs; if two retries are needed, the operation cost and review record cover eight attempts.

The sample is qualitative unless a statistician or platform experiment has sized it for a defined metric. Google Ads' current Demand Gen guidance illustrates the distinction: its campaign experiments hold all but one variable constant, recommend a 50% traffic split, require a selected success metric, and expose directional or conclusive confidence settings. Conversion-related metrics need at least 100 data points before results start calculating, and conversion-based bidding experiments need at least 50 conversions per arm to surface results; an inconclusive result remains informative. A handful of generated proofs cannot inherit those campaign-level claims.

Coverage count
Cases × planned repetitionsName the normal, edge, and failure-prone cases instead of presenting an unexplained total.
Attempt count
Planned outputs + retries + rejected outputsUse this count for cost and capacity reporting even when only selected proofs appear in a contact sheet.
Inference limit
Recipe readiness, not audience performanceMove performance questions into a properly configured platform experiment with its own sample and spend.
Model specimen

Freeze the variables that make rows comparable

Create a run card before generation. Freeze the approved brief, source assets, model and version, prompt template, aspect ratio, output duration or resolution, safety settings, review rubric, and finishing assumptions. Change one declared factor across rows, such as composition or motion direction. If a model swap also forces a duration or schema change, label the comparison exploratory; do not attribute the difference to one factor.

Keep the review order stable. Reviewers should first apply hard stops for rights, unsafe content, invented claims, source-fidelity failures, unreadable required text, or an unusable crop. Only surviving rows receive preference scores for composition, edit effort, or brand fit. Meta's current A/B guidance similarly starts with a hypothesis, recommends evenly split audiences and equal budgets for a fair comparison, and warns that informal on-and-off testing can produce overlapping audiences and unreliable results. Those campaign controls support the operating principle, while the internal proof remains a pre-launch review.

Frozen
Brief, inputs, model, format, rubric, and reviewer orderStore hashes or stable references for source files so a later row cannot silently use different material.
Changed
One named factorWhen several factors must move together, record the bundle and avoid a single-cause conclusion.
Output contact sheet

Publish a cost ledger beside every proof decision

A disclosed proof-set cost should separate provider usage, platform subscription or account costs, human review time, finishing work, and paid-media spend. Record the displayed rate at the time of the run, the units consumed by every attempt, taxes or account-specific charges when known, and the source URL used for the estimate. Do not convert credits into currency unless the provider publishes that conversion for the same interface.

Runway's current API documentation prices organization credits at $0.01 each and lists Gen-4 Turbo at 5 credits per output second, Gen-4.5 at 12 credits per second, Gen-4 Image Turbo at 2 credits per image, and Gen-4 Image at 5 credits for 720p or 8 for 1080p. At those API rates, twelve five-second Gen-4 Turbo attempts have a disclosed generation base of 300 credits, or $3.00, before tax, retries, input charges, labor, and finishing. This dated arithmetic is an operating example, not a quote for a future run.

A second current first-party price shows why the model identifier belongs in the ledger: fal lists Seedream 4.5 text-to-image at $0.04 per image and permits one to six images in a request. Twelve outputs therefore have a listed inference base of $0.48 at the retrieved rate. The page did not run either provider, so these figures disclose vendor rates and transparent arithmetic, not observed billing or comparable output quality.

Fit filter

Separate operator, reviewer, and scale authority

Assign three responsibilities even when one person holds more than one role. The operator executes the frozen run card and logs every attempt. The reviewer applies the rubric without rewriting the success rule after seeing attractive outputs. The scale authority accepts the evidence limits, cost total, unresolved defects, and next-batch ceiling. Record names and timestamps so a later team can tell who generated, who judged, and who accepted additional spend.

Use a change log between rounds. A rejected set should produce a named change to the prompt, source, model, format, or rubric, followed by a new set identifier. Do not overwrite failed rows or merge rounds into one favorable contact sheet. For video exploration, Runway explicitly recommends testing Gen-4 generations in lower-cost Turbo before moving to Gen-4 as needed; its current guide lists 5 credits per second for Turbo and 12 for Gen-4. That is a provider-specific escalation path, not evidence that Turbo quality is sufficient for every brief.

Operator
Runs the card and preserves all attemptsMay report execution errors but does not silently substitute a new recipe.
Reviewer
Applies hard stops before preferencesRecords pass, revise, or fail with a reason tied to the prewritten rubric.
Scale authority
Approves the next ceilingCan approve a smaller follow-up, request a new proof, or stop with no winner.
Related circuit

Use the creator-use-case hub when the team has not chosen the medium or generation route. Move to human-approved creative automation when the central problem is where approval gates belong around spend, delivery, publication, or another irreversible action. Use multi-format content packaging when an approved concept must become coordinated wide, square, vertical, still, and motion deliverables. These links lead to distinct next decisions; they do not enlarge the claim made by this proof set.

Canonical plate

Evidence limits and operational boundary

Publication evidence is narrow. OpenAI's customer story attributes an approximate 28-hour monthly reduction only to Nobuhisa Hirata's FamilyAlbum advertising-creative planning context and remains a vendor-published account, not an independent benchmark. Google Ads documents one-variable Demand Gen experiments, success metrics, confidence settings, and conversion data thresholds; Meta documents evenly split audiences, equal test budgets, and the risk of informal on-and-off testing. These sources support controlled decision boundaries, not the statistical adequacy or likely audience performance of a reviewed generation proof.

Cost evidence is bounded. Runway publishes its API credit conversion and model rates, fal publishes Seedream 4.5's per-image rate, and Runway recommends lower-cost Gen-4 Turbo for initial testing before Gen-4 as needed. These are vendor rates and guidance, not invoices. No generation, billing reconciliation, staff-time study, contact-sheet review, or live advertising experiment was performed. The page's operating procedures are a bounded team design grounded by cited controls and prices, not a benchmark, product comparison, or promised outcome.

Recent-source coverage for this refresh was degraded: optional X credentials were absent, Reddit ended partial after RSS failures, and YouTube retained zero final items. The engine returned 104 dated records, but none were relevance-qualified community evidence for proof-set sample sizes, operating methods, provider preferences, costs, adoption, reliability, creative quality, savings, campaign performance, or customer outcomes. Access gaps are not evidence that those communities were quiet.