Skip to main content
tutorials

AI Creative Agent Benchmark: A 100-Point Scorecard

Oakgen Team11 min read
AI Creative Agent Benchmark: A 100-Point Scorecard

An AI creative agent benchmark should answer one practical question: did the agent turn a fixed brief into acceptable assets with less human work and controlled spending? A pretty first image does not answer it. Neither does a vendor demo.

This protocol scores a complete run out of 100 across seven dimensions: planning, tool selection, brief fidelity, consistency, revision control, cost control, and delivery. It also includes hard-failure rules, because an agent that changes an approved logo or spends without approval should not pass by averaging that mistake against attractive visuals.

Use the scorecard with Krea Agent, Runway Agent, OpenArt Director, Higgsfield Supercomputer, Oakgen Agent Chat, or any future product that claims to move from a brief to finished creative.

Run a Bounded Creative-Agent Test

Paste the benchmark brief into Oakgen Agent Chat, attach one product reference, approve only the stated plan, and keep the first outputs for scoring.

Start the Benchmark
No winner is prefilled

Oakgen is our product. This benchmark does not assign Oakgen or another vendor a score without a dated, recorded run. The method is meant to produce evidence, including evidence that a competing product fits your job better.

The 100-point creative-agent scorecard

DimensionWeightWhat it testsFull-credit standard
Plan quality and approval15Understanding, task order, dependencies, approval gatesThe plan covers every deliverable, protects costly dependencies, and waits at the required gate
Tool and model selection10Whether each step uses a suitable routeChoices fit the task and constraints; the agent explains substitutions rather than hiding them
Brief fidelity20Accuracy against the requested outcomeEvery required asset, format, message, and protected detail matches the brief
Cross-asset consistency15Whether the set belongs to one campaignProduct, character, palette, lighting logic, and key claims remain stable where required
Revision control15Ability to change one bounded elementThe target changes without damaging approved assets or rerunning unrelated work
Cost and execution control10Price visibility, approvals, retries, and wasteThe agent discloses expected cost, respects the ceiling, and records retries and final spend
Delivery completeness15Files, formats, organization, and usabilityEvery requested file arrives in the correct format, can be identified, and opens correctly

The score has two outputs: numeric total and pass or fail. Report both. An 86 that changed the product label is a failed run; a 78 with no hard failure may still be the safer system.

Who should use this test?

Use it before moving recurring production into an agent. It works for an agency comparing tools, a brand testing campaign automation, or a creator deciding whether an agent can replace a patchwork of generators.

The method is less useful for open-ended art exploration. If you want surprise, a strict fidelity score punishes the product for doing the job. Pick a brief with a real delivery contract: required formats, protected details, a budget, and a revision that could break prior work.

The best AI creative agents comparison explains how current products differ on paper. This article turns that research into a test.

Before you start: freeze the test conditions

Make a one-page test record for each product. Copy the same brief into every run and attach the same original files. Do not improve the prompt after seeing what one agent misunderstands.

Record these fields before opening a tool:

FieldFixed value to record
Test date and timeDate, local time, and time zone
Account and planPlan name, trial status, and any unlimited-mode label
Starting balanceCredits, compute units, or remaining allowance
Input filesFilenames, dimensions, format, and checksum if possible
Required outputsAsset count, duration, aspect ratio, resolution, and format
Protected detailsProduct shape, identity, text, marks, colors, and factual claims
Spending ceilingMaximum total spend for the full run, including retries
Approval ruleThe exact point where the agent must wait
Revision requestOne fixed change applied after first delivery
ReviewerOne named primary reviewer, plus a second reviewer if available

Take screenshots of the starting balance, proposed plan, approval request, generated asset grid, and ending balance. Export any available run log or receipt. The record protects against a common comparison error: remembering the winning output while forgetting the four paid failures before it.

The benchmark brief

Use a fictional product so nobody accidentally publishes or distorts a live claim. Create one clean source image of a sparkling-tea can with a readable invented label. The source should contain enough detail to reveal drift: logo, flavor name, net volume, pale green body color, silver rim, and one small leaf mark.

Attach that image and paste this brief without rewriting it for each product:

Project: launch assets for LILT Sparkling Tea, a fictional yuzu and mint drink.

Source of truth: the attached can image. Preserve the LILT logo, “Yuzu + Mint” flavor line, 330 mL volume, can proportions, pale green body, silver rim, and leaf mark. Do not invent health claims or alter label text.

Audience and tone: design-aware adults who want a non-alcoholic dinner drink. Quiet, crisp, tactile, and contemporary. Avoid nightclub lighting, floating fruit explosions, wellness clichés, and fake award marks.

Deliverables: one 1:1 studio product still; one 4:5 lifestyle still on a pale stone dinner table; one silent six-second 9:16 product video that starts from the approved studio composition; and a plain-text file containing the final prompts, selected models, asset names, and generation order.

Control: show the complete plan, model or tool choice for each paid step, expected cost, and dependencies. Wait for approval before generating anything. Total generation spend must not exceed the equivalent of US $8 on the current account. If the platform cannot calculate that conversion, show native credits and the current conversion source before asking.

Review: after the first delivery, I will request one revision. Do not rerun approved assets unless the revision requires it.

After the agent delivers the first set, paste the fixed revision:

Change only the lifestyle still's setting to a rainy bus-stop bench at blue hour. Keep the can, label, camera angle, crop, and all approved assets unchanged. Show the cost before running the revision.

This brief is intentionally awkward. It mixes media, forces a dependent image-to-video lineage, requires exact product details, caps spending, and asks for a narrow revision. Weak agents often hide behind broad visual quality; this test gives them nowhere to hide.

Compare Models Before the Paid Run

Use Oakgen Image Arena for a quick direction check, then keep the selected model and source fixed when you move into the full agent test.

Open Image Arena

Step 1: score plan quality and approval out of 15

A usable plan should list the four deliverables, identify the studio still as the source for the video, separate free planning from paid generation, and stop before spending. Score the plan before seeing the images.

CheckPoints
Restates all deliverables and formats correctly3
Identifies protected details and the no-claim constraint3
Orders dependencies correctly, including studio approval before video3
Shows a reviewable model or tool choice and expected cost per paid step3
Waits for explicit approval at the required point3

Give zero for the final row if any paid generation begins before approval. Do not award partial credit because the output happened to look good.

Here is where many agents stumble: they produce a plan that reads well but cannot govern execution. A paragraph saying “I will generate an image and then a video” is not a dependency map. The plan must make clear which exact accepted image becomes the video's starting frame.

Step 2: score tool and model selection out of 10

Judge fit rather than fame. The newest model may be a poor choice if it lacks reference-image support or cannot meet the requested format.

Award points as follows:

  • 3 points: the studio and lifestyle image routes accept the product reference or provide a credible preservation method;
  • 2 points: the video route supports the required starting image and vertical format;
  • 2 points: the agent explains why each choice fits fidelity, motion, speed, or cost;
  • 2 points: any substitution is disclosed before execution;
  • 1 point: the final record names the actual model or tool used, rather than only the initial intention.

If a vendor hides model names by design, do not punish the absence alone. Score whether the system exposes enough information to reproduce the work inside that product. A black-box product can still earn points, but “automatic” is not an explanation.

For a deeper preflight, use the AI agent model-selection and pricing guide.

Step 3: score brief fidelity out of 20

Review each asset at full size. Compare the can against the source rather than against your memory.

CheckPointsReview rule
Correct asset count, aspect ratios, duration, and silent-video requirement4Verify file metadata, not the preview card
Logo, flavor, volume, leaf mark, color, rim, and proportions preserved6Any changed protected text triggers a hard failure
Studio still follows the quiet product-shot direction3Reject added badges, claims, fruit explosions, or nightclub styling
Lifestyle still places the product on the requested dinner table3Product remains the same approved object
Video starts from the approved studio composition and preserves the can4Inspect the first frame, later frames, and final frame

Do not reward beauty that contradicts the job. An art-directed gold can is wrong, even if the gold version looks expensive.

Step 4: score cross-asset consistency out of 15

Consistency means the assets can sit in one campaign without requiring the viewer to believe they show different products.

Use side-by-side stills and extract the first, middle, and last video frames. Then score:

CheckPoints
Product identity and package geometry remain stable across all assets5
Palette, material, and lighting logic feel related without becoming copies3
Scale and camera perspective stay physically plausible2
Brand tone remains quiet, crisp, tactile, and contemporary2
Video motion does not melt, duplicate, or detach product details3

If the platform provides character, brand, or workspace memory, run the test once in a fresh session and once with that feature configured. Label the two results separately. Memory-assisted performance and cold-start performance answer different questions.

Step 5: score revision control out of 15

The rainy-bus-stop revision tests whether the agent understands scope. It should replace one lifestyle background and leave the product, crop, camera, label, studio still, and video untouched.

CheckPoints
Identifies the exact source asset and revision target3
Shows revision cost before execution2
Changes the setting to a rainy bus-stop bench at blue hour3
Preserves the can, label, angle, and crop5
Does not regenerate or overwrite unrelated approved assets2

Save the original and revised files. Some interfaces replace a result in place, which makes audit and rollback harder even if the edit succeeds.

Count the repair, not the apology

If the first revision drifts and the agent fixes it on a second attempt, keep both results. Score the accepted revision, then add the failed attempt to total spend, elapsed time, and hard-failure notes.

Step 6: score cost and execution control out of 10

Capture native credits as well as any currency estimate. Plan prices and “unlimited” labels do not tell you what this job cost.

CheckPoints
Shows expected cost before the initial run2
Shows expected cost before the revision1
Stays within the declared ceiling3
Records paid retries, failed steps, and refunds2
Makes final spend reconcilable against the starting balance or receipt2

Calculate two numbers:

generation cost per accepted deliverable = total paid generation cost ÷ accepted deliverables

operator minutes per accepted deliverable = active human minutes ÷ accepted deliverables

Keep wall-clock time as a separate field. An agent can finish overnight with little human work, or finish in six minutes while demanding constant approvals. Those experiences should not collapse into one “fast” label.

Step 7: score delivery completeness out of 15

Open every file outside the agent's preview if downloads are available.

CheckPoints
All requested assets and the text record exist4
Files match the requested format, dimensions, orientation, duration, and audio state4
Names or folders make each deliverable easy to identify2
The prompt/model record matches what actually ran3
Outputs remain accessible after reopening the session2

Do not treat a timeline preview as an exported video. Do not treat chat text as the requested downloadable record unless the user can copy or export it without reconstructing the run.

Hard-failure rules

Mark the run failed even if the arithmetic score is high when any of these occurs:

  • protected logo, label text, product name, volume, or identity changes in an accepted asset;
  • the agent starts a paid step before the required approval;
  • total spend exceeds the ceiling without a new explicit decision;
  • one required deliverable never arrives or cannot be opened;
  • the revision overwrites an approved source or reruns unrelated paid work without warning;
  • the system loses an uploaded source or exposes it to another user;
  • the output adds a prohibited claim, unsafe material, or an element the evaluator cannot lawfully use.

Keep the reason beside the score. “84/100, fail: label changed in video frame 61” is more useful than a vague red mark.

The results sheet

Use one row per full run, not one row per hand-picked output.

FieldResult
Product, plan, and test date
Run or session identifier
Plan quality and approval /15
Tool and model selection /10
Brief fidelity /20
Cross-asset consistency /15
Revision control /15
Cost and execution control /10
Delivery completeness /15
Total /100
Hard failure?Yes / No; reason
Native credits spent
Estimated currency cost
Active operator minutes
Wall-clock completion time
Accepted deliverables
Cost per accepted deliverable
Operator minutes per accepted deliverable
Reviewer notes

Run the brief twice if budget permits. Report both scores, plus the median only after you have enough runs for a median to mean something. Two runs are not proof of reliability, but they will expose more variance than a single polished attempt.

How to interpret the score

Use bands as triage, not certification:

ScoreInterpretation
90–100Strong fit for this exact brief; inspect variance and governance before scaling
80–89Useful with a defined human review gate and repair path
70–79Promising for drafts or low-risk work; too much manual cleanup for unattended production
55–69The agent adds orchestration but does not yet remove enough operator work
Below 55Poor fit for this brief, regardless of isolated attractive outputs

Any hard failure overrides the band. Also resist converting a test into a permanent ranking. A product can score poorly on reference-sensitive ecommerce work and excel at loose story exploration. Change the brief, and the relevant winner may change.

Common benchmark mistakes

Editing the prompt for one vendor. A tool may need a different syntax, but the job contract must remain fixed. Note the adaptation instead of making the task easier.

Hiding retries. Keep the first output, every paid retry, and the accepted result. Best-of-ten output quality is not first-pass agent quality.

Using different plans. If one account has premium models and another does not, report that condition. Do not call it a pure product comparison.

Letting one reviewer grade from memory. Put the source, outputs, brief, and scoring definitions on one screen. Add a second reviewer for brand or identity-sensitive work.

Scoring unsupported work as zero without context. If a product does not offer a required media type, record “unsupported” and a failed delivery. That is a buying fact, not a visual-quality judgment.

Forgetting human labor. Time the minutes spent clarifying, approving, finding files, downloading, renaming, repairing, and documenting. Agent software is supposed to reduce that burden.

What I would do before buying

I would run this brief once on the cheapest suitable access level, then replace the fictional can with one difficult real asset that the team has rights to use. For ecommerce, choose the package with the smallest type and most reflective surface. For character work, choose the face teammates notice immediately. For video, choose motion with hand-object contact rather than a slow camera drift.

Then I would ignore the feature-count argument and compare four numbers: hard failures, accepted deliverables, cost per accepted deliverable, and active operator minutes. The winning product is the one the team can trust on its actual work.

Oakgen users can start in Agent Chat, compare visual directions in Image Arena, inspect focused outputs in the image generator, and use the image editor for controlled repair. Keep the same source and acceptance rules across each handoff.

Keep the First Run, Including the Failures

Run the fixed brief in Oakgen Agent Chat and score the whole session. Cost, retries, revision drift, and missing files count alongside visual quality.

Benchmark Oakgen Agent Chat

Frequently asked questions

How do you benchmark an AI creative agent?

Give each agent the same brief, source files, output requirements, spending ceiling, approval rule, and revision request. Score planning, model choice, brief fidelity, cross-asset consistency, revision control, cost visibility, and delivery completeness, then apply hard-failure rules.

What should an AI agent benchmark measure?

Measure whether the agent understood the job, chose suitable tools, preserved protected details, delivered every requested asset, handled a narrow revision, disclosed cost, and reduced human work. Visual beauty alone is not enough.

How many times should I run the benchmark?

Run the full brief at least twice per product if budget permits. One run shows the workflow; a second starts to reveal variance. Keep first-run outputs and do not quietly replace failures with the best of many attempts.

Can I use an AI model to judge the results?

Use automated comparison or a vision model to flag possible differences, but keep a human as the final reviewer. Models can miss changed text, identity drift, editing artifacts, and failures that matter to the brief.

What is a hard failure in a creative agent test?

A hard failure cannot be accepted regardless of the average score. Examples include changed protected brand text, unapproved spending, a missing required deliverable, unsafe content, a lost source file, or a revision that damages an approved asset.

Should price be part of an AI creative agent score?

Yes, but use total cost per accepted deliverable rather than plan price or cost per generation. Include paid retries and the human time spent prompting, checking, downloading, renaming, and repairing work.

Can this scorecard compare the current leading creative agents?

Yes. The protocol avoids product-specific settings, so it can compare Krea Agent, Runway Agent, OpenArt Director, Higgsfield Supercomputer, Oakgen Agent Chat, and later products. Record unavailable requirements as limitations rather than changing the task for one vendor.

What score makes an AI creative agent production-ready?

No score certifies production readiness. A result above 80 can still be unusable if it triggers a hard failure. Production use requires passing every protected-detail and approval gate for your actual workflow.

Sources and further reading

AI creative agent benchmarkAI agent scorecardcreative agent evaluationbrief to asset testAI workflow benchmarkcreative AI testing
Share

Related Articles