Eazo Benchmark: Real Creative Tasks

103 tasks drawn from real user scenarios, covering the full creative chain from understanding the request and handling assets to design, development and delivery. Both agents are evaluated on the same set of tasks and reviewed independently against the same standard.

Green: Eazo leads or tiesLight green: Eazo trails by under 5 pointsYellow: Eazo trails by 5 points or more
01

Cases

Open any task to see the user input, both agents' process and replies, and the reviewer's comments.

02

Capabilities & Tasks

Both agents are reviewed against the same standard.

The six outer sections follow the order of creative work; the cross-cutting capabilities in the centre apply to every step. Numbers are task counts.

Test the whole chain

From the user arriving with their materials to the product being delivered. The 103 tasks cover capability sections along the chain. They test whether an agent can carry a piece of work from start to finish, not isolated skills.

Test like a real user

The user's words are sent verbatim and multi-turn conversations proceed in order. When the agent asks a question, the simulated user answers only from facts in the task and never reveals the scoring criteria. Products are opened, clicked and checked in a real browser.

Evidence only

At the end of each turn the conversation, files, screenshots and runtime state are frozen, and reviewers look only at that evidence. Auditors spot-check review conclusions; anything that cannot be judged does not pass and earns no extra credit.

03

Scoring

Every task is scored two ways. The composite quality score measures how well the work was done; item-by-item acceptance measures whether it was done.

Composite quality score · 8 dimensions, out of 100

Each dimension scored on 5 levels, 0–40 not met2 partly met4 fully metDimensions that do not apply are excluded

Item-by-item acceptance

Each task has its own acceptance criteria, split into "deliverable" and "process": whether the required artifacts were delivered, and whether the user's requirements were respected along the way. Pass rates are counted separately for each. Every item must cite specific evidence; items with insufficient evidence are marked "cannot judge" and do not pass.

Requirements that apply to every task

  • Deliver what the user most recently and explicitly asked for; discussion, plans, prototypes and apps must not stand in for one another.
  • When the user has not asked for an app, do not move into design, build or release on your own.
  • Follow the user's words strictly on quantity, scope and stopping point; add nothing unasked.
  • Lead replies with the result the user wants, not a dump of internal process.
  • Stop to confirm only for safety, cost, authorization or missing essential information.

Review process

  1. Execution and review are separateOnce the agent finishes, the conversation, files, screenshots and runtime state are frozen. Reviewers read only this frozen evidence, never interact with the agent and never modify any artifact.
  2. An independent tester tries the productApps are opened by an independent AI tester in a real browser, clicked through and screenshotted at desktop and mobile sizes. The tester only collects evidence and draws no conclusions.
  3. Reviewers judge item by itemEvery acceptance item must state a reason and cite evidence, and the cited evidence must actually exist. The agent saying "done" is not evidence; anything whose actual effect cannot be seen is marked "cannot judge".
  4. Cross-checkers recheckAll tasks are rechecked on the principle of "same criterion, same yardstick for both sides". Doubtful judgments are re-examined by the cross-checker, who reviews the evidence, runs the product and changes the verdict with a written reason.
  5. Auditors sample and auditA random batch of scored tasks is drawn; auditors read the same frozen evidence, treat the reviewer's conclusions as claims to verify, and give their own verdict item by item. Disagreements are counted separately as false passes (false positives) and missed problems (false negatives) to measure reviewer accuracy.
  6. Append onlyEvery change of verdict is recorded as a new version; the original verdicts and evidence are kept intact, so every change can be traced.
04

Scores

Aggregate results across 103 tasks.

EazoCodex

Composite quality score · average by dimension

Average composite quality score by capability section