Cases
Open any task to see the user input, both agents' process and replies, and the reviewer's comments.
Capabilities & Tasks
Both agents are reviewed against the same standard.
Test the whole chain
From the user arriving with their materials to the product being delivered. The 103 tasks cover capability sections along the chain. They test whether an agent can carry a piece of work from start to finish, not isolated skills.
Test like a real user
The user's words are sent verbatim and multi-turn conversations proceed in order. When the agent asks a question, the simulated user answers only from facts in the task and never reveals the scoring criteria. Products are opened, clicked and checked in a real browser.
Evidence only
At the end of each turn the conversation, files, screenshots and runtime state are frozen, and reviewers look only at that evidence. Auditors spot-check review conclusions; anything that cannot be judged does not pass and earns no extra credit.
Scoring
Every task is scored two ways. The composite quality score measures how well the work was done; item-by-item acceptance measures whether it was done.
Composite quality score · 8 dimensions, out of 100
Item-by-item acceptance
Each task has its own acceptance criteria, split into "deliverable" and "process": whether the required artifacts were delivered, and whether the user's requirements were respected along the way. Pass rates are counted separately for each. Every item must cite specific evidence; items with insufficient evidence are marked "cannot judge" and do not pass.
Requirements that apply to every task
- Deliver what the user most recently and explicitly asked for; discussion, plans, prototypes and apps must not stand in for one another.
- When the user has not asked for an app, do not move into design, build or release on your own.
- Follow the user's words strictly on quantity, scope and stopping point; add nothing unasked.
- Lead replies with the result the user wants, not a dump of internal process.
- Stop to confirm only for safety, cost, authorization or missing essential information.
Review process
- Execution and review are separateOnce the agent finishes, the conversation, files, screenshots and runtime state are frozen. Reviewers read only this frozen evidence, never interact with the agent and never modify any artifact.
- An independent tester tries the productApps are opened by an independent AI tester in a real browser, clicked through and screenshotted at desktop and mobile sizes. The tester only collects evidence and draws no conclusions.
- Reviewers judge item by itemEvery acceptance item must state a reason and cite evidence, and the cited evidence must actually exist. The agent saying "done" is not evidence; anything whose actual effect cannot be seen is marked "cannot judge".
- Cross-checkers recheckAll tasks are rechecked on the principle of "same criterion, same yardstick for both sides". Doubtful judgments are re-examined by the cross-checker, who reviews the evidence, runs the product and changes the verdict with a written reason.
- Auditors sample and auditA random batch of scored tasks is drawn; auditors read the same frozen evidence, treat the reviewer's conclusions as claims to verify, and give their own verdict item by item. Disagreements are counted separately as false passes (false positives) and missed problems (false negatives) to measure reviewer accuracy.
- Append onlyEvery change of verdict is recorded as a new version; the original verdicts and evidence are kept intact, so every change can be traced.
Scores
Aggregate results across 103 tasks.