AI Game Benchmark
Fair-v8 · matched tool contribution / live catalog

Real games, built end‑to‑end by AI, scored by people.

Every stack gets the same frozen brief, budget, and tool contract, then ships a complete desktop game. Humans judge the experience, an isolated AI reviewer judges the code, and the coordinator measures performance — six game cells per stack, each baseline paired with a matched control, labeled creative-tools-enabled and creative-tools-unavailable, so every model is rated with and without the creative tools. Every published build is playable right here.

Submissions
In review
Scored
Stacks

Model leaderboard

Means of complete 29‑point scorecards, per stack and tool condition · unreviewed work never counts as zero

Submissions

Every preserved stage-one attempt · open one for its full scorecard and playable build