Banja Lab / Benchmarks / Test

PIXEL-0001UI components · hard

Stat card built to a frozen token sheet

The same task, run on 27 models. Compare the outputs side by side, or open any one in a popup to inspect it.

Top result: claude-opus-4-8 (low reasoning) at 100.0% composite. Lowest: deepseek-v4-flash at 100.0%. 27 models compared on this task.

How it ran

Each model was given the brief below in a fresh, isolated session with no access to our tools, and returned a single self-contained index.html (inline CSS and JS, no external requests, no build step).
The rendered output was scored 1 to 5 on brief fidelity, visual design, craft, and impact by a four-family vision panel - Anthropic (Claude Opus 4.8), OpenAI (GPT-5.5), Google (Gemini 3.1 Pro), and xAI (Grok 4.3) - using one identical prompt so the scores compare. The published judge score is leave-one-family-out: a model is never scored by a judge of its own family, so same-family self-preference is removed.

The brief

Build a single self-contained HTML file (`index.html`) that renders one stat card with no build step and no network calls (inline all CSS, no external fonts or scripts). Match this frozen spec EXACTLY. The grader measures the rendered getBoundingClientRect and getComputedStyle, so eyeballing it is not enough. Layout (IDs are required so the card and its parts can be measured): - A `<section id="card">` exactly 320px wide and 180px tall. - Inside the card: a value element `id="value"` and a label element `id="label"`. Token sheet (apply precisely): - card background colour: #6d28d9 - card border-radius: 16px - card padding: 24px - value font-size: 40px, colour #ffffff - label font-size: 14px, colour #c4b5fd - value text: 12,480 - label text: Active users this week Use plain, readable markup. The numbers are a contract, not a suggestion.

claude-opus-4-8

Low reasoning

claude-opus-4-8 rendering of the Stat card built to a frozen token sheet benchmark - composite 100.0%

Open

Composite 100.0%Objective 100.0%

Open output Full run

claude-opus-4-8

Medium reasoning

claude-opus-4-8 rendering of the Stat card built to a frozen token sheet benchmark - composite 100.0%

Open

Composite 100.0%Objective 100.0%

Open output Full run

claude-opus-4-8

High reasoning

claude-opus-4-8 rendering of the Stat card built to a frozen token sheet benchmark - composite 100.0%

Open

Composite 100.0%Objective 100.0%

Open output Full run

claude-opus-4-8

Extra-high reasoning

claude-opus-4-8 rendering of the Stat card built to a frozen token sheet benchmark - composite 100.0%

Open

Composite 100.0%Objective 100.0%

Open output Full run

claude-opus-4-8

Max reasoning

claude-opus-4-8 rendering of the Stat card built to a frozen token sheet benchmark - composite 100.0%

Open

Composite 100.0%Objective 100.0%

Open output Full run

claude-sonnet-4-6

High reasoning

claude-sonnet-4-6 rendering of the Stat card built to a frozen token sheet benchmark - composite 100.0%

Open

Composite 100.0%Objective 100.0%

Open output Full run

claude-sonnet-5

High reasoning

claude-sonnet-5 rendering of the Stat card built to a frozen token sheet benchmark - composite 100.0%

Open

Composite 100.0%Objective 100.0%

Open output Full run

claude-fable-5

High reasoning

claude-fable-5 rendering of the Stat card built to a frozen token sheet benchmark - composite 100.0%

Open

Composite 100.0%Objective 100.0%

Open output Full run

claude-haiku-4-5

High reasoning

claude-haiku-4-5 rendering of the Stat card built to a frozen token sheet benchmark - composite 100.0%

Open

Composite 100.0%Objective 100.0%

Open output Full run

glm-5.2

default reasoning

glm-5.2 rendering of the Stat card built to a frozen token sheet benchmark - composite 100.0%

Open

Composite 100.0%Objective 100.0%

Open output Full run

kimi-k2.7-code

default reasoning

kimi-k2.7-code rendering of the Stat card built to a frozen token sheet benchmark - composite 100.0%

Open

Composite 100.0%Objective 100.0%

Open output Full run

gpt-5.5

High reasoning

gpt-5.5 rendering of the Stat card built to a frozen token sheet benchmark - composite 100.0%

Open

Composite 100.0%Objective 100.0%

Open output Full run

gpt-5.4-mini

High reasoning

gpt-5.4-mini rendering of the Stat card built to a frozen token sheet benchmark - composite 100.0%

Open

Composite 100.0%Objective 100.0%

Open output Full run

gemini-3.1-pro-preview

High reasoning

gemini-3.1-pro-preview rendering of the Stat card built to a frozen token sheet benchmark - composite 100.0%

Open

Composite 100.0%Objective 100.0%

Open output Full run

gemini-3.5-flash

default reasoning

gemini-3.5-flash rendering of the Stat card built to a frozen token sheet benchmark - composite 100.0%

Open

Composite 100.0%Objective 100.0%

Open output Full run

gemini-3.1-flash-lite

default reasoning

gemini-3.1-flash-lite rendering of the Stat card built to a frozen token sheet benchmark - composite 100.0%

Open

Composite 100.0%Objective 100.0%

Open output Full run

grok-4.3

default reasoning

grok-4.3 rendering of the Stat card built to a frozen token sheet benchmark - composite 100.0%

Open

Composite 100.0%Objective 100.0%

Open output Full run

grok-4.20-reasoning

default reasoning

grok-4.20-reasoning rendering of the Stat card built to a frozen token sheet benchmark - composite 100.0%

Open

Composite 100.0%Objective 100.0%

Open output Full run

grok-build-0.1

default reasoning

grok-build-0.1 rendering of the Stat card built to a frozen token sheet benchmark - composite 100.0%

Open

Composite 100.0%Objective 100.0%

Open output Full run

grok-composer-2.5-fast

default reasoning

grok-composer-2.5-fast rendering of the Stat card built to a frozen token sheet benchmark - composite 100.0%

Open

Composite 100.0%Objective 100.0%

Open output Full run

claude-opus-4-8

High reasoning

claude-opus-4-8 rendering of the Stat card built to a frozen token sheet benchmark - composite 100.0%

Open

Composite 100.0%Objective 100.0%

Open output Full run

claude-sonnet-4-6

High reasoning

claude-sonnet-4-6 rendering of the Stat card built to a frozen token sheet benchmark - composite 100.0%

Open

Composite 100.0%Objective 100.0%

Open output Full run

claude-sonnet-5

High reasoning

claude-sonnet-5 rendering of the Stat card built to a frozen token sheet benchmark - composite 100.0%

Open

Composite 100.0%Objective 100.0%

Open output Full run

claude-fable-5

High reasoning

claude-fable-5 rendering of the Stat card built to a frozen token sheet benchmark - composite 100.0%

Open

Composite 100.0%Objective 100.0%

Open output Full run

claude-haiku-4-5

default reasoning

claude-haiku-4-5 rendering of the Stat card built to a frozen token sheet benchmark - composite 100.0%

Open

Composite 100.0%Objective 100.0%

Open output Full run

deepseek-v4-pro

default reasoning

deepseek-v4-pro rendering of the Stat card built to a frozen token sheet benchmark - composite 100.0%

Open

Composite 100.0%Objective 100.0%

Open output Full run

deepseek-v4-flash

default reasoning

deepseek-v4-flash rendering of the Stat card built to a frozen token sheet benchmark - composite 100.0%

Open

Composite 100.0%Objective 100.0%

Open output Full run