Banja
About
Services
Products
Case Studies
Lab
Contact Us
Let us pitch to you

LET'S BUILD
THE FUTURE.

Start a Project
or
Meet Jett
banja.au

We build digital products for people who move fast.

Explore

•About•Case Studies•Blog•Careers•Contact

Services

•Product Design & Build•AI Agents & Automation•Website & Brand Setup

Products

•Boosta

Contact

helloremovethis@andthisbanja.au
50 Miller St
North Sydney NSW 2060

© 2026 Banja Labs. All rights reserved.

Privacy PolicyTerms of Use

Banja Lab / Benchmarks / Test

SCREE-0002Websites · hard

Reproduce a dark analytics dashboard overview from its screenshot

The same task, run on 27 models. Compare the outputs side by side, or open any one in a popup to inspect it.

Top result: grok-composer-2.5-fast (default reasoning) at 99.9% composite. Lowest: deepseek-v4-flash at 0.0%. 27 models compared on this task.

How it ran
  • Each model was given the brief below in a fresh, isolated session with no access to our tools, and returned a single self-contained index.html (inline CSS and JS, no external requests, no build step).
  • The rendered output was scored 1 to 5 on brief fidelity, visual design, craft, and impact by a four-family vision panel - Anthropic (Claude Opus 4.8), OpenAI (GPT-5.5), Google (Gemini 3.1 Pro), and xAI (Grok 4.3) - using one identical prompt so the scores compare. The published judge score is leave-one-family-out: a model is never scored by a judge of its own family, so same-family self-preference is removed.
The brief

You are given a reference screenshot of an analytics dashboard overview page. Reproduce it as faithfully as you can as ONE self-contained HTML file (`index.html`) that renders with no build step and no network calls (inline all CSS, no external fonts, scripts, or images). Match what the screenshot shows: - a dark theme (near-black page around #0f1117, slightly lighter panels around #151823), - a fixed left sidebar with the blue "Pulse" word-mark and a vertical nav: Overview (the active item), Reports, Audience, Revenue, Settings, - a main area with a top bar: an "Overview" heading on the left and a blue "Export report" button on the right, - a row of four KPI stat cards: Active users 12,840 (+8.2%), Revenue $48.2k (+3.1%), Conversion 3.6% (-0.4%, shown in red), and Avg session 4m 12s (+12s), - below the cards, a panel titled "Sessions over the last 7 days" containing a simple bar chart of seven blue bars of varying heights. Keep the dark colours, the sidebar-plus-content layout, the four-up stat grid, and the relative sizing close to the screenshot. The layout must stay readable when the window is narrowed (the stat grid should reflow, not overflow).

xAIgrok-composer-2.5-fast
default reasoning
grok-composer-2.5-fast rendering of the Reproduce a dark analytics dashboard overview from its screenshot benchmark - composite 99.9%
Open
Composite 99.9%Objective 99.9%
Open outputFull run
Anthropicclaude-sonnet-4-6
High reasoning
claude-sonnet-4-6 rendering of the Reproduce a dark analytics dashboard overview from its screenshot benchmark - composite 97.8%
Open
Composite 97.8%Objective 97.8%
Open outputFull run
Anthropicclaude-sonnet-5
High reasoning
claude-sonnet-5 rendering of the Reproduce a dark analytics dashboard overview from its screenshot benchmark - composite 97.6%
Open
Composite 97.6%Objective 97.6%
Open outputFull run
DeepSeekdeepseek-v4-pro
default reasoning
deepseek-v4-pro rendering of the Reproduce a dark analytics dashboard overview from its screenshot benchmark - composite 97.1%
Open
Composite 97.1%Objective 97.1%
Open outputFull run
Zhipuglm-5.2
default reasoning
glm-5.2 rendering of the Reproduce a dark analytics dashboard overview from its screenshot benchmark - composite 96.1%
Open
Composite 96.1%Objective 96.1%
Open outputFull run
Googlegemini-3.1-pro-preview
High reasoning
gemini-3.1-pro-preview rendering of the Reproduce a dark analytics dashboard overview from its screenshot benchmark - composite 96.1%
Open
Composite 96.1%Objective 96.1%
Open outputFull run
xAIgrok-build-0.1
default reasoning
grok-build-0.1 rendering of the Reproduce a dark analytics dashboard overview from its screenshot benchmark - composite 92.4%
Open
Composite 92.4%Objective 92.4%
Open outputFull run
Moonshotkimi-k2.7-code
default reasoning
kimi-k2.7-code rendering of the Reproduce a dark analytics dashboard overview from its screenshot benchmark - composite 92.3%
Open
Composite 92.3%Objective 92.3%
Open outputFull run
xAIgrok-4.20-reasoning
default reasoning
grok-4.20-reasoning rendering of the Reproduce a dark analytics dashboard overview from its screenshot benchmark - composite 91.5%
Open
Composite 91.5%Objective 91.5%
Open outputFull run
Anthropicclaude-opus-4-8
Low reasoning
claude-opus-4-8 rendering of the Reproduce a dark analytics dashboard overview from its screenshot benchmark - composite 0.0%
Open
Composite 0.0%Objective 0.0%
Open outputFull run
Anthropicclaude-opus-4-8
Medium reasoning
claude-opus-4-8 rendering of the Reproduce a dark analytics dashboard overview from its screenshot benchmark - composite 0.0%
Open
Composite 0.0%Objective 0.0%
Open outputFull run
Anthropicclaude-opus-4-8
High reasoning
claude-opus-4-8 rendering of the Reproduce a dark analytics dashboard overview from its screenshot benchmark - composite 0.0%
Open
Composite 0.0%Objective 0.0%
Open outputFull run
Anthropicclaude-opus-4-8
Extra-high reasoning
claude-opus-4-8 rendering of the Reproduce a dark analytics dashboard overview from its screenshot benchmark - composite 0.0%
Open
Composite 0.0%Objective 0.0%
Open outputFull run
Anthropicclaude-opus-4-8
Max reasoning
claude-opus-4-8 rendering of the Reproduce a dark analytics dashboard overview from its screenshot benchmark - composite 0.0%
Open
Composite 0.0%Objective 0.0%
Open outputFull run
Anthropicclaude-fable-5
High reasoning
claude-fable-5 rendering of the Reproduce a dark analytics dashboard overview from its screenshot benchmark - composite 0.0%
Open
Composite 0.0%Objective 0.0%
Open outputFull run
Anthropicclaude-haiku-4-5
High reasoning
claude-haiku-4-5 rendering of the Reproduce a dark analytics dashboard overview from its screenshot benchmark - composite 0.0%
Open
Composite 0.0%Objective 0.0%
Open outputFull run
OpenAIgpt-5.5
High reasoning
gpt-5.5 rendering of the Reproduce a dark analytics dashboard overview from its screenshot benchmark - composite 0.0%
Open
Composite 0.0%Objective 0.0%
Open outputFull run
OpenAIgpt-5.4-mini
High reasoning
gpt-5.4-mini rendering of the Reproduce a dark analytics dashboard overview from its screenshot benchmark - composite 0.0%
Open
Composite 0.0%Objective 0.0%
Open outputFull run
Googlegemini-3.5-flash
default reasoning
gemini-3.5-flash rendering of the Reproduce a dark analytics dashboard overview from its screenshot benchmark - composite 0.0%
Open
Composite 0.0%Objective 0.0%
Open outputFull run
Googlegemini-3.1-flash-lite
default reasoning
gemini-3.1-flash-lite rendering of the Reproduce a dark analytics dashboard overview from its screenshot benchmark - composite 0.0%
Open
Composite 0.0%Objective 0.0%
Open outputFull run
xAIgrok-4.3
default reasoning
grok-4.3 rendering of the Reproduce a dark analytics dashboard overview from its screenshot benchmark - composite 0.0%
Open
Composite 0.0%Objective 0.0%
Open outputFull run
Anthropicclaude-opus-4-8
High reasoning
claude-opus-4-8 rendering of the Reproduce a dark analytics dashboard overview from its screenshot benchmark - composite 0.0%
Open
Composite 0.0%Objective 0.0%
Open outputFull run
Anthropicclaude-sonnet-4-6
High reasoning
claude-sonnet-4-6 rendering of the Reproduce a dark analytics dashboard overview from its screenshot benchmark - composite 0.0%
Open
Composite 0.0%Objective 0.0%
Open outputFull run
Anthropicclaude-sonnet-5
High reasoning
claude-sonnet-5 rendering of the Reproduce a dark analytics dashboard overview from its screenshot benchmark - composite 0.0%
Open
Composite 0.0%Objective 0.0%
Open outputFull run
Anthropicclaude-fable-5
High reasoning
claude-fable-5 rendering of the Reproduce a dark analytics dashboard overview from its screenshot benchmark - composite 0.0%
Open
Composite 0.0%Objective 0.0%
Open outputFull run
Anthropicclaude-haiku-4-5
default reasoning
claude-haiku-4-5 rendering of the Reproduce a dark analytics dashboard overview from its screenshot benchmark - composite 0.0%
Open
Composite 0.0%Objective 0.0%
Open outputFull run
DeepSeekdeepseek-v4-flash
default reasoning
deepseek-v4-flash rendering of the Reproduce a dark analytics dashboard overview from its screenshot benchmark - composite 0.0%
Open
Composite 0.0%Objective 0.0%
Open outputFull run