Skip to main content

Design Creativity Bench

Today’s models can build functional, feature-rich interfaces, yet designers and developers often find the results recognisably AI-generated. The main reason, we think, is repetition: models explore a narrow range of creative possibilities and return to the same choices of layout, structure and visual theme.

Design Creativity Bench measures how much models repeat themselves. It has 80 open-ended UI tasks, each written for two product domains, such as a page for reviewing expense reports and one for reviewing insurance claims. Layout and style are left entirely to the designer, and human experts work on the same 160 briefs as a reference. Designs are compared with a similarity score that focuses on structural decisions like navigation, page layout and content format, scores colour separately, and has been validated against human judgements. The score measures originality, how distinct a model’s designs are from other models’, and creative range, how much they change between domains. A checklist for each brief measures appropriateness: whether the design actually does what was asked.

Creativity Score

GPT-6 Astra
GPT-6.1 Sol
Grok 4.7
Gemini 3.8 Flash
Qwen3.8 Max
GPT-5.6 Sol
Claude Opus 5.5
Claude Fable 5.1
Muse Spark 1.3
DeepSeek V4.1 Flash
Grok 4.6
GLM-5.3
Kimi K3
GLM-5.3-Flash
MiniMax M3

Creativity Score vs. Cost ($)

Creativity Score2030405060708090100
Human reference 92.3
$0.02$0.03$0.05$0.10$0.20$0.50$1$2$6
Median cost ($) · log scale

The default output of LLMs is substantially more repetitive than human design.

As LLM adoption in UI design accelerates, we are entering an era of design monoculture - what the internet increasingly calls “AI slop.” Our data backs this up: The overall score, which weights both diversity and appropriateness, reflects the exact same story. GPT-6 Astra tops the models at 62.9, but the human reference sits far ahead at 84.6.

When we asked different models to build a screen for the exact same brief, they clustered around a highly repetitive default, with a median score of 0.591 on distinctiveness. Introduce a human to the mix, and distinctiveness jumps to 0.767. Worse, models don't adapt to context. Ask a model to design for a clinic versus a pharmacy, and its designs barely change (0.581 median distinctiveness). A human designer intuitively shifts their approach to match the product, resulting in a distinctiveness score of 0.902. The models have the capacity for range, but without steering, their default state is a sea of sameness.

Creative range0.50.60.70.80.90.550.600.650.700.750.80+4.8 SD originality+6.1 SD creative range15 modelsHuman reference
Originality

Models by use case

Best Overall

OpenAI

GPT-6 Astra

Release date: Sep 2026

  • Creativity Score65.5
  • Originality0.646
  • Creative Range0.472
  • Appropriateness99.1%
  • Median Cost ($)$2.564

Models are very good at building exactly what the brief asks for.

The uniformity of AI design isn’t a sign of failure. In fact, the models we tested were remarkably good at doing precisely what they were told. Every one of them satisfied at least 90% of a brief’s requirements, and the best of them (Claude Opus 5.5 and GPT-6 Astra) outperform human designers at following instructions, with 99.1% appropriateness.

When models do falter, their mistakes tend to be small and specific. In 31% of cases, the failure involved leaving out a control tied to a specific part of the brief. In 27% of cases, the models generated sample data that didn’t quite add up. We also noticed that with each new generation, models tend to become slightly better at following instructions, and slightly less capable of producing a variety of designs. The latest GPT-Sol, Claude Opus, and Grok models, for instance, all outdo their predecessors in building what they were asked for, but they have the same or less creative range. The question now is not whether LLMs can design an interface that matches the brief, but whether they can do it with human-level spark - either through human direction or with increased creative capability of their own.

Full results

Model
Human reference92.30.7640.90298.0%——
65.50.6460.47299.1%$2.56413.8
63.70.6420.48798.8%$0.46015.8
62.20.6430.60397.8%$0.59318.8
60.60.6370.62897.5%$0.1230.9
58.10.6230.61797.6%$0.30918.0
51.30.5870.50598.2%$0.5984.4
51.00.5570.57499.2%$1.55713.1
48.60.5580.58998.5%$5.32913.9
47.20.5680.56597.8%$0.0801.1
43.60.5760.59696.3%$0.0515.3
43.40.5910.62595.5%$0.1003.5
41.90.5790.63495.6%$0.0432.5
40.40.5470.61596.8%$0.57017.0
38.50.5530.61196.1%$0.02210.0
34.00.5870.60293.4%$0.0242.5