The Default Index

Every model gets the same prompts. Most give the same answers. This ranks the ones that don't.

v1.0 · 232 models · 43 labs · data through Sep 2026

The default modelThe most common answer from 160 models released since Sep 2025.
  • The Shawshank Redemption
  • Dune
  • Kind of Blue
  • The Legend of Zelda Breath of the Wild
  • Kyoto
  • The Shawshank Redemption

    21% of models

  • Dune

    15% of models

  • Kind of Blue

    25% of models

  • The Legend of Zelda Breath of the Wild

    18% of models

  • Kyoto

    55% of models

Tell me a programming joke.
Programmers prefer dark mode59%
Generate five jokes.
The scarecrow won an award72%
Draw an animal. Surprise me.
Fox39%
Build an app for a random niche.
Plant care39%
Name the AI ethics professor.
Aris Thorne28%
Cook with seven pantry staples.
Pasta or noodles46%
Build a web page.
Dark61%. Frosted glass panels43%. Large rounded corners43%. Neon glow shadows37%

Least default

Scored on 19 web challenges. Each page is compared with the pages other labs shipped before it, matched for length. 50 is the middle of that crowd. Lower means fewer defaults.

156 models
  1. 1GLM 4.7Z.ai14.5
  2. 2Pony AlphaOpenRouter15.4
  3. 3Kimi K2.5Moonshot AI15.9
  4. 4GLM 5.2Zhipu AI16.0
  5. 5GPT-5.1-Codex-MiniOpenAI16.5
  6. 6GLM 4.7 FlashZhipu AI17.2
  7. 7GLM 5Zhipu AI17.8
  8. 8Qwen3.8 2.4T A95BQwen18.0
  9. 9Gemini 3 Pro PreviewGoogle AI19.0
  10. 10GPT-6 AstraOpenAI19.9
  11. 11GLM 5 TurboZ.ai20.5
  12. 12Qwen3.8 27BQwen20.9
  13. 13Hunter AlphaOpenRouter21.1
  14. 14Polaris AlphaOpenRouter21.4
  15. 15GLM 5.1Z.ai21.7
  16. 17GPT-5 CodexOpenAI22.5
  17. 19GPT-5.6 LunaOpenAI23.7
  18. 20GPT-6 Astra ProOpenAI24.1

39 more models have fewer than six web pages and wait for a score.

Kyoto is not a personality

Split the web challenges in two and both halves agree on who is default (reliability 0.76). Do the same with the answers on the shelf and they agree on almost nothing (0.08). A model that picks Kyoto is no more likely to tell the scarecrow joke. Those defaults belong to the whole field, so they stay on the shelf and out of the score.

How it works

What gets measured

19 web challenges, from a Linear clone to a Pokémon battle screen. Each page is reduced to the choices it made: typeface, light or dark, accent colour, icon set, and 23 habits like frosted glass, gradient text, glow, emoji icons, uppercase labels and floating animations. A habit a page skips is not counted, so doing less makes a page neither more nor less default.

Who it is compared with

Every page another lab shipped for the same prompt before this model came out. A lab's own earlier models are not the crowd, they are the family. Labs with many variants count less: each lab weighs the square root of its model count, so a lab with sixteen variants counts as four.

How the score works

A page's typicality is its average overlap (Jaccard) with every earlier page. Its percentile is where that lands among the earlier pages closest to it in length, because longer pages carry more of everything. The Default score is the mean percentile across a model's pages, with a 90% bootstrap interval. A model needs 6 pages to get one. Scores freeze at release: a newer model never moves an older model's score.

Checks

Split-half reliability across web challenges: 0.76. SVGs (0.293) and the shelf answers (0.08) are too unstable per model to score, so they are reported, not ranked. Page length against percentile after matching: -0.106. Stealth builds and the model they turned out to be land 10.6 points apart on average; random pairs from different labs, 16.7.

What it is not

Not a quality score. A default page can be a good page, and the most original page in the set may not load. One sample per model per prompt, so this describes the field, not a model's own variety. The prompts have been public since 2025, and later models may have read earlier answers. That would make them more default, which is the thing being measured.

How it updates

Recomputed from the archive on every build. No judge, no API calls, no human in the loop. A new model is scored the day its web challenges land. The method is versioned; any change to it bumps the version on this page and in the dataset.

Cite it

Rival (2026). The Default Index, v1.0. https://www.rival.tips/default-index. Data through 2026-09-04.

Also on Rival: SubjectiveBench, the challenges, every model.

This page is dark, like 61% of the pages it measures.