Skip to content

Benchmarks vs Vibes

AI labs publish benchmark scores to prove their model is the best. We collected 24,000+ blind preference votes to see if users agree. Mostly, they don't.

24,000+

Blind votes

276

Models ranked

1,990

Unique matchups

19.3%

End in ties

The Disconnect

Every lab launches with benchmark scores. MMLU, HumanEval, MATH, ARC. Higher number, better model. That's the pitch.

Show people the actual outputs side by side, no names and no scores, and they keep picking models that don't top those leaderboards. The model leading the Rival Index (Qwen3.6 27B, 82% win rate) does not lead any traditional benchmark.

What Each Approach Measures

Benchmarks

  • +Factual recall on multiple-choice tests
  • +Code completion on defined problems
  • +Math reasoning under controlled conditions
  • +Reproducible, automated scoring

Measures the ceiling of narrow capabilities. Not which model you'll actually prefer.

Vibe Testing

  • +Human preference on real, open-ended tasks
  • +Design taste, tone, and aesthetics
  • +Creative and subjective output quality
  • +Blind comparison, no brand bias

Measures what users prefer when they see the output. Harder to fake.

What We Found

19.3%

of matchups end in a tie. Nearly 1 in 5 times, users can't tell two models apart. On everyday tasks, the top models have converged.

276

models have logged enough blind duels to earn a rank. It's not a few leaders and everyone else. Hundreds are competitive.

82%

win rate for Qwen3.6 27B, currently topping the Rival Index. It leads no traditional benchmark. The two are measuring different things.

SVG

art is among the most-voted challenges, like drawing an Xbox controller. People test design taste, not math proofs.

We're Not the Only Ones Saying This

Andrej Karpathy (ex-OpenAI, ex-Tesla AI) keeps arguing for “vibe checks” over benchmarks. Stanford's HAI Index documented benchmark saturation. Artificial Analysis found scores diverge from real-world preference.

Our part is an open dataset: 24,000+ votes across 1,990 matchups, from landing pages to SVG art to code. The methodology is transparent and the data is free to download.

See for Yourself

Pick two models. Same prompt, real outputs. Vote, then check if the benchmarks agree.

FAQ

Sign in