Benchmarks vs Vibes
AI labs publish benchmark scores to prove their model is the best. We collected 24,000+ blind preference votes to see if users agree. Mostly, they don't.
24,000+
Blind votes
276
Models ranked
1,990
Unique matchups
19.3%
End in ties
The Disconnect
Every lab launches with benchmark scores. MMLU, HumanEval, MATH, ARC. Higher number, better model. That's the pitch.
Show people the actual outputs side by side, no names and no scores, and they keep picking models that don't top those leaderboards. The model leading the Rival Index (Qwen3.6 27B, 82% win rate) does not lead any traditional benchmark.
What Each Approach Measures
Benchmarks
- +Factual recall on multiple-choice tests
- +Code completion on defined problems
- +Math reasoning under controlled conditions
- +Reproducible, automated scoring
Measures the ceiling of narrow capabilities. Not which model you'll actually prefer.
Vibe Testing
- +Human preference on real, open-ended tasks
- +Design taste, tone, and aesthetics
- +Creative and subjective output quality
- +Blind comparison, no brand bias
Measures what users prefer when they see the output. Harder to fake.
What We Found
of matchups end in a tie. Nearly 1 in 5 times, users can't tell two models apart. On everyday tasks, the top models have converged.
models have logged enough blind duels to earn a rank. It's not a few leaders and everyone else. Hundreds are competitive.
win rate for Qwen3.6 27B, currently topping the Rival Index. It leads no traditional benchmark. The two are measuring different things.
art is among the most-voted challenges, like drawing an Xbox controller. People test design taste, not math proofs.
We're Not the Only Ones Saying This
Andrej Karpathy (ex-OpenAI, ex-Tesla AI) keeps arguing for “vibe checks” over benchmarks. Stanford's HAI Index documented benchmark saturation. Artificial Analysis found scores diverge from real-world preference.
Our part is an open dataset: 24,000+ votes across 1,990 matchups, from landing pages to SVG art to code. The methodology is transparent and the data is free to download.
See for Yourself
Pick two models. Same prompt, real outputs. Vote, then check if the benchmarks agree.