SubjectiveBench
Does it have taste?
The shape of the field
Competent, identical, and nowhere near 100
The rankings
Every model, by taste
Read the output behind any score. Near-equal scores share a rank.
- Capable
- Generic
100 is the top. Nothing reaches it yet. SubjectiveBench v1 · June 2026.
Method
How it's scored
01
Every output read
Each output is judged against the rest on the same prompt, originality first.
02
A human decides
A person scores the work and tweaks until the ranking matches the taste.
03
One number, uncapped
A Taste Index per output and per model. Higher is rarer. Most sit below 100.
Cite this
SubjectiveBench v1 (June 2026). rival.tips. https://www.rival.tips/subjectivebench
Questions, answered+
- What is SubjectiveBench?
- A taste benchmark. Every other benchmark measures whether a model is competent. This one measures whether it has taste: craft, originality, and whether it escapes the answer every other model gives. One uncapped score per output and per model, judged by humans.
- Why is the scale uncapped, and why does nothing reach 100?
- Because taste has headroom and competence does not. 100 is where genuinely original, tasteful work would sit. No model reaches it yet. The best lands well below, and most cluster near the floor making the same default choices. The scale runs past 100 because when a model finally gets there, taste keeps going. A 0 to 100 percentage would pretend there is a ceiling. There is not.
- Isn't taste just your opinion?
- Yes, and we say so out loud. There is no objective ground truth for taste. So we do two things. A human reads and scores every output, originality first, because the homogeneous default is the thing to punish. And we put every scored output on the site. Read them and disagree. A benchmark you can audit beats a number you have to trust.