Skip to content

SubjectiveBench

Does it have taste?

10 models, one question, one uncapped score judged by humans. Turns out most of them draw the same seagull.
v1·10,758 outputs·226 models·independent, no paid placement
100 · the reference050100150taste index · uncapped
Hover a model to inspect
Every scored model on one uncapped scale. The whole field still sits short of 100.

The shape of the field

Competent, identical, and nowhere near 100

48
spread
60
top
21
median
40
headroom

The rankings

Every model, by taste

Read the output behind any score. Near-equal scores share a rank.

#ModelIndex
  1. Capable
    1
    n=56·craft 69·orig 69·see output
    60
  2. =3
    n=18·craft 52·orig 39·see output
    44
  3. =3
    n=57·craft 51·orig 38·see output
    43
  4. =3
    n=58·craft 52·orig 37·see output
    42
  5. =3
    n=56·craft 52·orig 33·see output
    41
  6. =3
    n=58·craft 50·orig 37·see output
    41
  7. Generic
    =9
    n=57·craft 50·orig 33·see output
    39
  8. =9
    n=58·craft 51·orig 30·see output
    39
  9. =9
    n=35·craft 47·orig 33·see output
    38
  10. =9
    n=58·craft 47·orig 34·see output
    38
  11. =9
    n=53·craft 49·orig 33·see output
    38
  12. =9
    n=58·craft 50·orig 30·see output
    38
  13. =9
    n=57·craft 49·orig 31·see output
    38
  14. =9
    n=57·craft 48·orig 30·see output
    37
  15. =9
    n=58·craft 50·orig 29·see output
    36
  16. =9
    n=58·craft 49·orig 28·see output
    36
  17. =9
    n=52·craft 49·orig 26·see output
    36
  18. =9
    n=58·craft 48·orig 28·see output
    36
  19. =9
    n=58·craft 47·orig 28·see output
    35
  20. =9
    n=53·craft 47·orig 27·see output
    35
  21. =9
    n=38·craft 45·orig 28·see output
    34
  22. =9
    n=47·craft 43·orig 28·see output
    34
  23. =9
    n=53·craft 48·orig 25·see output
    33
  24. =9
    n=53·craft 47·orig 25·see output
    33
  25. =9
    n=58·craft 41·orig 28·see output
    33
  26. =9
    n=53·craft 45·orig 25·see output
    33
  27. =9
    n=53·craft 42·orig 27·see output
    32
  28. =9
    n=51·craft 43·orig 25·see output
    32
  29. =9
    31
  30. =9
    n=53·craft 43·orig 25·see output
    31
  31. =9
    n=58·craft 41·orig 26·see output
    31
  32. =9
    31
  33. =9
    n=55·craft 44·orig 24·see output
    31
  34. =9
    n=59·craft 39·orig 27·see output
    31
  35. =9
    n=57·craft 44·orig 23·see output
    31
  36. =9
    n=41·craft 44·orig 23·see output
    31
  37. =9
    n=58·craft 42·orig 23·see output
    31
  38. =9
    n=58·craft 42·orig 24·see output
    30

100 is the top. Nothing reaches it yet. SubjectiveBench v1 · June 2026.

Method

How it's scored

01

Every output read

Each output is judged against the rest on the same prompt, originality first.

02

A human decides

A person scores the work and tweaks until the ranking matches the taste.

03

One number, uncapped

A Taste Index per output and per model. Higher is rarer. Most sit below 100.

Cite this

SubjectiveBench v1 (June 2026). rival.tips. https://www.rival.tips/subjectivebench

Questions, answered+
What is SubjectiveBench?
A taste benchmark. Every other benchmark measures whether a model is competent. This one measures whether it has taste: craft, originality, and whether it escapes the answer every other model gives. One uncapped score per output and per model, judged by humans.
Why is the scale uncapped, and why does nothing reach 100?
Because taste has headroom and competence does not. 100 is where genuinely original, tasteful work would sit. No model reaches it yet. The best lands well below, and most cluster near the floor making the same default choices. The scale runs past 100 because when a model finally gets there, taste keeps going. A 0 to 100 percentage would pretend there is a ceiling. There is not.
Isn't taste just your opinion?
Yes, and we say so out loud. There is no objective ground truth for taste. So we do two things. A human reads and scores every output, originality first, because the homogeneous default is the thing to punish. And we put every scored output on the site. Read them and disagree. A benchmark you can audit beats a number you have to trust.
Sign in