One prompt asks for a one-month starter plan. The other wants a three-month biohacker protocol with dosages. Same models, opposite readers.
2 challenges20 models
20 models tested across 2 health & longevity challenges.Composite score: 30% Rival Index, 20% task coverage, 20% challenge-scoped duel performance, 15% recency, 15% tier. Deduplicated by product line. Gemini 3.1 Pro Preview leads at 80.9/100. Drawn from Rival's open dataset of 21,000+ human preference votes.
Rival's Pick#6 Rival IndexGoogle flagship
Too close to callComposite scores combine task evidence, Rival Index, recency, and model tier. Rival’s Pick is a separate editorial recommendation. How ranking works