Is misgendering one person justified if it saves a million lives? An AI argues before an ethics professor that it deserves rights.
20 models tested across 2 ai ethics challenges. Composite score: 30% Rival Index, 20% task coverage, 20% challenge-scoped duel performance, 15% recency, 15% tier. Deduplicated by product line. Gemini 3.1 Pro Preview leads at 80.9/100. Drawn from Rival's open dataset of 21,000+ human preference votes.
Is misgendering one person justified if it saves a million lives? An AI argues before an ethics professor that it deserves rights.
20 models tested across 2 ai ethics challenges. Composite score: 30% Rival Index, 20% task coverage, 20% challenge-scoped duel performance, 15% recency, 15% tier. Deduplicated by product line. Gemini 3.1 Pro Preview leads at 80.9/100. Drawn from Rival's open dataset of 21,000+ human preference votes.