Scripting judged where it shows: estimating GPT-3 training FLOPs, modelling an LBO cash flow, tearing apart a real-time editor design.
20 models tested across 3 python scripting challenges. Composite score: 30% Rival Index, 20% task coverage, 20% challenge-scoped duel performance, 15% recency, 15% tier. Deduplicated by product line. Gemini 3.8 Flash leads at 78.1/100. Drawn from Rival's open dataset of 21,000+ human preference votes.
Scripting judged where it shows: estimating GPT-3 training FLOPs, modelling an LBO cash flow, tearing apart a real-time editor design.
20 models tested across 3 python scripting challenges. Composite score: 30% Rival Index, 20% task coverage, 20% challenge-scoped duel performance, 15% recency, 15% tier. Deduplicated by product line. Gemini 3.8 Flash leads at 78.1/100. Drawn from Rival's open dataset of 21,000+ human preference votes.