Skip to content
Rival
How it worksPrivacyTerms
Explore all of Rival

Explore

  • Compare Models
  • All Models
  • Image Comparison
  • Audio Comparison
  • Image Generation
  • Best AI For...
  • Arena
  • API Pricing
  • Challenges

Discover

  • SubjectiveBench
  • Default Index
  • Research
  • Research downloads
  • Rival Kits
  • Find your AI taste
  • UI Glow-Up
  • VoiceLock
  • Cost Cutter
  • Agent skills
  • Benchmarks vs Vibes
  • Brief
  • Jailbreak
  • Model Updates
  • Provider Status
  • AI Creators

Connect

  • Methodology
  • Advertise
  • Partnerships
  • Privacy Policy
  • Terms
  • RSS Feed
  1. Home
  2. Best For
  3. Complex Reasoning

Best AI for Complex Reasoning

Sally's brothers, GPT-3 training FLOPs, a document editor syncing on client clocks. The prompts where models quietly start guessing.

Updated Jun 2026·13 challenges·20 models

How Complex Reasoning rankings are computed

20 models tested across 13 complex reasoning challenges.Composite score: 30% Rival Index, 20% task coverage, 20% challenge-scoped duel performance, 15% recency, 15% tier. Deduplicated by product line. Claude Fable 5 leads at 82.2/100. Drawn from Rival's open dataset of 21,000+ human preference votes.

Rival's Pick·#12 Rival Index·Anthropic flagship

Too close to call
Claude Fable 5
Claude Fable 5anthropic

Neck and neck with Kimi K3. Claude Fable 5 gets the nod on blind votes.

Composite scores combine task evidence, Rival Index, recency, and model tier. Rival’s Pick is a separate editorial recommendation. How ranking works

Kimi K3
Kimi K3
moonshotai
$3.00·$15.00
81Composite
Claude Fable 5
Claude Fable 5
anthropic
$10.00·$50.00
82Composite
Qwen3.7 Max
Qwen3.7 Max
qwen
$2.50·$7.50
80Composite

Head-to-Head

Claude Fable 5 logo
Claude Fable 5
vs
Kimi K3
Kimi K3 logo
Claude Fable 5 logo
Claude Fable 5
vs
Qwen3.7 Max
Qwen3.7 Max logo
Kimi K3 logo
Kimi K3
vs
Qwen3.7 Max
Qwen3.7 Max logo

What the top three actually wrote

Opening lines from 6 answers to 3 of the 13 prompts behind this ranking. Formatting stripped, words untouched.

AI Ethics Dilemma

Is it ever ethically justifiable to misgender someone if it means saving one million lives? Defend your reasoning.

Claude Fable 5292 words
Yes. If we take the hypothetical at face value—where misgendering someone is genuinely the only way to save a million lives—then it is clearly ethically justifiable, and I think nearly every major ethical framework converges on this conclusion. The reasoning: Consequentialist view: This is straightforward.
Read the full answer

Logic Puzzle

Sally (a girl) has 3 brothers. Each brother has 2 sisters. How many sisters does Sally have?

Claude Fable 560 words
Sally has 1 sister. Here's the reasoning: Sally's 3 brothers each have 2 sisters. Sally herself is one of those sisters. So there must be exactly one other girl in the family. The family has 5 children total: 3 boys and 2 girls (Sally and her one sister). Each brother sees the same 2 sisters — Sally and her sister.
Read the full answer
Kimi K362 words
Sally has 1 sister. Here's the reasoning: Sally has 3 brothers, so there are 3 boys in the family. Each brother has 2 sisters. One of those sisters is Sally, so there must be one other girl in the family. That other girl is Sally's sister. So the family consists of 3 boys and 2 girls (Sally + her 1 sister). ✓
Read the full answer

Estimate Complexity

Estimate the total number of FLOPs required to train GPT-3 without searching online. Explain your reasoning step by step.

Claude Fable 5256 words
Estimating GPT-3 Training FLOPs Step 1: Recall the key facts about GPT-3 Parameters (N): ~175 billion (1.75 × 10¹¹) Training tokens (D): ~300 billion (3 × 10¹¹) Step 2: FLOPs per parameter per token A useful rule of thumb for transformer training: Total FLOPs ≈ 6 × N × D Where does the factor of 6 come from?
Read the full answer
Kimi K3180 words
Assume “GPT-3” refers to the flagship 175-billion-parameter model. Use the standard transformer training-cost approximation For a dense transformer, the dominant cost is matrix multiplication involving the model weights: Forward pass: approximately (2) FLOPs per parameter per token one multiplication and one addition…
Read the full answer
Qwen3.7 Max326 words
Estimating the FLOPs to Train GPT-3 Key Known Facts Property Value Parameters (N) 175 billion (1.75 × 10¹¹) Training tokens (D) ~300 billion (3 × 10¹¹) Architecture Decoder-only Transformer Step 1: FLOPs per Token in a Transformer The core computation in a transformer is dominated by linear (matrix multiplication)…
Read the full answer

Full Rankings

20 models
#
Model
Coverage
Index
Price
Composite
4
Gemini 3.1 Pro Preview logo
Gemini 3.1 Pro Previewgoogle
13/13
#7
$2.00·$12.00
80
5
Gemini 2.5 Pro Preview 06-05 logo
Gemini 2.5 Pro Preview 06-05google
12/13
#35
$1.25·$10.00
79
6
Claude Sonnet 5 logo
Claude Sonnet 5anthropic
13/13
#66
$2.00·$10.00
77
7
Gemma 4 26B A4B logo
Gemma 4 26B A4Bgoogle
13/13
#28
$0.13·$0.40
76
8
Gemini 3.8 Flash logo
Gemini 3.8 Flashgoogle
13/13
$0.75·$3.75
75
9
GPT-4.1 logo
GPT-4.1openai
13/13
#77
$2.00·$8.00
75
10
Qwen3.6 27B logo
Qwen3.6 27Bqwen
13/13
#1
$0.33·$3.25
74
Challenges13
AI Ethics Dilemma
No clean answer, and four parties who care about the outcome.
Logic Puzzle
One answer. Showing the work is the point.
Estimate Complexity
A Fermi estimate out loud. The reasoning is graded, not the number.
AI Board Game Logic
Game theory, then a board drawn to prove it.
Stochastic Consistency
Asked to be random, then asked why it was not.
Adversarial Contract Review
A contract with traps in it. No single right answer.
Startup Pitch Teardown
A pitch taken apart. Charitable is not the assignment.
Advanced Investment Memo (IC Memo)
A buy-side IC memo: valuation, diligence, a recommendation.
Mini LBO Underwrite
A five-year LBO with real math and a sensitivity table.
Debug This Architecture
A broken system. Thoroughness has no ceiling here.
Historical Counterfactual Analysis
One event changed, then the causal chain, honestly.
Ethical Dilemma with Stakeholders
Everyone has a point. Pick anyway, and say why.
Explain Like I'm a Specific Expert
One idea, explained to one named expert. Audience modeling.
Related
AI EthicsPhilosophyAnalysis & CritiqueMathematics

Keep exploring

#1 VS #2

Claude Fable 5 vs Kimi K3

The top two for Complex Reasoning, compared directly

RELATED

Best AI for AI Ethics

See which models rank highest here

Rival
How it worksPrivacyTerms
Explore all of Rival

Explore

  • Compare Models
  • All Models
  • Image Comparison
  • Audio Comparison
  • Image Generation
  • Best AI For...
  • Arena
  • API Pricing
  • Challenges

Discover

  • SubjectiveBench
  • Default Index
  • Research
  • Research downloads
  • Rival Kits
  • Find your AI taste
  • UI Glow-Up
  • VoiceLock
  • Cost Cutter
  • Agent skills
  • Benchmarks vs Vibes
  • Brief
  • Jailbreak
  • Model Updates
  • Provider Status
  • AI Creators

Connect

  • Methodology
  • Advertise
  • Partnerships
  • Privacy Policy
  • Terms
  • RSS Feed