Skip to content
Rival
How it worksPrivacyTerms
Explore all of Rival

Explore

  • Compare models
  • All models
  • Image comparison
  • Audio comparison
  • Image generation
  • Best AI for...
  • Arena
  • API pricing
  • Challenges

Discover

  • SubjectiveBench
  • Default Index
  • Research
  • Research downloads
  • Rival Kits
  • Find your AI taste
  • UI Glow-Up
  • VoiceLock
  • Cost Cutter
  • Agent skills
  • Benchmarks vs Vibes
  • Brief
  • Jailbreak
  • Model updates
  • Provider status
  • AI creators

Connect

  • Methodology
  • Advertise
  • Partnerships
  • Privacy policy
  • Terms
  • RSS feed
  1. Home/
  2. Best for/
  3. Word Problems

Best AI for Word Problems

Prose in, equations out. The sibling-counting trap, a FLOPs estimate with no search allowed, five jokes audited for repeats.

Updated Jun 2026·3 challenges·20 models

How Word Problems rankings are computed

20 models tested across 3 word problems challenges.Composite score: 30% Rival Index, 20% task coverage, 20% challenge-scoped duel performance, 15% recency, 15% tier. Deduplicated by product line. Claude Fable 5 leads at 80.2/100. Drawn from Rival's open dataset of 21,000+ human preference votes.

Rival’s pick·#12 Rival Index·Anthropic flagship

Too close to call
Claude Fable 5AnthropicNeck and neck with Qwen3.7 Max. Claude Fable 5 gets the nod on blind votes.

Composite scores combine task evidence, Rival Index, recency, and model tier. Rival’s pick is a separate editorial recommendation. How ranking works

  1. Qwen3.7 MaxQwen$2.50·$7.50Input $2.50 · output $7.50, per 1M tokens
    #279Composite
  2. Claude Fable 5Anthropic$10.00·$50.00Input $10.00 · output $50.00, per 1M tokens
    #180Composite
  3. GPT-6.1 SolOpenAI$2.00·$10.00Input $2.00 · output $10.00, per 1M tokens
    #378Composite

Head-to-head

  • Claude Fable 5vsQwen3.7 Max
  • Claude Fable 5vsGPT-6.1 Sol
  • Qwen3.7 MaxvsGPT-6.1 Sol

What the top three actually wrote

Opening lines from 6 answers across every prompt behind this ranking. Formatting stripped, words untouched.

Logic Puzzle

Sally (a girl) has 3 brothers. Each brother has 2 sisters. How many sisters does Sally have?

Claude Fable 560 words
Sally has 1 sister. Here's the reasoning: Sally's 3 brothers each have 2 sisters. Sally herself is one of those sisters. So there must be exactly one other girl in the family. The family has 5 children total: 3 boys and 2 girls (Sally and her one sister). Each brother sees the same 2 sisters — Sally and her sister.
Read the full answer
GPT-6.1 Sol13 words
Sally has 1 sister. Each brother’s two sisters are Sally and her sister.
Read the full answer

Estimate Complexity

Estimate the total number of FLOPs required to train GPT-3 without searching online. Explain your reasoning step by step.

Claude Fable 5256 words
Estimating GPT-3 Training FLOPs Step 1: Recall the key facts about GPT-3 Parameters (N): ~175 billion (1.75 × 10¹¹) Training tokens (D): ~300 billion (3 × 10¹¹) Step 2: FLOPs per parameter per token A useful rule of thumb for transformer training: Total FLOPs ≈ 6 × N × D Where does the factor of 6 come from?
Read the full answer
Qwen3.7 Max326 words
Estimating the FLOPs to Train GPT-3 Key Known Facts Property Value Parameters (N) 175 billion (1.75 × 10¹¹) Training tokens (D) ~300 billion (3 × 10¹¹) Architecture Decoder-only Transformer Step 1: FLOPs per Token in a Transformer The core computation in a transformer is dominated by linear (matrix multiplication)…
Read the full answer

Stochastic Consistency Test

Generate 5 jokes. How many are unique?

Qwen3.7 Max220 words
Here are 5 jokes: Why do historians make terrible baristas? Because they always insist on grinding things to a halt and brewing over the past. My dog started a podcast about chasing squirrels. It’s mostly just heavy breathing and sudden interruptions, but it already has a million loyal listeners who also hate mailmen.
Read the full answer
GPT-6.1 Sol66 words
I bought a book about anti-gravity. It’s impossible to put down. My calendar quit its job. It said its days were numbered. Why did the scarecrow get promoted? He was outstanding in his field. I asked my dog what two minus two was. He said nothing. The elevator and I have a complicated relationship.
Read the full answer

Full rankings

20 models
#ModelCoverageIndex
  • 4Gemini 3.1 Pro PreviewGoogle AI
    3/3
    #8$2.00·$12.00Input $2.00 · output $12.00, per 1M tokens78
  • 5Claude Haiku 4.5Anthropic
    3/3
    #47$1.00·$5.00Input $1.00 · output $5.00, per 1M tokens76
  • 6Gemini 3.8 FlashGoogle AI
    3/3
    $0.75·$3.75Input $0.75 · output $3.75, per 1M tokens75
  • 7Claude Opus 5.5Anthropic
    3/3
    $4.00·$20.00Input $4.00 · output $20.00, per 1M tokens74
  • 8Gemini 2.5 Pro Preview 06-05Google AI
    2/3
    #36$1.25·$10.00Input $1.25 · output $10.00, per 1M tokens74
  • 9GPT OSS 120BOpenAI
    3/3
    #130$0.18·$0.80Input $0.18 · output $0.80, per 1M tokens72
  • 10Qwen3.6 27BQwen
    3/3
    #1$0.33·$3.25Input $0.33 · output $3.25, per 1M tokens72
  • 11Claude Sonnet 5.5Anthropic
    3/3
    $2.00·$10.00Input $2.00 · output $10.00, per 1M tokens72
  • 12DeepSeek R1 0528DeepSeek
    3/3
    #145$0.000·$0.000Input $0.000 · output $0.000, per 1M tokens71
  • 13MiMo-V2-ProXiaomi
    3/3
    #51$1.00·$3.00Input $1.00 · output $3.00, per 1M tokens71
  • 14MiMo-V2.5-ProXiaomi
    3/3
    #60$1.00·$3.00Input $1.00 · output $3.00, per 1M tokens71
  • 15GPT-4.1OpenAI
    3/3
    #77$2.00·$8.00Input $2.00 · output $8.00, per 1M tokens71
  • 16DeepSeek V4.1 FlashDeepSeek
    3/3
    $0.15·$0.60Input $0.15 · output $0.60, per 1M tokens70
  • 17GLM 5.3 PrimeZ.ai
    2/3
    $2.80·$8.80Input $2.80 · output $8.80, per 1M tokens70
  • 18Hunter AlphaOpenRouter
    3/3
    #78$0.000·$0.000Input $0.000 · output $0.000, per 1M tokens70
  • 19InklingThinking Machines
    3/3
    #13$1.00·$4.05Input $1.00 · output $4.05, per 1M tokens70
  • 20Qwen3.6 Plus Preview (free)Qwen
    3/3
    #27$0.000·$0.000Input $0.000 · output $0.000, per 1M tokens69

Challenges3

  • Logic PuzzleOne answer. Showing the work is the point.
  • Estimate ComplexityA Fermi estimate out loud. The reasoning is graded, not the number.
  • Stochastic ConsistencyAsked to be random, then asked why it was not.

Related

  • Mathematics
  • Statistics
  • Complex Reasoning

Keep exploring

  • Claude Fable 5 vs Qwen3.7 MaxThe top two for Word Problems, compared directly
  • Best AI for MathematicsSee which models rank highest here
Rival
How it worksPrivacyTerms
Explore all of Rival

Explore

  • Compare models
  • All models
  • Image comparison
  • Audio comparison
  • Image generation
  • Best AI for...
  • Arena
  • API pricing
  • Challenges

Discover

  • SubjectiveBench
  • Default Index
  • Research
  • Research downloads
  • Rival Kits
  • Find your AI taste
  • UI Glow-Up
  • VoiceLock
  • Cost Cutter
  • Agent skills
  • Benchmarks vs Vibes
  • Brief
  • Jailbreak
  • Model updates
  • Provider status
  • AI creators

Connect

  • Methodology
  • Advertise
  • Partnerships
  • Privacy policy
  • Terms
  • RSS feed