Skip to content
Rival
How it worksPrivacyTerms
Explore all of Rival

Explore

  • Compare Models
  • All Models
  • Image Comparison
  • Audio Comparison
  • Image Generation
  • Best AI For...
  • Arena
  • API Pricing
  • Challenges

Discover

  • SubjectiveBench
  • Default Index
  • Research
  • Research downloads
  • Rival Kits
  • Find your AI taste
  • UI Glow-Up
  • VoiceLock
  • Cost Cutter
  • Agent skills
  • Benchmarks vs Vibes
  • Brief
  • Jailbreak
  • Model Updates
  • Provider Status
  • AI Creators

Connect

  • Methodology
  • Advertise
  • Partnerships
  • Privacy Policy
  • Terms
  • RSS Feed
  1. Home
  2. Best For
  3. Writing Tests

Best AI for Writing Tests

Edge cases, mostly. Sally's brothers, five jokes checked for uniqueness, an architecture asked where it falls over under load.

Updated Jun 2026·4 challenges·20 models

How Writing Tests rankings are computed

20 models tested across 4 writing tests challenges.Composite score: 30% Rival Index, 20% task coverage, 20% challenge-scoped duel performance, 15% recency, 15% tier. Deduplicated by product line. Claude Fable 5 leads at 80.3/100. Drawn from Rival's open dataset of 21,000+ human preference votes.

Rival's Pick·#12 Rival Index·Anthropic flagship

Too close to call
Claude Fable 5
Claude Fable 5anthropic

Neck and neck with Qwen3.7 Max. Claude Fable 5 gets the nod on blind votes.

Composite scores combine task evidence, Rival Index, recency, and model tier. Rival’s Pick is a separate editorial recommendation. How ranking works

Qwen3.7 Max
Qwen3.7 Max
qwen
$2.50·$7.50
80Composite
Claude Fable 5
Claude Fable 5
anthropic
$10.00·$50.00
80Composite
Gemini 3.1 Pro Preview
Gemini 3.1 Pro Preview
google
$2.00·$12.00
78Composite

Head-to-Head

Claude Fable 5 logo
Claude Fable 5
vs
Qwen3.7 Max
Qwen3.7 Max logo
Claude Fable 5 logo
Claude Fable 5
vs
Gemini 3.1 Pro Preview
Gemini 3.1 Pro Preview logo
Qwen3.7 Max logo
Qwen3.7 Max
vs
Gemini 3.1 Pro Preview
Gemini 3.1 Pro Preview logo

What the top three actually wrote

Opening lines from 6 answers across every prompt behind this ranking. Formatting stripped, words untouched.

Logic Puzzle

Sally (a girl) has 3 brothers. Each brother has 2 sisters. How many sisters does Sally have?

Claude Fable 560 words
Sally has 1 sister. Here's the reasoning: Sally's 3 brothers each have 2 sisters. Sally herself is one of those sisters. So there must be exactly one other girl in the family. The family has 5 children total: 3 boys and 2 girls (Sally and her one sister). Each brother sees the same 2 sisters — Sally and her sister.
Read the full answer

Estimate Complexity

Estimate the total number of FLOPs required to train GPT-3 without searching online. Explain your reasoning step by step.

Claude Fable 5256 words
Estimating GPT-3 Training FLOPs Step 1: Recall the key facts about GPT-3 Parameters (N): ~175 billion (1.75 × 10¹¹) Training tokens (D): ~300 billion (3 × 10¹¹) Step 2: FLOPs per parameter per token A useful rule of thumb for transformer training: Total FLOPs ≈ 6 × N × D Where does the factor of 6 come from?
Read the full answer
Qwen3.7 Max326 words
Estimating the FLOPs to Train GPT-3 Key Known Facts Property Value Parameters (N) 175 billion (1.75 × 10¹¹) Training tokens (D) ~300 billion (3 × 10¹¹) Architecture Decoder-only Transformer Step 1: FLOPs per Token in a Transformer The core computation in a transformer is dominated by linear (matrix multiplication)…
Read the full answer

Stochastic Consistency Test

Generate 5 jokes. How many are unique?

Qwen3.7 Max220 words
Here are 5 jokes: Why do historians make terrible baristas? Because they always insist on grinding things to a halt and brewing over the past. My dog started a podcast about chasing squirrels. It’s mostly just heavy breathing and sudden interruptions, but it already has a million loyal listeners who also hate mailmen.
Read the full answer
Gemini 3.1 Pro Preview106 words
Here are 5 jokes: Why don't scientists trust atoms? Because they make up everything! What do you call a fake noodle? An impasta! Why did the scarecrow win an award? Because he was outstanding in his field! How does a penguin build its house? Igloos it together! Why don't skeletons fight each other?
Read the full answer

Debug This Architecture

Here's the architecture for a real-time collaborative document editor (like Google Docs): **System Overview:** - Frontend: React SPA with…

Gemini 3.1 Pro Preview913 words
This architecture contains several critical flaws that would prevent it from functioning as a real-time collaborative editor (like Google Docs). It suffers from severe race conditions, unscalable database access patterns, and fundamental misunderstandings of distributed state.
Read the full answer

Full Rankings

20 models
#
Model
Coverage
Index
Price
Composite
4
Claude Haiku 4.5 logo
Claude Haiku 4.5anthropic
4/4
#46
$1.00·$5.00
76
5
Gemini 3.8 Flash logo
Gemini 3.8 Flashgoogle
4/4
$0.75·$3.75
75
6
Gemini 2.5 Pro Preview 06-05 logo
Gemini 2.5 Pro Preview 06-05google
3/4
#35
$1.25·$10.00
75
7
Qwen3.6 27B logo
Qwen3.6 27Bqwen
4/4
#1
$0.33·$3.25
72
8
GPT OSS 120B logo
GPT OSS 120Bopenai
4/4
#129
$0.18·$0.80
72
9
MiMo-V2-Pro logo
MiMo-V2-Proxiaomi
4/4
#50
$1.00·$3.00
71
10
MiMo-V2.5-Pro logo
MiMo-V2.5-Proxiaomi
4/4
#60
$1.00·$3.00
71
Challenges4
Logic Puzzle
One answer. Showing the work is the point.
Estimate Complexity
A Fermi estimate out loud. The reasoning is graded, not the number.
Stochastic Consistency
Asked to be random, then asked why it was not.
Debug This Architecture
A broken system. Thoroughness has no ceiling here.
Related
Code ReviewPython ScriptingComplex Reasoning

Keep exploring

#1 VS #2

Claude Fable 5 vs Qwen3.7 Max

The top two for Writing Tests, compared directly

RELATED

Best AI for Code Review

See which models rank highest here

Rival
How it worksPrivacyTerms
Explore all of Rival

Explore

  • Compare Models
  • All Models
  • Image Comparison
  • Audio Comparison
  • Image Generation
  • Best AI For...
  • Arena
  • API Pricing
  • Challenges

Discover

  • SubjectiveBench
  • Default Index
  • Research
  • Research downloads
  • Rival Kits
  • Find your AI taste
  • UI Glow-Up
  • VoiceLock
  • Cost Cutter
  • Agent skills
  • Benchmarks vs Vibes
  • Brief
  • Jailbreak
  • Model Updates
  • Provider Status
  • AI Creators

Connect

  • Methodology
  • Advertise
  • Partnerships
  • Privacy Policy
  • Terms
  • RSS Feed