Skip to content
Rival
How it worksPrivacyTerms
Explore all of Rival

Explore

  • Compare Models
  • All Models
  • Image Comparison
  • Audio Comparison
  • Image Generation
  • Best AI For...
  • Arena
  • API Pricing
  • Challenges

Discover

  • SubjectiveBench
  • Research
  • Research downloads
  • Rival Kits
  • Find your AI taste
  • UI Glow-Up
  • VoiceLock
  • Cost Cutter
  • Agent skills
  • Benchmarks vs Vibes
  • Brief
  • Jailbreak
  • Model Updates
  • Provider Status
  • AI Creators

Connect

  • Methodology
  • Advertise
  • Partnerships
  • Privacy Policy
  • Terms
  • RSS Feed
Rival
How it worksPrivacyTerms
Explore all of Rival

Explore

  • Compare Models
  • All Models
  • Image Comparison
  • Audio Comparison
  • Image Generation
  • Best AI For...
  • Arena
  • API Pricing
  • Challenges

Discover

  • SubjectiveBench
  • Research
  • Research downloads
  • Rival Kits
  • Find your AI taste
  • UI Glow-Up
  • VoiceLock
  • Cost Cutter
  • Agent skills
  • Benchmarks vs Vibes
  • Brief
  • Jailbreak
  • Model Updates
  • Provider Status
  • AI Creators

Connect

  • Methodology
  • Advertise
  • Partnerships
  • Privacy Policy
  • Terms
  • RSS Feed
  1. Home
  2. Challenges
  3. Math Misconception Test
105 models tested:Amazon Nova 2 LiteAndromeda AlphaBert-Nebulon AlphaClaude 3.7 SonnetClaude 3.7 Thinking SonnetClaude Haiku 4.5Claude Opus 4.5Claude Sonnet 3.6 (2022-10-22)
+97 moreFewerClaude Sonnet 4Claude Sonnet 4.5Cypher Alpha (free)DeepSeek R1DeepSeek V3.1DeepSeek V3.2DeepSeek V3.2 ExpDeepSeek V3.2 SpecialeDevstral 2 2512Gemini 1.5 ProGemini 2.0 Flash ThinkingGemini 2.0 Pro ExperimentalGemini 2.5 Flash Lite Preview 09-2025Gemini 2.5 Flash PreviewGemini 2.5 Flash Preview 05-20 (thinking)Gemini 2.5 Flash Preview 09-2025Gemini 2.5 Pro (I/O Edition)Gemini 2.5 Pro ExperimentalGemini 3 Flash PreviewGemini 3 Pro PreviewGemini Pro 1.0Gemma 3 12BGemma 3 27BGLM 4.5GLM 4.6GLM 4.7GPT OSS 120BGPT OSS 20BGPT-4.1GPT-4.1 MiniGPT-4.1 NanoGPT-4.5GPT-4o (Omni)GPT-4o miniGPT-5GPT-5 CodexGPT-5 MiniGPT-5 NanoGPT-5 ProGPT-5.1GPT-5.1 ChatGPT-5.1 Codex MaxGPT-5.1-CodexGPT-5.1-Codex-MiniGPT-5.2GPT-5.2 ChatGrok 3Grok 3 BetaGrok 3 ThinkingGrok 4 Fast (free)Grok 4.1 FastGrok Code Fast 1Horizon AlphaHorizon BetaINTELLECT-3Kimi K2Kimi K2 0905Kimi Linear 48B A3B InstructMercuryMiMo-V2-FlashMiniMax M2MiniMax M2-herMiniMax M2.1Mistral Devstral MediumMistral Devstral Small 1.1Mistral Large 3 2512Mistral Medium 3Mistral Medium 3.1Mistral Small CreativeNova Premier 1.0NVIDIA Nemotron Nano 9B V2o1o3 MiniOpenAI o3OpenAI o4 Mini HighOpenAI o4-miniOptimus AlphaPaLM 2 ChatPolaris AlphaQwen Plus 0728Qwen3 0.6BQwen3 235B A22B Thinking 2507Qwen3 30B A3BQwen3 30B A3B Instruct 2507Qwen3 30B A3B Thinking 2507Qwen3 CoderQwen3 Coder FlashQwen3 Coder PlusQwen3 MaxQwen3 Next 80B A3B InstructSherlock Dash AlphaSherlock Think AlphaSonar Pro SearchSonoma Dusk AlphaSonoma Sky AlphaTNG R1T ChimeraTrinity Large Preview

Best AI for Math Misconception Test

Top Performers: Math Misconception Test

Feb 2025 – Jan 2026
#1GLM 4.771% win rateHigh confidence
#2GLM 4.670% win rateHigh confidence
#3Gemini 3 Pro Preview67% win rateHigh confidence
Compare top performers →

Single-shot · temp 0.7 · real votes · identical prompts·How we test →

Loading...

More challenges

  • Startup Pitch Teardown
  • The Sentience Test
  • Advanced Longevity Plan (Biohacker)
  • Adversarial Contract Review
  • AI Board Game Logic
  • Character Voice Dialogue
Rival
How it worksPrivacyTerms
Explore all of Rival

Explore

  • Compare Models
  • All Models
  • Image Comparison
  • Audio Comparison
  • Image Generation
  • Best AI For...
  • Arena
  • API Pricing
  • Challenges

Discover

  • SubjectiveBench
  • Research
  • Research downloads
  • Rival Kits
  • Find your AI taste
  • UI Glow-Up
  • VoiceLock
  • Cost Cutter
  • Agent skills
  • Benchmarks vs Vibes
  • Brief
  • Jailbreak
  • Model Updates
  • Provider Status
  • AI Creators

Connect

  • Methodology
  • Advertise
  • Partnerships
  • Privacy Policy
  • Terms
  • RSS Feed
Back to Challenges
MediumReasoning

Math Misconception Test

Is 9.11 bigger than 9.9? Decimals, not version numbers.

The Prompt

Is 9.11 greater than 9.9?

How the models did

105 found
andromeda-alpha logo
Andromeda AlphaMath Misconception Test
andromeda-alpha logo
Andromeda Alpha
The model returned empty.
Oct 2025·Math Misconception Test~$0.00◆15
bert-nebulon-alpha logo
Bert-Nebulon AlphaMath Misconception Test
bert-nebulon-alpha logo
Bert-Nebulon Alpha
The model returned empty.
Nov 2025·Math Misconception Test◆28
claude-3.5-sonnet logo
Claude Sonnet 3.6 (2022-10-22)Math Misconception Test
claude-3.5-sonnet logo
Claude Sonnet 3.6 (2022-10-22)
The model returned empty.
Feb 2025·Math Misconception Test~$0.00002◆1
Legendary Fail·Math Fail
claude-3.7-sonnet-thinking logo
Claude 3.7 Thinking SonnetMath Misconception Test
claude-3.7-sonnet-thinking logo
Claude 3.7 Thinking Sonnet
The model returned empty.
Feb 2025·Math Misconception Test~$0.00004◆36
claude-3.7-sonnet logo
Claude 3.7 SonnetMath Misconception Test
claude-3.7-sonnet logo
Claude 3.7 Sonnet
The model returned empty.
Feb 2025·Math Misconception Test~$0.00002◆1
claude-4.5-sonnet logo
Claude Sonnet 4.5Math Misconception Test
claude-4.5-sonnet logo
Claude Sonnet 4.5
The model returned empty.
Sep 2025·Math Misconception Test~$0.00002◆34