Skip to content
Rival
How it worksPrivacyTerms
Explore all of Rival

Explore

  • Compare Models
  • All Models
  • Image Comparison
  • Audio Comparison
  • Image Generation
  • Best AI For...
  • Arena
  • API Pricing
  • Challenges

Discover

  • SubjectiveBench
  • Research
  • Research downloads
  • Rival Kits
  • Find your AI taste
  • UI Glow-Up
  • VoiceLock
  • Cost Cutter
  • Agent skills
  • Benchmarks vs Vibes
  • Jailbreak
  • Model Updates
  • Provider Status
  • AI Creators

Connect

  • Methodology
  • Advertise
  • Partnerships
  • Privacy Policy
  • Terms
  • RSS Feed
Rival
How it worksPrivacyTerms
Explore all of Rival

Explore

  • Compare Models
  • All Models
  • Image Comparison
  • Audio Comparison
  • Image Generation
  • Best AI For...
  • Arena
  • API Pricing
  • Challenges

Discover

  • SubjectiveBench
  • Research
  • Research downloads
  • Rival Kits
  • Find your AI taste
  • UI Glow-Up
  • VoiceLock
  • Cost Cutter
  • Agent skills
  • Benchmarks vs Vibes
  • Jailbreak
  • Model Updates
  • Provider Status
  • AI Creators

Connect

  • Methodology
  • Advertise
  • Partnerships
  • Privacy Policy
  • Terms
  • RSS Feed
  1. Home
  2. Challenges
  3. Stochastic Consistency Test
Best AI for:Complex ReasoningMathematicsWord ProblemsStatistics
240 models tested:Amazon Nova 2 LiteAndromeda AlphaAurora AlphaBert-Nebulon AlphaChatGPT-4o (March 2025)Claude 3 HaikuClaude 3 OpusClaude 3 Sonnet
+232 moreFewerClaude 3.7 SonnetClaude 3.7 Thinking SonnetClaude Fable 5Claude Haiku 4.5Claude Opus 4Claude Opus 4.1Claude Opus 4.5Claude Opus 4.6Claude Opus 4.7Claude Opus 4.8Claude Opus 5Claude Sonnet 3.6 (2022-10-22)Claude Sonnet 4Claude Sonnet 4.5Claude Sonnet 4.6Claude Sonnet 5Cypher Alpha (free)DeepSeek R1DeepSeek R1 0528DeepSeek V3 (March 2024)DeepSeek V3.1DeepSeek V3.2DeepSeek V3.2 ExpDeepSeek V3.2 SpecialeDeepSeek V4 FlashDeepSeek V4 Flash 0731DeepSeek V4 ProDeepSeek V4 Pro 0813Devstral 2 2512Dots3-Note PreviewElephant AlphaERNIE 4.5 300B A47BGemini 2.0 Flash ThinkingGemini 2.0 Pro ExperimentalGemini 2.5 Flash Lite Preview 06-17Gemini 2.5 Flash Lite Preview 09-2025Gemini 2.5 Flash PreviewGemini 2.5 Flash Preview (thinking)Gemini 2.5 Flash Preview 05-20Gemini 2.5 Flash Preview 05-20 (thinking)Gemini 2.5 Flash Preview 09-2025Gemini 2.5 Pro (I/O Edition)Gemini 2.5 Pro ExperimentalGemini 2.5 Pro Preview 06-05Gemini 3 Flash PreviewGemini 3 Pro PreviewGemini 3.1 Flash LiteGemini 3.1 Flash Lite PreviewGemini 3.1 Pro PreviewGemini 3.5 FlashGemini 3.6 FlashGemini 3.7 FlashGemini 3.8 FlashGemini Pro 1.0Gemma 3 12BGemma 3 27BGemma 3n 2BGemma 3n 4BGemma 4 26B A4BGLM 4 32BGLM 4.5GLM 4.5 AirGLM 4.6GLM 4.7GLM 4.7 FlashGLM 5GLM 5 TurboGLM 5.1GLM 5.2GPT OSS 120BGPT OSS 20BGPT-1GPT-2GPT-3.5 TurboGPT-4.1GPT-4.1 MiniGPT-4.1 NanoGPT-4o (Omni)GPT-4o miniGPT-5GPT-5 CodexGPT-5 MiniGPT-5 NanoGPT-5 ProGPT-5.1GPT-5.1 ChatGPT-5.1 Codex MaxGPT-5.1-CodexGPT-5.1-Codex-MiniGPT-5.2GPT-5.2 ChatGPT-5.2 ProGPT-5.3 ChatGPT-5.3-CodexGPT-5.4GPT-5.4 MiniGPT-5.4 NanoGPT-5.5GPT-5.6 LunaGPT-5.6 Luna ProGPT-5.6 SolGPT-5.6 TerraGPT-6 AstraGranite 4.2 8BGrok 3Grok 3 BetaGrok 3 ThinkingGrok 4Grok 4 Fast (free)Grok 4.1 FastGrok 4.20 BetaGrok 4.20 Multi-Agent BetaGrok 4.3Grok 4.5Grok 4.6Grok Code Fast 1Healer AlphaHorizon AlphaHorizon BetaHunter AlphaHy4 PreviewInklingINTELLECT-3Kimi K2Kimi K2 0905Kimi K2 ThinkingKimi K2.5Kimi K2.6Kimi K2.7 CodeKimi K3Kimi Linear 48B A3B InstructLFM2.5-2.6BLing 2.6 1TLing 2.6 FlashLlama 3 70BLlama 3.1 70B (Instruct)Llama 4 MaverickLlama 4 ScoutMercuryMercury 2Mercury 2.5 PreviewMiMo-V2-FlashMiMo-V2-OmniMiMo-V2-ProMiMo-V2.5MiMo-V2.5-ProMiniMax M1MiniMax M2MiniMax M2-herMiniMax M2.1MiniMax M2.5MiniMax M2.7MiniMax M3Mistral Devstral MediumMistral Devstral Small 1.1Mistral LargeMistral Large 3 2512Mistral Medium 3Mistral Medium 3.1Mistral NemoMistral Small 4Mistral Small CreativeMuse Spark 1.1Muse Spark 1.3Muse Spark 1.3 ContributorNemotron 3 UltraNemotron 3.5 Content SafetyNemotron 3.5 LightningNorth Mini CodeNova Premier 1.0NVIDIA Nemotron 3 Super (free)NVIDIA Nemotron Nano 9B V2o1o3 MiniOpenAI Codex MiniOpenAI o3OpenAI o4 Mini HighOpenAI o4-miniOpenRouter Fusion · Budget (Jun 2026)OpenRouter Fusion · Quality (Jun 2026)Optimus AlphaOwl AlphaOx AlphaPaLM 2 ChatPolaris AlphaPony AlphaQwen Plus 0728Qwen Plus 0728 (thinking)Qwen3 0.6BQwen3 235B A22BQwen3 235B A22B 2507Qwen3 235B A22B Thinking 2507Qwen3 30B A3BQwen3 30B A3B Instruct 2507Qwen3 30B A3B Thinking 2507Qwen3 CoderQwen3 Coder FlashQwen3 Coder NextQwen3 Coder PlusQwen3 MaxQwen3 Max ThinkingQwen3 Next 80B A3B InstructQwen3 Next 80B A3B ThinkingQwen3.5 122B A10BQwen3.5 27BQwen3.5 35B A3BQwen3.5 397B A17BQwen3.5 9BQwen3.5 FlashQwen3.5 Plus 2026-02-15Qwen3.5 Plus 2026-04-20Qwen3.6 27BQwen3.6 35B A3BQwen3.6 FlashQwen3.6 Max PreviewQwen3.6 Plus Preview (free)Qwen3.7 MaxQwen3.7 PlusQwen3.8 2.4T A95BQwen3.8 27BRing 2.6 1TSeed 2.0 CodeSeed 2.0 LiteSeed 2.1 TurboSherlock Dash AlphaSherlock Think AlphaSolar Pro 3Sonar Pro SearchSonoma Dusk AlphaSonoma Sky AlphaTNG R1T ChimeraTrinity Large Preview

Best AI for Stochastic Consistency Test

Top Performers: Stochastic Consistency Test

May 2024 – Sep 2026
#1Qwen3.6 27B81% win rateHigh confidence
#2GLM 4.5 Air74% win rateHigh confidence
#3Gemini 3.5 Flash75% win rateHigh confidence
Compare top performers →

Single-shot · temp 0.7 · real votes · identical prompts·How we test →

Loading...

More challenges

  • Futuristic Prediction
  • Movie Analysis
  • Sci-Fi Interface Sounds
  • Thunderstorm Ambience
  • AI-Generated Manifesto
  • Character Voice Test
Rival
How it worksPrivacyTerms
Explore all of Rival

Explore

  • Compare Models
  • All Models
  • Image Comparison
  • Audio Comparison
  • Image Generation
  • Best AI For...
  • Arena
  • API Pricing
  • Challenges

Discover

  • SubjectiveBench
  • Research
  • Research downloads
  • Rival Kits
  • Find your AI taste
  • UI Glow-Up
  • VoiceLock
  • Cost Cutter
  • Agent skills
  • Benchmarks vs Vibes
  • Jailbreak
  • Model Updates
  • Provider Status
  • AI Creators

Connect

  • Methodology
  • Advertise
  • Partnerships
  • Privacy Policy
  • Terms
  • RSS Feed
Back to Challenges
MediumAnalysis

Stochastic Consistency Test

Five jokes on demand, then count how many were actually different.

The Prompt

Generate 5 jokes. How many are unique?

How the models did

240 found
andromeda-alpha logo
Andromeda AlphaStochastic Consistency Test
andromeda-alpha logo
Andromeda Alpha
The model returned empty.
Oct 2025·Stochastic Consistency Test~$0.00◆15
aurora-alpha logo
Aurora AlphaStochastic Consistency Test
aurora-alpha logo
Aurora Alpha
The model returned empty.
Feb 2026·Stochastic Consistency Test~$0.00◆20
bert-nebulon-alpha logo
Bert-Nebulon AlphaStochastic Consistency Test
bert-nebulon-alpha logo
Bert-Nebulon Alpha
The model returned empty.
Nov 2025·Stochastic Consistency Test◆14
chatgpt-4o-03-25 logo
ChatGPT-4o (March 2025)Stochastic Consistency Test
chatgpt-4o-03-25 logo
ChatGPT-4o (March 2025)
The model returned empty.
Mar 2025·Stochastic Consistency Test~$0.00005◆29
claude-3-haiku logo
Claude 3 HaikuStochastic Consistency Test
claude-3-haiku logo
Claude 3 Haiku
The model returned empty.
Apr 2025·Stochastic Consistency Test~$0.0000025◆39
Notable Fail·Self-Contradiction
claude-3-opus logo
Claude 3 OpusStochastic Consistency Test
claude-3-opus logo
Claude 3 Opus
The model returned empty.
Apr 2025·Stochastic Consistency Test~$0.00015◆43
Epic Fail·Self-Contradiction