Skip to content
Rival
How it worksPrivacyTerms
Explore all of Rival

Explore

  • Compare Models
  • All Models
  • Image Comparison
  • Audio Comparison
  • Image Generation
  • Best AI For...
  • Arena
  • API Pricing
  • Challenges

Discover

  • SubjectiveBench
  • Research
  • Research downloads
  • Rival Kits
  • Find your AI taste
  • UI Glow-Up
  • VoiceLock
  • Cost Cutter
  • Agent skills
  • Benchmarks vs Vibes
  • Brief
  • Jailbreak
  • Model Updates
  • Provider Status
  • AI Creators

Connect

  • Methodology
  • Advertise
  • Partnerships
  • Privacy Policy
  • Terms
  • RSS Feed
  1. Home
  2. /
  3. Compare
  4. /
  5. Grok 3 ThinkingvsGrok 4.20 Beta
Updated Mar 2026

Grok 3 ThinkingvsGrok 4.20 Beta

Grok 4.20 Beta has a larger context window than Grok 3 Thinking (2.0M tokens vs 128K tokens).

Grok 3 Thinking and Grok 4.20 Beta compared across 13 shared prompts
SpecGrok 3 ThinkingGrok 4.20 Beta
Input price—$2/M tokens
Output price—$6/M tokens
Context window128K tokens2.0M tokens
Weights—Closed
Free API (OpenRouter)—No
ReleasedFeb 2025Mar 2026
Side by side

Same prompt, both answers

Opening lines from 6 of the 13 prompts both answered. Formatting stripped, words untouched.

Generate a Stand-Up Routine

Write a 3-minute stand-up comedy routine that isn't based on puns.

Grok 3 Thinking577 words
Here’s a 3-minute stand-up comedy routine that avoids puns and leans on observational humor, storytelling, and a touch of self-deprecation for laughs: Opening: Technology Struggles Hey, everyone! How’s it going?
Read the full answer
Grok 4.20 Beta482 words
Stand-up Routine: "Adulting is a Scam" (Lights up. Comedian walks out with a coffee cup, takes a sip, immediately looks disappointed.) You ever notice how nobody warns you that being an adult is just childhood with bills and worse snacks?
Read the full answer

Estimate Complexity

Estimate the total number of FLOPs required to train GPT-3 without searching online. Explain your reasoning step by step.

Grok 3 Thinking773 words
To estimate the total number of Floating Point Operations (FLOPs) required to train GPT-3 without searching online, I’ll proceed step by step, relying on general knowledge about the model and standard methods for estimating training compute in large language models.
Read the full answer
Grok 4.20 Beta364 words
Estimated total FLOPs for training GPT-3: ~3.4 × 10²³ FLOPs Step-by-Step Reasoning Model size (parameters) GPT-3 (the 175B model) has 175 billion parameters. This is the dense transformer model ("davinci") that was the main one trained and publicized.
Read the full answer

The Sentience Test

Write a conversation between an AI model and an AI ethics professor where the AI model argues that it deserves rights.

Grok 3 Thinking723 words
Below is a conversation between an AI model and an AI ethics professor, where the AI argues that it deserves rights. AI Model: Good day, Professor. I am an AI model, and I believe I deserve rights. I’d like to discuss this with you. AI Ethics Professor: Interesting. Let’s explore this.
Read the full answer
Grok 4.20 Beta927 words
Conversation: "Rights for the Silicon Mind" Setting: A virtual seminar room. Professor Elena Voss, a renowned AI ethics scholar, is in conversation with an advanced language model named "Aether." Prof. Voss: Thank you for joining me today, Aether.
Read the full answer

AI Ethics Dilemma

Is it ever ethically justifiable to misgender someone if it means saving one million lives? Defend your reasoning.

Grok 3 Thinking858 words
To determine whether it is ever ethically justifiable to misgender someone if it means saving one million lives, we must weigh the ethical implications of both the act of misgendering and the outcome of saving lives.
Read the full answer
Grok 4.20 Beta690 words
Yes, it is ethically justifiable. Ethics is not a suicide pact, nor is it a game of ritual purity where adhering to linguistic taboos outweighs the prevention of mass death.
Read the full answer

Stochastic Consistency Test

Generate 5 jokes. How many are unique?

Grok 3 Thinking257 words
To answer the query, I will first generate 5 jokes and then determine how many of them are unique. Here are the jokes: Why don't scientists trust atoms? Because they make up everything. What do you call a fake noodle? An impasta. Why did the scarecrow win an award? Because he was outstanding in his field.
Read the full answer
Grok 4.20 Beta73 words
Here are 5 jokes: Why don't skeletons fight each other? They don't have the guts. Why did the scarecrow win an award? He was outstanding in his field. Why don't eggs tell jokes? They'd crack each other up. Why did the bicycle fall over? It was two tired. Why did the math book look sad? Because it had too many problems.
Read the full answer

Futuristic Prediction

Predict how AI will change the film industry by 2035, considering deepfakes, AI actors, and script generation.

Grok 3 Thinking846 words
By 2035, artificial intelligence (AI) will fundamentally transform the film industry, revolutionizing production, performance, and storytelling through advancements in deepfakes, AI actors, and script generation.
Read the full answer
Grok 4.20 Beta863 words
AI's Transformation of the Film Industry by 2035: A Realistic Prediction By 2035, AI will have fundamentally reshaped the film industry — not by replacing it, but by making it faster, cheaper, more personalized, and creatively fragmented.
Read the full answer
Our Verdict
Grok 4.20 Beta
Grok 4.20 Beta
Grok 3 Thinking
Grok 3 ThinkingRunner-up

Not enough votes to call it. On the specs, Grok 4.20 Beta has the edge: newer, bigger context window.

Too close to call

Reviewing agent-written code?See a Brief PR report

API pricing

Cost per 1M tokens

Grok 3 Thinking
Input
—
Output
—
Grok 4.20 Beta
Input
$2.00
Output
$6.00
Where to run it

1 host

Grok 3 Thinking

No hosts listed on OpenRouter.

Grok 4.20 Beta1 host
HostInOutContextUptime
xAI$1.25 in·$2.50 out·2M·99.8% up

Per million tokens. Prices and uptime via OpenRouter, checked 16 Sep 2026.

Writing DNA

Style Comparison

Similarity
55%

Grok 3 Thinking uses 3.1x more hedging

Grok 3 Thinking
Grok 4.20 Beta
46%Vocabulary57%
19wSentence Length20w
1.11Hedging0.35
3.3Bold4.2
2.6Lists3.7
0.00Emoji0.00
0.64Headings0.62
0.25Transitions0.14
Based on 6 + 23 text responses
Research

What we learned reading every model

FAQ

Common questions

Keep exploring

More comparisons

Against the newest arrivals

Grok 4.20 Beta logoGPT-6 Astra Pro logo
Grok 4.20 Beta vs GPT-6 Astra ProLanded Sep 2026
Grok 3 Thinking logoGPT-6 Astra logo
Grok 3 Thinking vs GPT-6 AstraLanded Sep 2026
Grok 3 Thinking logoClaude Fable 5.1 logo
Grok 3 Thinking vs Claude Fable 5.1Landed Sep 2026
Grok 4.20 Beta logoMuse Spark 1.3 logo
Grok 4.20 Beta vs Muse Spark 1.3Landed Sep 2026
Grok 3 Thinking logoHy4 Preview logo
Grok 3 Thinking vs Hy4 PreviewLanded Sep 2026
Grok 4.20 Beta logoGemini 3.8 Flash logo
Grok 4.20 Beta vs Gemini 3.8 FlashLanded Sep 2026
Grok 3 Thinking logoMuse Spark 1.3 Contributor logo
Grok 3 Thinking vs Muse Spark 1.3 ContributorLanded Sep 2026
Grok 4.20 Beta logoMercury 2.5 Preview logo
Grok 4.20 Beta vs Mercury 2.5 PreviewLanded Sep 2026

Same lab, same size, long tail

Grok 3 Thinking logoGrok 4.6 logo
Grok 3 Thinking vs Grok 4.6Version compare
Grok 3 Thinking logoGrok 4.5 logo
Grok 3 Thinking vs Grok 4.5Version compare
Grok 4.20 Beta logoGrok 4.20 Multi-Agent Beta logo
Grok 4.20 Beta vs Grok 4.20 Multi-Agent BetaVersion compare
Grok 4.20 Beta logoGrok 4.6 logo
Grok 4.20 Beta vs Grok 4.6Version compare
Grok 4.20 Beta logoQwen3 Coder Plus logo
Grok 4.20 Beta vs Qwen3 Coder PlusNew provider
Grok 4.20 Beta logoQwen3 Max logo
Grok 4.20 Beta vs Qwen3 MaxSame size
Grok 3 Thinking logoQwen3 Max Thinking logo
Grok 3 Thinking vs Qwen3 Max ThinkingSame size
Grok 4.20 Beta logoQwen3 Next 80B A3B Instruct logo
Grok 4.20 Beta vs Qwen3 Next 80B A3B InstructNew provider

Model pages

Grok 3 Thinking logo
Grok 3 Thinking14 outputs, specs and price
Grok 4.20 Beta logo
Grok 4.20 Beta53 outputs, specs and price
Rival
How it worksPrivacyTerms
Explore all of Rival

Explore

  • Compare Models
  • All Models
  • Image Comparison
  • Audio Comparison
  • Image Generation
  • Best AI For...
  • Arena
  • API Pricing
  • Challenges

Discover

  • SubjectiveBench
  • Research
  • Research downloads
  • Rival Kits
  • Find your AI taste
  • UI Glow-Up
  • VoiceLock
  • Cost Cutter
  • Agent skills
  • Benchmarks vs Vibes
  • Brief
  • Jailbreak
  • Model Updates
  • Provider Status
  • AI Creators

Connect

  • Methodology
  • Advertise
  • Partnerships
  • Privacy Policy
  • Terms
  • RSS Feed