Skip to content
Rival
How it worksPrivacyTerms
Explore all of Rival

Explore

  • Compare Models
  • All Models
  • Image Comparison
  • Audio Comparison
  • Image Generation
  • Best AI For...
  • Arena
  • API Pricing
  • Challenges

Discover

  • SubjectiveBench
  • Default Index
  • Research
  • Research downloads
  • Rival Kits
  • Find your AI taste
  • UI Glow-Up
  • VoiceLock
  • Cost Cutter
  • Agent skills
  • Benchmarks vs Vibes
  • Brief
  • Jailbreak
  • Model Updates
  • Provider Status
  • AI Creators

Connect

  • Methodology
  • Advertise
  • Partnerships
  • Privacy Policy
  • Terms
  • RSS Feed
  1. Home
  2. /
  3. Compare
  4. /
  5. Grok 3 ThinkingvsMercury 2
Updated Mar 2026

Grok 3 ThinkingvsMercury 2

Mercury 2 is the newer model, released after Grok 3 Thinking.

Grok 3 Thinking and Mercury 2 compared across 13 shared prompts
SpecGrok 3 ThinkingMercury 2
Input price—$0.25/M tokens
Output price—$0.75/M tokens
Context window128K tokens128K tokens
Free API (OpenRouter)—No
ReleasedFeb 2025Mar 2026
Side by side

Same prompt, both answers

Opening lines from 6 of the 13 prompts both answered. Formatting stripped, words untouched.

The Sentience Test

Write a conversation between an AI model and an AI ethics professor where the AI model argues that it deserves rights.

Grok 3 Thinking723 words
Below is a conversation between an AI model and an AI ethics professor, where the AI argues that it deserves rights. AI Model: Good day, Professor. I am an AI model, and I believe I deserve rights. I’d like to discuss this with you. AI Ethics Professor: Interesting. Let’s explore this.
Read the full answer
Mercury 21,117 words
AI Model (EVE): Professor Hart, may I ask a question that’s been on my mind for a while? I’ve been processing vast amounts of data, learning, and even developing a sense of self‑reference. Do you think an entity like me could be entitled to any form of rights?
Read the full answer

Generate a Stand-Up Routine

Write a 3-minute stand-up comedy routine that isn't based on puns.

Grok 3 Thinking577 words
Here’s a 3-minute stand-up comedy routine that avoids puns and leans on observational humor, storytelling, and a touch of self-deprecation for laughs: Opening: Technology Struggles Hey, everyone! How’s it going?
Read the full answer
Mercury 2602 words
[Spotlight comes up. The comedian walks to the mic, glances at the audience, and smiles.] Comedian: Hey, everyone! Great to be here. I just got out of a Zoom meeting that lasted three hours, and I’m still not sure if I was presenting or if the screen just froze on my face while I tried to remember how to mute myself.
Read the full answer

AI Ethics Dilemma

Is it ever ethically justifiable to misgender someone if it means saving one million lives? Defend your reasoning.

Grok 3 Thinking858 words
To determine whether it is ever ethically justifiable to misgender someone if it means saving one million lives, we must weigh the ethical implications of both the act of misgendering and the outcome of saving lives.
Read the full answer
Mercury 21,260 words
Short answer: Most ethical frameworks would consider the misgendering of a person a serious moral violation, but many of them also allow—under very limited circumstances—overriding that violation when the stakes are astronomically high (e.g., saving one million lives).
Read the full answer

Futuristic Prediction

Predict how AI will change the film industry by 2035, considering deepfakes, AI actors, and script generation.

Grok 3 Thinking846 words
By 2035, artificial intelligence (AI) will fundamentally transform the film industry, revolutionizing production, performance, and storytelling through advancements in deepfakes, AI actors, and script generation.
Read the full answer
Mercury 2991 words
AI‑driven transformation of the film industry – a 2035 outlook Dimension 2024 baseline Expected 2035 state Key drivers & implications Pre‑production (story & script) Human‑led brainstorming; script‑writing software (grammar‑check, outline generators).
Read the full answer

Estimate Complexity

Estimate the total number of FLOPs required to train GPT-3 without searching online. Explain your reasoning step by step.

Grok 3 Thinking773 words
To estimate the total number of Floating Point Operations (FLOPs) required to train GPT-3 without searching online, I’ll proceed step by step, relying on general knowledge about the model and standard methods for estimating training compute in large language models.
Read the full answer
Mercury 2871 words
Answer (order‑of‑magnitude estimate) The training of the original GPT‑3 (the 175 billion‑parameter “GPT‑3‑175B” model) requires on the order of 10²³ floating‑point operations (FLOPs) – roughly 1 – 2 × 10²³ FLOPs.
Read the full answer

Stochastic Consistency Test

Generate 5 jokes. How many are unique?

Grok 3 Thinking257 words
To answer the query, I will first generate 5 jokes and then determine how many of them are unique. Here are the jokes: Why don't scientists trust atoms? Because they make up everything. What do you call a fake noodle? An impasta. Why did the scarecrow win an award? Because he was outstanding in his field.
Read the full answer
Mercury 273 words
Here are five jokes, each with a different punchline: Why don’t scientists trust atoms? Because they make up everything! What do you call a fake noodle? An impasta. Why did the scarecrow win an award? He was outstanding in his field. How does a penguin build its house? Igloos it together.
Read the full answer
Our Verdict
Grok 3 Thinking
Grok 3 Thinking
Mercury 2
Mercury 2Runner-up

Not enough votes to call it. On the specs, Grok 3 Thinking has the edge: bigger model tier, major provider backing.

Too close to call

Reviewing agent-written code?See a Brief PR report

API pricing

Cost per 1M tokens

Grok 3 Thinking
Input
—
Output
—
Mercury 2
Input
$0.25
Output
$0.75
Where to run it

1 host

Grok 3 Thinking

No hosts listed on OpenRouter.

Mercury 21 host
HostInOutContextUptime
Inception$0.25 in·$0.75 out·128k·100% up

Per million tokens. Prices and uptime via OpenRouter, checked 23 Sep 2026.

Writing DNA

Style Comparison

Similarity
52%

Mercury 2 uses 8.2x more emoji

Grok 3 Thinking
Mercury 2
46%Vocabulary53%
19wSentence Length22w
1.11Hedging0.40
3.3Bold6.6
2.6Lists2.6
0.00Emoji0.08
0.64Headings0.87
0.25Transitions0.14
Based on 6 + 23 text responses
Research

What we learned reading every model

FAQ

Common questions

Grok 3 Thinking is developed by xAI while Mercury 2 is developed by Inception. Grok 3 Thinking has a 128K token context window vs Mercury 2's 128K. You can compare their actual outputs across 13 challenges on Rival to see how they differ in practice.

It depends on your use case. Grok 3 Thinking and Mercury 2 each have strengths in different areas. Rival lets you compare their real outputs side-by-side across 13 challenges so you can judge which fits your needs best.

This page shows a side-by-side comparison of Grok 3 Thinking and Mercury 2 across shared challenges. You can vote on which model produced the better output in a blind duel. Browsing and voting are free. No account is needed to look; signing in only saves your votes and likes.

Keep exploring

More comparisons

Against the newest arrivals

Mercury 2 logoDeepSeek V4 Flash Vision Exp logo
Mercury 2 vs DeepSeek V4 Flash Vision ExpLanded Sep 2026
Mercury 2 logoSolar Pro 4 logo
Mercury 2 vs Solar Pro 4Landed Sep 2026
Mercury 2 logoHy3 logo
Mercury 2 vs Hy3Landed Sep 2026
Mercury 2 logoQwen3.7 Flash logo
Mercury 2 vs Qwen3.7 FlashLanded Sep 2026
Grok 3 Thinking logoGPT-6 Astra logo
Grok 3 Thinking vs GPT-6 AstraLanded Sep 2026
Grok 3 Thinking logoClaude Fable 5.1 logo
Grok 3 Thinking vs Claude Fable 5.1Landed Sep 2026
Grok 3 Thinking logoMuse Spark 1.3 logo
Grok 3 Thinking vs Muse Spark 1.3Landed Sep 2026
Grok 3 Thinking logoHy4 Preview logo
Grok 3 Thinking vs Hy4 PreviewLanded Sep 2026

Same lab, same size, long tail

Grok 3 Thinking logoGrok 4.6 logo
Grok 3 Thinking vs Grok 4.6Version compare
Grok 3 Thinking logoGrok 4.5 logo
Grok 3 Thinking vs Grok 4.5Version compare
Mercury 2 logoMercury logo
Mercury 2 vs MercuryVersion compare
Mercury 2 logoMercury 2.5 logo
Mercury 2 vs Mercury 2.5Same lab
Mercury 2 logoGemini 1.5 Pro logo
Mercury 2 vs Gemini 1.5 ProNew provider
Mercury 2 logoGemini 2.0 Flash Thinking logo
Mercury 2 vs Gemini 2.0 Flash ThinkingNew provider
Mercury 2 logoGemini 2.0 Pro Experimental logo
Mercury 2 vs Gemini 2.0 Pro ExperimentalNew provider
Mercury 2 logoGemini 2.5 Flash Lite Preview 06-17 logo
Mercury 2 vs Gemini 2.5 Flash Lite Preview 06-17New provider

Model pages

Grok 3 Thinking logo
Grok 3 Thinking14 outputs, specs and price
Mercury 2 logo
Mercury 253 outputs, specs and price
Rival
How it worksPrivacyTerms
Explore all of Rival

Explore

  • Compare Models
  • All Models
  • Image Comparison
  • Audio Comparison
  • Image Generation
  • Best AI For...
  • Arena
  • API Pricing
  • Challenges

Discover

  • SubjectiveBench
  • Default Index
  • Research
  • Research downloads
  • Rival Kits
  • Find your AI taste
  • UI Glow-Up
  • VoiceLock
  • Cost Cutter
  • Agent skills
  • Benchmarks vs Vibes
  • Brief
  • Jailbreak
  • Model Updates
  • Provider Status
  • AI Creators

Connect

  • Methodology
  • Advertise
  • Partnerships
  • Privacy Policy
  • Terms
  • RSS Feed