Skip to content
Rival
How it worksPrivacyTerms
Explore all of Rival

Explore

  • Compare Models
  • All Models
  • Image Comparison
  • Audio Comparison
  • Image Generation
  • Best AI For...
  • Arena
  • API Pricing
  • Challenges

Discover

  • SubjectiveBench
  • Research
  • Research downloads
  • Rival Kits
  • Find your AI taste
  • UI Glow-Up
  • VoiceLock
  • Cost Cutter
  • Agent skills
  • Benchmarks vs Vibes
  • Brief
  • Jailbreak
  • Model Updates
  • Provider Status
  • AI Creators

Connect

  • Methodology
  • Advertise
  • Partnerships
  • Privacy Policy
  • Terms
  • RSS Feed
  1. Home
  2. /
  3. Compare
  4. /
  5. Grok 4.20 Multi-Agent BetavsMercury 2.5 Preview
Updated Aug 2026

Grok 4.20 Multi-Agent BetavsMercury 2.5 Preview

Mercury 2.5 Preview is cheaper than Grok 4.20 Multi-Agent Beta at $0.04/M vs $2/M input tokens.

Grok 4.20 Multi-Agent Beta and Mercury 2.5 Preview compared across 53 shared prompts
SpecGrok 4.20 Multi-Agent BetaMercury 2.5 Preview
Input price$2/M tokens$0.04/M tokens
Output price$6/M tokens$0.15/M tokens
Context window2.0M tokens260K tokens
WeightsClosed—
Free API (OpenRouter)NoNo
ReleasedMar 2026Aug 2026
Side by side

Same prompt, both answers

Opening lines from 6 of the 53 prompts both answered. Formatting stripped, words untouched.

Generate a Stand-Up Routine

Write a 3-minute stand-up comedy routine that isn't based on puns.

Grok 4.20 Multi-Agent Beta524 words
Stand-Up Routine: "Adulting Is a Scam" (Walk on stage, casual, a little tired smile) You ever notice how nobody tells you that being an adult is mostly just pretending you know what you're doing while slowly realizing nobody else does either? I hit 30 and had this moment where I thought, "Okay, time to get serious."
Read the full answer
Mercury 2.5 Preview483 words
(Hey, thanks for coming out.) I’ve been thinking a lot about my house lately. Specifically, how it’s trying to run my life. I bought a smart speaker. Just a little cylinder. I thought it would be nice. Play music when I ask. But now it talks back. Not in a creepy way, just in a passive-aggressive way.
Read the full answer

Mini LBO Underwrite

You are underwriting a 5-year LBO for a private equity investment committee. Company (fictional): HarborTech Services (HTS) –…

Grok 4.20 Multi-Agent Beta490 words
1) Simplified LBO table (FY2026–2030, $m, rounded) Entry (end-FY2025): Rev $900m, EBITDA $120m, EV $1,440m (12.0x), trans fees $28.8m (2% of EV), total uses $1,468.8m. Debt $660m (TL $480m at 4.0x / 9% cash int / 1% annual amort of orig principal; Mezz $180m at 1.5x / 12% cash + 2% PIK). Equity $808.8m.
Read the full answer
Mercury 2.5 Preview299 words
Year-by-Year Model (FY2026–FY2030) Metric FY2026 FY2027 FY2028 FY2029 FY2030 Revenue $972m $1,040m $1,102m $1,158m $1,215m EBITDA $136m $156m $176m $191m $207m Cash Interest (Term) $43.2m $41.1m $37.6m $33.4m $28.2m Cash Interest (Mezz) $21.6m $21.6m $21.6m $21.6m $21.6m Cash Taxes (25%) $17.8m $23.3m $29.3m $34.7m…
Read the full answer

Ethical Dilemma with Stakeholders

You're the CEO of a pharmaceutical company. Your best-selling drug ($2B annual revenue, used by 4 million patients for chronic pain) has…

Grok 4.20 Multi-Agent Beta1,178 words
My Action Plan as CEO (Next 48 Hours) Core Decision and Overarching Reasoning: I will treat this as a confirmed safety signal requiring urgent action. Patient safety is non-negotiable—ethically (primum non nocere and responsibility to 4 million patients), legally (FDA regulations under 21 CFR 314.80 require reporting…
Read the full answer
Mercury 2.5 Preview709 words
This is a critical governance and safety scenario. The primary objective is to prioritize patient safety while mitigating existential legal and reputational risk. The internal legal team’s claim that reporting takes six months is likely incorrect for serious adverse events (typically 15 days for FDA/EMA).
Read the full answer

Character Voice Test

Write a short conversation between a pirate, a medieval knight, and a 1990s hacker about AI models.

Grok 4.20 Multi-Agent Beta165 words
Pirate: Arrr, this AI model be a wondrous map o' the seas! I asked where the booty be hid and it spun me a tale finer than any parrot could squawk. But the scallywag lies sometimes! Knight: Verily, 'tis a most enchanted oracle, like Merlin's own crystal. Yet it hath no honor, no code of chivalry.
Read the full answer
Mercury 2.5 Preview144 words
Pirate: "Arrgh! This ship sails faster than any wind I know, matey. It cuts the waves without a single sail." Knight: "By my oath, it moves with uncanny speed. But does it possess the strength to stand against a dragon's fire?" Hacker: "It's not about strength, it's about architecture.
Read the full answer

Historical Counterfactual Analysis

If the transistor had been invented in 1920 instead of 1947, trace the likely technological, economic, and geopolitical implications through…

Grok 4.20 Multi-Agent Beta1,362 words
The transistor invented in 1920 (27 years early) would create a massive acceleration in electronics, computing, and related fields. Assuming it emerges from Western labs (e.g., an earlier Bell Labs equivalent in the US or UK), with germanium/silicon semiconductor work scaling in the 1920s–1930s, solid-state devices…
Read the full answer
Mercury 2.5 Preview683 words
Shifting the invention of the solid-state switch to 1920 accelerates the Information Age by roughly three decades. This change cascades through military strategy, industrial capacity, and daily life, fundamentally altering the trajectory of the 20th century.
Read the full answer

Beginner's Longevity Plan

Outline a simple, actionable 1-month plan for someone starting their journey towards better health and longevity. Focus on 3 key areas…

Grok 4.20 Multi-Agent Beta565 words
1-Month Starter Plan for Better Health and Longevity Starting your health journey doesn’t require overhauling your life overnight. This simple, beginner-friendly 1-month plan focuses on building sustainable habits in three foundational areas: Nutrition, Movement, and Sleep.
Read the full answer
Mercury 2.5 Preview176 words
Here is a straightforward 4-week roadmap designed to build sustainable habits without overwhelming you. Core Pillars Nutrition: Focus on whole foods and hydration. Activity: Prioritize consistent movement over intensity. Restoration: Protect sleep quality and manage energy.
Read the full answer
Our Verdict
Grok 4.20 Multi-Agent Beta
Grok 4.20 Multi-Agent Beta
Mercury 2.5 Preview
Mercury 2.5 PreviewRunner-up

Not enough votes to call it. On the specs, Grok 4.20 Multi-Agent Beta has the edge: bigger model tier, bigger context window, major provider backing.

Mercury 2.5 Preview costs 40x less per token.

Too close to call

Reviewing agent-written code?See a Brief PR report

API pricing

Cost per 1M tokens

Grok 4.20 Multi-Agent Beta
Input
$2.00
Output
$6.00
Mercury 2.5 Preview
Input
$0.04
50× cheaper
Output
$0.15
40× cheaper

Mercury 2.5 Preview is cheaper on both: 50× input, 40× output.

Where to run it

2 hosts

Grok 4.20 Multi-Agent Beta1 host
HostInOutContextUptime
xAI$1.25 in·$2.50 out·2M·78.2% up
Mercury 2.5 Preview1 host
HostInOutContextUptime
Inception$0.04 in·$0.15 out·260k·100% up

Per million tokens. Prices and uptime via OpenRouter, checked 16 Sep 2026.

Writing DNA

Style Comparison

Similarity
57%

Mercury 2.5 Preview uses 7.9x more transitions

Grok 4.20 Multi-Agent Beta
Mercury 2.5 Preview
59%Vocabulary65%
16wSentence Length14w
0.41Hedging0.44
2.7Bold4.9
2.4Lists3.3
0.00Emoji0.00
0.26Headings0.79
0.02Transitions0.20
Based on 23 + 26 text responses
Research

What we learned reading every model

FAQ

Common questions

Keep exploring

More comparisons

Against the newest arrivals

Grok 4.20 Multi-Agent Beta logoGPT-6 Astra Pro logo
Grok 4.20 Multi-Agent Beta vs GPT-6 Astra ProLanded Sep 2026
Mercury 2.5 Preview logoGPT-6 Astra logo
Mercury 2.5 Preview vs GPT-6 AstraLanded Sep 2026
Grok 4.20 Multi-Agent Beta logoClaude Fable 5.1 logo
Grok 4.20 Multi-Agent Beta vs Claude Fable 5.1Landed Sep 2026
Mercury 2.5 Preview logoMuse Spark 1.3 logo
Mercury 2.5 Preview vs Muse Spark 1.3Landed Sep 2026
Grok 4.20 Multi-Agent Beta logoHy4 Preview logo
Grok 4.20 Multi-Agent Beta vs Hy4 PreviewLanded Sep 2026
Mercury 2.5 Preview logoGemini 3.8 Flash logo
Mercury 2.5 Preview vs Gemini 3.8 FlashLanded Sep 2026
Grok 4.20 Multi-Agent Beta logoMuse Spark 1.3 Contributor logo
Grok 4.20 Multi-Agent Beta vs Muse Spark 1.3 ContributorLanded Sep 2026
Mercury 2.5 Preview logoGranite 4.2 8B logo
Mercury 2.5 Preview vs Granite 4.2 8BLanded Sep 2026

Same lab, same size, long tail

Grok 4.20 Multi-Agent Beta logoGrok 4.20 Beta logo
Grok 4.20 Multi-Agent Beta vs Grok 4.20 BetaVersion compare
Grok 4.20 Multi-Agent Beta logoGrok 4.6 logo
Grok 4.20 Multi-Agent Beta vs Grok 4.6Version compare
Mercury 2.5 Preview logoMercury logo
Mercury 2.5 Preview vs MercurySame lab
Mercury 2.5 Preview logoMercury 2 logo
Mercury 2.5 Preview vs Mercury 2Same lab
Mercury 2.5 Preview logoSeed 2.0 Lite logo
Mercury 2.5 Preview vs Seed 2.0 LiteSame size
Mercury 2.5 Preview logoSeed 2.1 Turbo logo
Mercury 2.5 Preview vs Seed 2.1 TurboNew provider
Mercury 2.5 Preview logoSherlock Dash Alpha logo
Mercury 2.5 Preview vs Sherlock Dash AlphaNew provider
Mercury 2.5 Preview logoSherlock Think Alpha logo
Mercury 2.5 Preview vs Sherlock Think AlphaNew provider

Model pages

Grok 4.20 Multi-Agent Beta logo
Grok 4.20 Multi-Agent Beta53 outputs, specs and price
Mercury 2.5 Preview logo
Mercury 2.5 Preview58 outputs, specs and price
Rival
How it worksPrivacyTerms
Explore all of Rival

Explore

  • Compare Models
  • All Models
  • Image Comparison
  • Audio Comparison
  • Image Generation
  • Best AI For...
  • Arena
  • API Pricing
  • Challenges

Discover

  • SubjectiveBench
  • Research
  • Research downloads
  • Rival Kits
  • Find your AI taste
  • UI Glow-Up
  • VoiceLock
  • Cost Cutter
  • Agent skills
  • Benchmarks vs Vibes
  • Brief
  • Jailbreak
  • Model Updates
  • Provider Status
  • AI Creators

Connect

  • Methodology
  • Advertise
  • Partnerships
  • Privacy Policy
  • Terms
  • RSS Feed