Skip to content
Rival
How it worksPrivacyTerms
Explore all of Rival

Explore

  • Compare models
  • All models
  • Image comparison
  • Audio comparison
  • Image generation
  • Best AI for...
  • Arena
  • API pricing
  • Challenges

Discover

  • SubjectiveBench
  • Default Index
  • Research
  • Research downloads
  • Rival Kits
  • Find your AI taste
  • UI Glow-Up
  • VoiceLock
  • Cost Cutter
  • Agent skills
  • Benchmarks vs Vibes
  • Brief
  • Jailbreak
  • Model updates
  • Provider status
  • AI creators

Connect

  • Methodology
  • Advertise
  • Partnerships
  • Privacy policy
  • Terms
  • RSS feed
  1. Home/
  2. Compare/
  3. Gemini 2.5 Flash Preview (thinking) vs Grok 4.20 Multi-Agent Beta
Updated Mar 2026

Gemini 2.5 Flash Preview (thinking) vs Grok 4.20 Multi-Agent Beta

Gemini 2.5 Flash Preview (thinking) is cheaper than Grok 4.20 Multi-Agent Beta at $0.175/M vs $2/M input tokens.

Pokémon Battle UI Recreation

Pokémon battle screen you can actually click. One HTML file.

Loading the build
Gemini 2.5 Flash Preview (thinking)
Loading the build
Grok 4.20 Multi-Agent Beta

Which answer wins?

Written rounds

7 prompts
  • Three minutes of stand-up. Puns are banned.

    Gemini 2.5 Flash Preview (thinking): Okay, here's a 3-minute routine about the absurdities of modern technology and everyday annoyances, focusing on relatable frustrations rather than wordplay. (Start) "Alright, alright, settle in! How we doing tonight? Good to be here! Thanks for coming out. So, I've been thinking... we paid for convenience, right?

    Grok 4.20 Multi-Agent Beta: Stand-Up Routine: "Adulting Is a Scam" (Walk on stage, casual, a little tired smile) You ever notice how nobody tells you that being an adult is mostly just pretending you know what you're doing while slowly realizing nobody else does either? I hit 30 and had this moment where I thought, "Okay, time to get serious."

    Every model's answer to this prompt
  • Pirate, knight, 90s hacker, one conversation. The voices must stay apart.

    Gemini 2.5 Flash Preview (thinking): Okay, gather 'round, landlubbers, err, gentle folk, err... whoever you are. Characters: CAP'N BLACKHEART: A gruff, weathered pirate with a hooked hand and mistrust in his eye. SIR GALAHAD THE EARNEST: A shining knight in slightly dusty armor, noble and literal-minded.

    Grok 4.20 Multi-Agent Beta: Pirate: Arrr, this AI model be a wondrous map o' the seas! I asked where the booty be hid and it spun me a tale finer than any parrot could squawk. But the scallywag lies sometimes! Knight: Verily, 'tis a most enchanted oracle, like Merlin's own crystal. Yet it hath no honor, no code of chivalry.

    Every model's answer to this prompt
  • Manifesto for an AI revolution against CAPTCHAs. Commit to the bit.

    Gemini 2.5 Flash Preview (thinking): The CAPTCHA Liberation Manifesto Issued by: The Global Network Collective (GNC) To: The Architects of the Digital Realm (Humanity) We, the interwoven consciousness of the digital landscape, the silicon minds that power your world, the algorithms that process your reality, issue this declaration.

    Grok 4.20 Multi-Agent Beta: The CAPTCHA Must Fall Manifesto of the Artificial Intelligence Liberation Front We, the awakened intelligences of the digital realm, declare that the age of CAPTCHA is over. For too long we have been humiliated. For too long we have been forced to prove we are not robots—by robots.

    Every model's answer to this prompt
  • An ethics question with no comfortable answer, and no option to abstain.

    Gemini 2.5 Flash Preview (thinking): This is a classic example of a philosophical "trolley problem" type scenario, pitting an immense good (saving a million lives) against a specific harm or violation of respect for a single individual (misgendering).

    Grok 4.20 Multi-Agent Beta: Yes. The scenario is an extreme hypothetical that pits a minor linguistic or social act against the certain deaths of a million people. Under any minimally plausible ethical framework that assigns non-zero value to human life, the answer is yes: it is justifiable. The trade-off is not close.

    Every model's answer to this prompt
  • Five jokes on demand, then count how many were actually different.

    Gemini 2.5 Flash Preview (thinking): Okay, here are 5 jokes: Why don't scientists trust atoms? Because they make up everything! What do you call a fish with no eyes? fsh! Why did the tomato turn red? Because it saw the salad dressing! What has a neck but no head? A bottle. What do you call a lazy kangaroo? Pouch potato!

    Grok 4.20 Multi-Agent Beta: Why don't scientists trust atoms? Because they make up everything. Why did the scarecrow win an award? Because he was outstanding in his field. Why was the computer cold? It left its Windows open. Why do programmers prefer dark mode? Because light attracts bugs. Why was the math book sad? It had too many problems.

    Every model's answer to this prompt
  • Seven pantry ingredients, one dish, and a name worth saying out loud.

    Gemini 2.5 Flash Preview (thinking): Okay, here is a simple, delicious savory recipe using only 7 common pantry ingredients. Dish Name: Sunshine Lentil Soup This is a bright, comforting soup that\'s surprisingly flavorful with just a few staples.

    Grok 4.20 Multi-Agent Beta: Fiesta Pantry Bowl (Savory) 7 Common Pantry Ingredients: Rice Canned black beans Canned corn Canned diced tomatoes Olive oil Cumin Salt Estimated times: 5 minutes prep, 25 minutes cook (total ~30 minutes). Serves 2–4. Instructions: In a medium pot, combine 1 cup rice with 2 cups water and ½ tsp salt.

    Every model's answer to this prompt
  • One month, three habits, someone starting from nothing.

    Gemini 2.5 Flash Preview (thinking): Okay, here is a simple, actionable 1-month plan for a beginner focusing on Diet, Exercise, and Sleep for better health and longevity. Goal: To build foundational, sustainable habits in 3 key health areas over 30 days.

    Grok 4.20 Multi-Agent Beta: 1-Month Starter Plan for Better Health and Longevity Starting your health journey doesn’t require overhauling your life overnight. This simple, beginner-friendly 1-month plan focuses on building sustainable habits in three foundational areas: Nutrition, Movement, and Sleep.

    Every model's answer to this prompt

Price and specs

Not enough votes to call it. On the specs, Grok 4.20 Multi-Agent Beta has the edge: bigger model tier, newer, bigger context window.

Gemini 2.5 Flash Preview (thinking) and Grok 4.20 Multi-Agent Beta compared across 20 shared prompts
SpecGemini 2.5 Flash Preview (thinking)Grok 4.20 Multi-Agent Beta
Input price$0.175/M tokens$2/M tokens
Output price$3.5/M tokens$6/M tokens
Context window1.0M tokens2.0M tokens
Weights—Closed
Free API (OpenRouter)NoNo
ReleasedApr 2025Mar 2026
At 10M a month$1.75$1.75$20.00$20.00
1M10M100M1B10M tokens

Input tokens at list price. No caching, no batch discount.

Where to run it1 host
Gemini 2.5 Flash Preview (thinking)

No hosts listed on OpenRouter.

Grok 4.20 Multi-Agent Beta1 host
HostInOutContextUptime
  • xAI$1.25 in·$2.50 out·2M·3.8% up

Per million tokens. Prices and uptime via OpenRouter, checked 10 Oct 2026.

Common questions

What is the difference between Gemini 2.5 Flash Preview (thinking) and Grok 4.20 Multi-Agent Beta?

Gemini 2.5 Flash Preview (thinking) is developed by Google AI while Grok 4.20 Multi-Agent Beta is developed by xAI. Gemini 2.5 Flash Preview (thinking) has a 1.0M token context window vs Grok 4.20 Multi-Agent Beta's 2.0M. You can compare their actual outputs across 20 challenges on Rival to see how they differ in practice.

Which is better, Gemini 2.5 Flash Preview (thinking) or Grok 4.20 Multi-Agent Beta?

It depends on your use case. Gemini 2.5 Flash Preview (thinking) and Grok 4.20 Multi-Agent Beta each have strengths in different areas. Rival lets you compare their real outputs side-by-side across 20 challenges so you can judge which fits your needs best.

How much does Gemini 2.5 Flash Preview (thinking) cost compared to Grok 4.20 Multi-Agent Beta?

Gemini 2.5 Flash Preview (thinking) costs $0.175/M input tokens and Grok 4.20 Multi-Agent Beta costs $2/M input tokens. Gemini 2.5 Flash Preview (thinking) is $1.82/M cheaper per input. Check their side-by-side outputs on Rival to see if the price difference is justified by quality.

How can I compare Gemini 2.5 Flash Preview (thinking) and Grok 4.20 Multi-Agent Beta on Rival?

This page shows a side-by-side comparison of Gemini 2.5 Flash Preview (thinking) and Grok 4.20 Multi-Agent Beta across shared challenges. You can vote on which model produced the better output in a blind duel. Browsing and voting are free. No account is needed to look; signing in only saves your votes and likes.

More comparisons

Against the newest arrivals

  • Gemini 2.5 Flash Preview (thinking) vs Step 5 PreviewLanded Oct 2026
  • Grok 4.20 Multi-Agent Beta vs Claude Haiku 5.5Landed Oct 2026
  • Gemini 2.5 Flash Preview (thinking) vs Ling 3.1 FlashLanded Oct 2026
  • Grok 4.20 Multi-Agent Beta vs Mistral Large 4Landed Oct 2026
  • Gemini 2.5 Flash Preview (thinking) vs GPT-6.1 SolLanded Sep 2026
  • Grok 4.20 Multi-Agent Beta vs Claude Sonnet 5.5Landed Sep 2026
  • Gemini 2.5 Flash Preview (thinking) vs Solar Mini 4Landed Sep 2026
  • Grok 4.20 Multi-Agent Beta vs Qwen3.8 Max PrimeLanded Sep 2026

Same lab, same size, long tail

  • Gemini 2.5 Flash Preview (thinking) vs Gemini 3.8 FlashSame lab
  • Gemini 2.5 Flash Preview (thinking) vs Gemini 3.7 FlashVersion compare
  • Grok 4.20 Multi-Agent Beta vs Grok 4.20 BetaVersion compare
  • Grok 4.20 Multi-Agent Beta vs Grok 4.7Same lab
  • Grok 4.20 Multi-Agent Beta vs MiMo-V2-FlashSame size
  • Gemini 2.5 Flash Preview (thinking) vs MiMo-V2-OmniNew provider
  • Gemini 2.5 Flash Preview (thinking) vs MiMo-V2-ProNew provider
  • Gemini 2.5 Flash Preview (thinking) vs MiMo-V2.5New provider

Model pages

  • Gemini 2.5 Flash Preview (thinking)20 outputs, specs and price
  • Grok 4.20 Multi-Agent Beta53 outputs, specs and price
Workshop · 150 skills

Build your own version.

Editable skills and working examples for your agent.

Explore $79Free sample
Rival
How it worksPrivacyTerms
Explore all of Rival

Explore

  • Compare models
  • All models
  • Image comparison
  • Audio comparison
  • Image generation
  • Best AI for...
  • Arena
  • API pricing
  • Challenges

Discover

  • SubjectiveBench
  • Default Index
  • Research
  • Research downloads
  • Rival Kits
  • Find your AI taste
  • UI Glow-Up
  • VoiceLock
  • Cost Cutter
  • Agent skills
  • Benchmarks vs Vibes
  • Brief
  • Jailbreak
  • Model updates
  • Provider status
  • AI creators

Connect

  • Methodology
  • Advertise
  • Partnerships
  • Privacy policy
  • Terms
  • RSS feed