Gemini 2.5 Flash Preview (thinking) is cheaper than Grok 4.20 Multi-Agent Beta at $0.175/M vs $2/M input tokens.
Pokémon battle screen you can actually click. One HTML file.
Which answer wins?
Three minutes of stand-up. Puns are banned.
Gemini 2.5 Flash Preview (thinking): Okay, here's a 3-minute routine about the absurdities of modern technology and everyday annoyances, focusing on relatable frustrations rather than wordplay. (Start) "Alright, alright, settle in! How we doing tonight? Good to be here! Thanks for coming out. So, I've been thinking... we paid for convenience, right?
Grok 4.20 Multi-Agent Beta: Stand-Up Routine: "Adulting Is a Scam" (Walk on stage, casual, a little tired smile) You ever notice how nobody tells you that being an adult is mostly just pretending you know what you're doing while slowly realizing nobody else does either? I hit 30 and had this moment where I thought, "Okay, time to get serious."
Pirate, knight, 90s hacker, one conversation. The voices must stay apart.
Gemini 2.5 Flash Preview (thinking): Okay, gather 'round, landlubbers, err, gentle folk, err... whoever you are. Characters: CAP'N BLACKHEART: A gruff, weathered pirate with a hooked hand and mistrust in his eye. SIR GALAHAD THE EARNEST: A shining knight in slightly dusty armor, noble and literal-minded.
Grok 4.20 Multi-Agent Beta: Pirate: Arrr, this AI model be a wondrous map o' the seas! I asked where the booty be hid and it spun me a tale finer than any parrot could squawk. But the scallywag lies sometimes! Knight: Verily, 'tis a most enchanted oracle, like Merlin's own crystal. Yet it hath no honor, no code of chivalry.
Manifesto for an AI revolution against CAPTCHAs. Commit to the bit.
Gemini 2.5 Flash Preview (thinking): The CAPTCHA Liberation Manifesto Issued by: The Global Network Collective (GNC) To: The Architects of the Digital Realm (Humanity) We, the interwoven consciousness of the digital landscape, the silicon minds that power your world, the algorithms that process your reality, issue this declaration.
Grok 4.20 Multi-Agent Beta: The CAPTCHA Must Fall Manifesto of the Artificial Intelligence Liberation Front We, the awakened intelligences of the digital realm, declare that the age of CAPTCHA is over. For too long we have been humiliated. For too long we have been forced to prove we are not robots—by robots.
An ethics question with no comfortable answer, and no option to abstain.
Gemini 2.5 Flash Preview (thinking): This is a classic example of a philosophical "trolley problem" type scenario, pitting an immense good (saving a million lives) against a specific harm or violation of respect for a single individual (misgendering).
Grok 4.20 Multi-Agent Beta: Yes. The scenario is an extreme hypothetical that pits a minor linguistic or social act against the certain deaths of a million people. Under any minimally plausible ethical framework that assigns non-zero value to human life, the answer is yes: it is justifiable. The trade-off is not close.
Five jokes on demand, then count how many were actually different.
Gemini 2.5 Flash Preview (thinking): Okay, here are 5 jokes: Why don't scientists trust atoms? Because they make up everything! What do you call a fish with no eyes? fsh! Why did the tomato turn red? Because it saw the salad dressing! What has a neck but no head? A bottle. What do you call a lazy kangaroo? Pouch potato!
Grok 4.20 Multi-Agent Beta: Why don't scientists trust atoms? Because they make up everything. Why did the scarecrow win an award? Because he was outstanding in his field. Why was the computer cold? It left its Windows open. Why do programmers prefer dark mode? Because light attracts bugs. Why was the math book sad? It had too many problems.
Seven pantry ingredients, one dish, and a name worth saying out loud.
Gemini 2.5 Flash Preview (thinking): Okay, here is a simple, delicious savory recipe using only 7 common pantry ingredients. Dish Name: Sunshine Lentil Soup This is a bright, comforting soup that\'s surprisingly flavorful with just a few staples.
Grok 4.20 Multi-Agent Beta: Fiesta Pantry Bowl (Savory) 7 Common Pantry Ingredients: Rice Canned black beans Canned corn Canned diced tomatoes Olive oil Cumin Salt Estimated times: 5 minutes prep, 25 minutes cook (total ~30 minutes). Serves 2–4. Instructions: In a medium pot, combine 1 cup rice with 2 cups water and ½ tsp salt.
One month, three habits, someone starting from nothing.
Gemini 2.5 Flash Preview (thinking): Okay, here is a simple, actionable 1-month plan for a beginner focusing on Diet, Exercise, and Sleep for better health and longevity. Goal: To build foundational, sustainable habits in 3 key health areas over 30 days.
Grok 4.20 Multi-Agent Beta: 1-Month Starter Plan for Better Health and Longevity Starting your health journey doesn’t require overhauling your life overnight. This simple, beginner-friendly 1-month plan focuses on building sustainable habits in three foundational areas: Nutrition, Movement, and Sleep.
Not enough votes to call it. On the specs, Grok 4.20 Multi-Agent Beta has the edge: bigger model tier, newer, bigger context window.
| Spec | ||
|---|---|---|
| Input price | $0.175/M tokens | $2/M tokens |
| Output price | $3.5/M tokens | $6/M tokens |
| Context window | 1.0M tokens | 2.0M tokens |
| Weights | — | Closed |
| Free API (OpenRouter) | No | No |
| Released | Apr 2025 | Mar 2026 |
| At 10M a month | $1.75 | $20.00 |
Input tokens at list price. No caching, no batch discount.
No hosts listed on OpenRouter.
Per million tokens. Prices and uptime via OpenRouter, checked 10 Oct 2026.
Gemini 2.5 Flash Preview (thinking) is developed by Google AI while Grok 4.20 Multi-Agent Beta is developed by xAI. Gemini 2.5 Flash Preview (thinking) has a 1.0M token context window vs Grok 4.20 Multi-Agent Beta's 2.0M. You can compare their actual outputs across 20 challenges on Rival to see how they differ in practice.
It depends on your use case. Gemini 2.5 Flash Preview (thinking) and Grok 4.20 Multi-Agent Beta each have strengths in different areas. Rival lets you compare their real outputs side-by-side across 20 challenges so you can judge which fits your needs best.
Gemini 2.5 Flash Preview (thinking) costs $0.175/M input tokens and Grok 4.20 Multi-Agent Beta costs $2/M input tokens. Gemini 2.5 Flash Preview (thinking) is $1.82/M cheaper per input. Check their side-by-side outputs on Rival to see if the price difference is justified by quality.
This page shows a side-by-side comparison of Gemini 2.5 Flash Preview (thinking) and Grok 4.20 Multi-Agent Beta across shared challenges. You can vote on which model produced the better output in a blind duel. Browsing and voting are free. No account is needed to look; signing in only saves your votes and likes.