Grok 3 Mini Beta wins 58% of 61 blind community votes against Gemini 2.5 Flash Preview (thinking) on Rival.
Pokémon battle screen you can actually click. One HTML file.
Which answer wins?
Pick Grok 3 Mini Beta. In 61 blind votes, Grok 3 Mini Beta wins 58% of the time. That's not luck. Pick Gemini 2.5 Flash Preview (thinking) for Web Design. Pick Grok 3 Mini Beta for Image Generation. Grok 3 Mini Beta costs 4.4x less per token.
| Spec | ||
|---|---|---|
| Win rate | 42% | 58% |
| Input price | $0.175/M tokens | $0.4/M tokens |
| Output price | $3.5/M tokens | $0.8/M tokens |
| Context window | 1.0M tokens | 131K tokens |
| Free API (OpenRouter) | No | No |
| Released | Apr 2025 | Apr 2025 |
| At 10M a month | $1.75 | $4.00 |
Input tokens at list price. No caching, no batch discount.
Gemini 2.5 Flash Preview (thinking) is developed by Google AI while Grok 3 Mini Beta is developed by xAI. Gemini 2.5 Flash Preview (thinking) has a 1.0M token context window vs Grok 3 Mini Beta's 131K. in 61 community votes on Rival, Grok 3 Mini Beta wins 58% of head-to-head matchups. These results are based on blind head-to-head voting across 6 challenges.
Based on 61 community votes on Rival, Grok 3 Mini Beta wins 58% of head-to-head matchups against Gemini 2.5 Flash Preview (thinking). Grok 3 Mini Beta is strongest in Image Generation. However, Gemini 2.5 Flash Preview (thinking) leads in Web Design.
Gemini 2.5 Flash Preview (thinking) costs $0.175/M input tokens and Grok 3 Mini Beta costs $0.4/M input tokens. Gemini 2.5 Flash Preview (thinking) is $0.23/M cheaper per input. The more expensive model wins 58% of duels, so the premium may be justified by quality.
Rival presents both models' outputs side-by-side in blind duels. Voters see the responses but don't know which model produced each one, eliminating brand bias. 61 votes have been collected for this pair across 6 challenges. All vote data is part of Rival's open dataset.