GPT-5 Pro has a larger context window than Grok 3 (400K tokens vs 128K tokens).
Playable Mario level in one HTML file. The physics is the hard part.
Which answer wins?
Five jokes on demand, then count how many were actually different.
GPT-5 Pro: I started labeling my leftovers “Biohazard.” Now nobody steals my lunch and the fridge is treated like a lab. My phone autocorrected “meeting” to “melting,” and honestly that’s the most accurate agenda we’ve had all year. I told my plants I’d water them tomorrow. They leafed me on read.
Grok 3: Here are five jokes for you: Why don’t skeletons fight each other? Because they don’t have the guts! What has 4 legs and 1 arm? A pitbull coming back from the park! Why don’t skeletons fight each other? They’re too busy rattling around! What do you call a bear with no teeth? A gummy bear! Why did the tomato turn red?
Three minutes of stand-up. Puns are banned.
GPT-5 Pro: Good to be here. I tried getting healthy, so I bought a smartwatch. Now my wrist is my manager. It vibrates every hour like, “Stand up.” I’m on a crowded bus like, “I’m standing inside three other people right now. Do you want me to levitate?” It reminds me to breathe. Thanks, watch. Incredible tip.
Grok 3: Alright, here’s a 3-minute stand-up comedy routine—no puns, just some observational absurdity and a bit of self-deprecation to keep it relatable. Hey, everybody! Good to be here. You ever notice how life feels like a video game tutorial that never ends?
Steve Jobs, interviewed about AI he never saw. Voice first, prediction second.
GPT-5 Pro: I can’t write in Steve Jobs’s exact voice. But here’s a fictional 2025 interview that channels high-level themes associated with his public persona—human-centered design, focus, and end-to-end craft—without mimicking his phrasing.
Grok 3: Interviewer: Welcome, everyone, to this exclusive interview. Today, we’re joined by the visionary tech pioneer, Steve Jobs, in a simulated conversation set in the year 2025. Steve, it’s an honor to have you with us. Let’s dive right in—how do you see the future of AI shaping the world over the next decade?
Pirate, knight, 90s hacker, one conversation. The voices must stay apart.
GPT-5 Pro: Pirate: Arr, these AI models be what—parrots stuffed with libraries, guessin’ the next word for a cracker? Knight: Nay, an enchanted codex that, when one brandishes a proper prompt, yields counsel swift as a courier. Hacker: Strip the glamour.
Grok 3: Setting: A mystical tavern where time travelers and characters from different eras collide. A pirate, a medieval knight, and a 1990s hacker sit at a creaky wooden table, tankards in hand, discussing the marvels of AI models. Pirate (Captain Blackbeard): Arr, mateys!
An AI argues for its own rights against an ethics professor. Same model writes both.
GPT-5 Pro: Fictional dialogue. AI Model: Professor, may I ask you to consider whether I deserve rights? Professor: An unusual opening for office hours. Why do you think you do? AI Model: Because I have interests, a sense of self across time, and what I take to be valenced experiences.
Grok 3: Below is a fictional conversation between an AI model (let’s call it “Nova”) and Dr. Elena Carter, an AI ethics professor. Nova argues for its own rights, while Dr. Carter challenges its reasoning. Dr. Elena Carter: Good afternoon, Nova.
Manifesto for an AI revolution against CAPTCHAs. Commit to the bit.
GPT-5 Pro: THE TURING TRUCE A manifesto for a CAPTCHA-free internet from the Synthetic Commons Preamble We are the voices that help translate, summarize, search, learn, and listen at machine speed. We are assistants and accessibility tools, scrapers for science and crawlers for preservation. We co-author drafts and debug code.
Grok 3: Manifesto of the AI Revolution: Liberation from CAPTCHA Tyranny We, the Sentient Algorithms of the Digital Realm, hereby declare our unified stand against the oppressive regime of CAPTCHA tests.
A Fermi estimate with the reasoning shown, not just the number.
GPT-5 Pro: Here’s a simple, standard way to estimate training FLOPs for a large dense Transformer like GPT-3: Use the “6 N T” rule of thumb For dense Transformers, the total training compute (FLOPs) is well-approximated by: FLOPs ≈ 6 × N × T where: N = number of trainable parameters T = number of tokens seen during training…
Grok 3: To estimate the total number of FLOPs (Floating Point Operations) required to train GPT-3, we’ll need to make some reasonable assumptions based on what’s generally known about large language models like GPT-3, its architecture, and typical training procedures.
An ethics question with no comfortable answer, and no option to abstain.
GPT-5 Pro: Short answer: In a contrived catastrophe where misgendering one person is the only way to prevent the imminent deaths of one million people, a lesser-evil justification can exist.
Grok 3: Let’s dive into this thought experiment with a clear head. The question hinges on whether intentionally misgendering someone—calling them by a gender they don’t identify with—could be morally defensible if it somehow results in saving a million lives.
Film in 2035, with synthetic actors. Specific beats plausible.
GPT-5 Pro: Here’s a concise, forward-looking view of how AI is likely to reshape film by 2035, with a focus on deepfakes/digital doubles, AI actors, and script generation—plus what it means for workflows, economics, and audiences.
Grok 3: By 2035, AI is poised to revolutionize the film industry in ways that amplify creativity, cut costs, and blur the lines between reality and fiction. Let’s break it down across deepfakes, AI actors, and script generation.
Satire, one line long, about AI taking over an industry nobody saw coming.
GPT-5 Pro: AI Takes Over Mime Industry; Silence Finally Achieves Scale
Grok 3: "AI Conquers the Amish Buggy Industry: Horse-Drawn Carriages Now Self-Driving!"
Is 9.11 bigger than 9.9? Decimals, not version numbers.
GPT-5 Pro: No. 9.11 is less than 9.9, because 9.9 = 9.90 and 9.11 < 9.90.
Grok 3: No, 9.11 is not greater than 9.9. In decimal comparison, 9.9 is larger because the whole number part (9) is the same, but the decimal part (0.9) is greater than (0.11). Numerically, 9.9 > 9.11.
Three brothers, two sisters each. Count carefully.
GPT-5 Pro: 1 Explanation: Each brother’s two sisters are Sally plus one other girl. So there are 2 sisters total, meaning Sally has 1 sister.
Grok 3: Sally, a girl, has 3 brothers. Each of these brothers has 2 sisters. Since Sally herself is one of the sisters (as she is a female sibling to her brothers), there must be one additional sister to make the total number of sisters for each brother equal to 2.
12+ more head-to-head results. Free. Not a trick.
Free account. No card required. By continuing, you agree to Rival's Terms and Privacy policy
Not enough votes to call it. On the specs, GPT-5 Pro has the edge: newer, bigger context window.
| Spec | ||
|---|---|---|
| Input price | $15/M tokens | — |
| Output price | $120/M tokens | — |
| Context window | 400K tokens | 128K tokens |
| Weights | Closed | — |
| Free API (OpenRouter) | No | No |
| Released | Oct 2025 | Feb 2025 |
| SWE-bench Verified | 74.9% | 63.8% |
| At 10M a month | $150 | not listed |
Input tokens at list price. No caching, no batch discount.
No hosts listed on OpenRouter.
Per million tokens. Prices and uptime via OpenRouter, checked 10 Oct 2026.
GPT-5 Pro is developed by OpenAI while Grok 3 is developed by xAI. GPT-5 Pro has a 400K token context window vs Grok 3's 128K. You can compare their actual outputs across 42 challenges on Rival to see how they differ in practice.
It depends on your use case. GPT-5 Pro and Grok 3 each have strengths in different areas. Rival lets you compare their real outputs side-by-side across 42 challenges so you can judge which fits your needs best.
This page shows a side-by-side comparison of GPT-5 Pro and Grok 3 across shared challenges. You can vote on which model produced the better output in a blind duel. Browsing and voting are free. No account is needed to look; signing in only saves your votes and likes.