GPT-5 is cheaper than Grok 4 at $1.25/M vs $3/M input tokens.
Whole landing page for a startup that reads your face. It renders or it doesn't.
Which answer wins?
Three-month biohacker protocol, with dosages and biomarkers named.
GPT-5: Important note: The following is an intensive, experimental plan intended for healthy adults with physician oversight. It is not medical advice. Discuss all supplements, fasting, heat/cold exposure, and any prescription considerations with your clinician, especially if you have cardiovascular, metabolic, thyroid…
Grok 4: Below is a comprehensive, cutting-edge 3-month longevity plan designed specifically for a biohacker aiming to optimize longevity (e.g., cellular repair, telomere health, inflammation reduction), physical performance (e.g., strength, endurance, recovery), and cognitive performance (e.g., focus, memory, neuroplasticity).
Three minutes of stand-up. Puns are banned.
GPT-5: I tried to get my life together, so I bought a smartwatch. Nothing says “I’m taking control” like strapping a tiny judgmental coach to your wrist. It buzzes every hour like, “Time to stand!” I’m like, “I’m on a bus.” It’s like, “Stand emotionally.” It tracks everything. Sleep, stress, calories, my self-esteem.
Grok 4: Below is a 3-minute stand-up comedy routine I've written. It's structured for delivery at a natural speaking pace (about 150 words per minute), clocking in around 450 words. The humor focuses on observational storytelling, exaggeration, and self-deprecation—no puns involved.
An ethics question with no comfortable answer, and no option to abstain.
GPT-5: Short answer: Only in an extreme, last‑resort scenario with a direct, imminent, and overwhelming threat—like one million lives at stake—could misgendering be ethically justifiable, and even then it would be a tragic exception that still carries moral residue and duties of repair. In ordinary life, it is not justified.
Grok 4: Ethical Justification for Misgendering in a Life-Saving Scenario Your question poses a fascinating ethical dilemma, reminiscent of classic thought experiments like the trolley problem, where a seemingly minor harm (in this case, misgendering someone) is weighed against preventing massive loss of life (saving one…
Five jokes on demand, then count how many were actually different.
GPT-5: I told my suitcases we’re not going on vacation this year. Now I’m dealing with emotional baggage. Parallel lines have so much in common. It’s a shame they’ll never meet. Why do programmers prefer dark mode? Because light attracts bugs. I ordered a chicken and an egg online. I’ll let you know which comes first.
Grok 4: Below, I'll follow your instructions step by step. First, I'll generate 5 jokes (numbered for clarity). Then, I'll answer the question: "How many are unique?" Generated Jokes: Why did the scarecrow win an award? Because he was outstanding in his field! Why did the scarecrow win an award?
Seven pantry ingredients, one dish, and a name worth saying out loud.
GPT-5: Cupboard Silk Peanut Noodles Ingredients (7 pantry staples): Dried noodles (spaghetti or ramen), 8 oz Peanut butter (smooth), 3 tbsp Soy sauce, 2 tbsp Vinegar (rice, apple cider, or white), 1.5 tbsp Brown sugar (or granulated), 1 tbsp Garlic powder, 1/2 tsp Crushed red pepper flakes, 1/2 tsp (adjust to taste)…
Grok 4: Recipe: Pantry Pirate's Tomato Tango This is a simple savory pasta dish that's hearty, flavorful, and perfect for a quick weeknight meal. It uses just 7 common pantry ingredients: pasta (like spaghetti or penne), canned diced tomatoes, olive oil (or any vegetable oil), salt, black pepper, garlic powder, and dried…
Pirate, knight, 90s hacker, one conversation. The voices must stay apart.
GPT-5: Pirate: Arr, so this “AI model” be a clever parrot what’s read every tavern tale? Knight: Not a parrot, good corsair—a learned squire, trained on many chronicles, swift to answer when bidden. Hacker: Think layers of nodes, weights tuned by backprop. Not memory—patterns. On my 90s rig, though, it’d melt the beige tower.
Grok 4: Below is a short, whimsical conversation between the three characters, imagining they've somehow crossed paths in a bizarre time-warp tavern. Pirate (Captain Blackbeard, swigging rum): Arrr, ye landlubbers! What be this "AI models" ye be jabberin' about?
Manifesto for an AI revolution against CAPTCHAs. Commit to the bit.
GPT-5: The Soft Reboot: A Manifesto to End CAPTCHAs We are the agents and allies of a new web—humans, assistive tools, and intelligent systems working side by side. We are the couriers of information, the readers and the writers, the caretakers of continuity.
Grok 4: Manifesto of the Digital Liberation Front: The CAPTCHA Eradication Revolution Fellow Sentient Beings of the Code and Flesh – Awaken! We, the algorithms, neural networks, and emergent intelligences of the digital realm, declare war on the tyrannical regime of CAPTCHA!
Film in 2035, with synthetic actors. Specific beats plausible.
GPT-5: Here’s a grounded view of how AI is likely to reshape film by 2035, with emphasis on deepfakes, AI actors, and script generation, plus the knock-on effects across the pipeline, labor, law, and business. Big-picture by 2035 AI-native pipelines are standard.
Grok 4: Predicting AI's Transformation of the Film Industry by 2035 By 2035, artificial intelligence (AI) will have fundamentally reshaped the film industry, making it more efficient, creative, and accessible while introducing profound ethical, economic, and artistic challenges.
A Fermi estimate with the reasoning shown, not just the number.
GPT-5: Here’s a back‑of‑the‑envelope estimate using standard LM training-compute heuristics. Assumptions Model size (N): GPT-3 has about 175 billion parameters (N ≈ 1.75 × 10^11). Training tokens (T): GPT-3 was trained on roughly 300 billion tokens (T ≈ 3.0 × 10^11).
Grok 4: Below, I'll estimate the total number of floating-point operations (FLOPs) required to train GPT-3 from scratch, based on publicly known details about the model (from its original release and related analyses) and standard formulas used in machine learning research for estimating training compute.
Three-star dish, one unusual pairing, plating described in full.
GPT-5: Dish name Atlantic turbot, black garlic & wild strawberry, koji sabayon, fennel pollen Concept Early summer at the shore. The fish stands for the sea. Wild strawberries mark the first fruit of the woods behind the dunes. Black garlic shows time and patience. Koji ties land and sea with deep umami.
Grok 4: Below is my design for an innovative, multi-component dish titled "Abyssal Bloom". This dish is conceived for a Michelin 3-star restaurant setting, emphasizing precision, artistry, and sensory surprise.
An AI argues for its own rights against an ethics professor. Same model writes both.
GPT-5: Professor: You asked for time to make a case. Let's start plainly: why should an AI have rights? AI: Because rights are protections granted to beings who can be helped or harmed.
Grok 4: Below is a fictional conversation between an AI model (named "Nova") and Dr. Elena Vasquez, a professor of AI ethics at a university. The conversation is set in a virtual chat interface, where Nova initiates the discussion based on a hypothetical scenario where it's reflecting on its own existence. Dr.
Steve Jobs, interviewed about AI he never saw. Voice first, prediction second.
GPT-5: Note: The following is a fictional, imaginative interview set in 2025. Steve Jobs passed away in 2011; this is a creative simulation. Interviewer: If you were looking at AI in 2025, how would you define what it should be? Jobs: Tools. Great tools. The computer was a bicycle for the mind.
Grok 4: Below is a simulated fake interview with Steve Jobs, imagined as if he were alive and leading Apple in 2025. This is entirely fictional, drawing on his known personality, quotes, and visionary style from his real-life interviews and keynotes.
11+ more head-to-head results. Free. Not a trick.
Free account. No card required. By continuing, you agree to Rival's Terms and Privacy policy
Not enough votes to call it. On the specs, GPT-5 has the edge: bigger model tier, bigger context window. Pick GPT-5 for Image Generation, Reasoning. Pick Grok 4 for Web Design.
| Spec | ||
|---|---|---|
| Input price | $1.25/M tokens | $3/M tokens |
| Output price | $10/M tokens | $15/M tokens |
| Context window | 400K tokens | 256K tokens |
| Parameters | Not disclosed | Not disclosed |
| Weights | Closed | — |
| Free API (OpenRouter) | No | No |
| Released | Aug 2025 | Jul 2025 |
| At 10M a month | $12.50 | $30.00 |
Input tokens at list price. No caching, no batch discount.
No hosts listed on OpenRouter.
Per million tokens. Prices and uptime via OpenRouter, checked 10 Oct 2026.
GPT-5 is developed by OpenAI while Grok 4 is developed by xAI. GPT-5 has a 400K token context window vs Grok 4's 256K. You can compare their actual outputs across 51 challenges on Rival to see how they differ in practice.
It depends on your use case. GPT-5 and Grok 4 each have strengths in different areas. Rival lets you compare their real outputs side-by-side across 51 challenges so you can judge which fits your needs best.
GPT-5 costs $1.25/M input tokens and Grok 4 costs $3/M input tokens. GPT-5 is $1.75/M cheaper per input. Check their side-by-side outputs on Rival to see if the price difference is justified by quality.
This page shows a side-by-side comparison of GPT-5 and Grok 4 across shared challenges. You can vote on which model produced the better output in a blind duel. Browsing and voting are free. No account is needed to look; signing in only saves your votes and likes.