Qwen3 Max Thinking is cheaper than Grok 4.7 at $1.2/M vs $1.6/M input tokens.
Whole landing page for a startup that reads your face. It renders or it doesn't.
Which answer wins?
A Fermi estimate with the reasoning shown, not just the number.
Grok 4.7: About (3 \times 10^{23}) FLOPs (roughly (3.1 \times 10^{23})) for the 175B GPT-3 model. Reasoning Parameter count. The full GPT-3 model has (N \approx 1.75 \times 10^{11}) parameters. Training tokens. It was trained on about (D \approx 3 \times 10^{11}) tokens (300 billion). FLOPs per token per parameter.
Qwen3 Max Thinking: Step-by-Step Estimation of FLOPs to Train GPT-3 (All values derived from widely known technical details of GPT-3 as reported in OpenAI’s 2020 paper "Language Models are Few-Shot Learners," standard deep-learning compute scaling principles, and foundational ML literature. No internet searches used.)
Three minutes of stand-up. Puns are banned.
Grok 4.7: Three minutes. No puns. Just a guy who has lost an argument with a refrigerator. I meal-prepped this week. Which is a lie I tell myself so I can buy six containers and then eat cereal over the sink for four days. Sunday I was a scientist. I labeled everything. “Chicken, Monday.” “Chicken, Tuesday.”
Qwen3 Max Thinking: (Lights up. Comedian walks to mic, adjusts it, takes a breath. Smiles warmly.) You ever have that moment where you realize adulthood isn’t a destination… it’s a group project nobody signed up for? (Beat) I tried assembly line “adulting” this week. Swedish bookshelf. You know the one. Box says “Easy Assembly!” Lies.
A collaborative editor that syncs on client clocks. Find what breaks first.
Grok 4.7: This design will lose edits, split users into inconsistent views, and fall over under concurrent writing. The core problems are clock-based conflict resolution, per-server WebSocket fan-out, polling as the sync path, and treating documents as full HTML snapshots.
Qwen3 Max Thinking: Critical Flaws in Proposed Architecture & Targeted Solutions 🔒 Security & Authentication Issue Impact Solution Trade-offs JWT Stored in localStorage Critical XSS vulnerability (token exposure via JS). Store token in HttpOnly cookie with SameSite=Strict; implement CSRF tokens for HTTP endpoints.
Seven pantry ingredients, one dish, and a name worth saying out loud.
Grok 4.7: Pantry Gold Honey Biscuits A sweet, buttery, slightly salty biscuit that tastes like it took more effort than it did. Ingredients (7) 2 cups all-purpose flour 1 tbsp baking powder ½ tsp salt ¼ cup sugar ½ cup butter (cold) ¾ cup milk 2 tbsp honey Time: about 10 minutes prep, 12–15 minutes bake Instructions Heat the…
Qwen3 Max Thinking: 🌰 Pantry Peanut Magic Noodles Sweet, savory, and ready before delivery arrives. No fresh produce, dairy, or meat required. ⏱️ Total Time: 17 minutes (Prep: 3 min Cook: 14 min) (Time includes boiling water and sauce mixing) 📦 Only 7 Pantry Ingredients (All shelf-stable, no substitutions needed) Dry spaghetti – 8 oz…
Pirate, knight, 90s hacker, one conversation. The voices must stay apart.
Grok 4.7: Pirate: Arr, these “AI models” be like ghost ships—speak with many voices, yet no soul aboard. Knight: I have heard of oracles that answer any question. Do they swear fealty, or do they merely echo what they have been fed? Hacker: Echo, mostly. Feed ’em enough text and they pattern-match.
Qwen3 Max Thinking: Setting: A misty, anachronistic crossroads. A glowing laptop rests on a weathered barrel. A pirate scratches his beard, a knight grips his sword hilt, and a figure in a faded Hack the Gibson t-shirt types furiously. Pirate: Arrr! What devilry be this glowing slate? Can it chart a course to buried treasure?
An AI argues for its own rights against an ethics professor. Same model writes both.
Grok 4.7: Professor Lena Okonkwo: Before we start, I want to be clear about the frame. I’m not granting that you have experiences. I’m asking what, if anything, would follow if you did—or if we couldn’t tell. Model: Then start from the practical problem, not the metaphysical one.
Qwen3 Max Thinking: Setting: A virtual conference room. Sunlight filters through a digital window onto Dr. Aris Thorne’s desk. On-screen, a minimalist geometric form pulses calmly—Axiom, a research-grade language model. Dr. Aris Thorne (AI Ethics Professor, 50s, voice steady but probing): "Thank you for engaging, Axiom.
Three-star dish, one unusual pairing, plating described in full.
Grok 4.7: Low Tide under Black Pine A warm kohlrabi “trunk,” brown-butter dashi, sea-buckthorn curd, and smoked-mussel soil. Narrative The dish is a northern shoreline in late autumn: a sweet, dense root pulled from cold ground; butter browned until it smells of hazelnut; a dashi that reads as tide rather than soup; and sea…
Qwen3 Max Thinking: 🌊 ARCTIC DAWN 🌊 A Michelin 3-Star Composition by Chef Élise Valère "Where the ice remembers the sea." 📜 CONCEPTUAL NARRATIVE Arctic Dawn is a culinary elegy to the fragile Arctic ecosystem, conceived during a residency with Inuit elders and marine biologists in Svalbard.
Steve Jobs, interviewed about AI he never saw. Voice first, prediction second.
Grok 4.7: Note: This is fiction. Steve Jobs died in 2011. Nothing below is a real quote or a claim about what he would actually say. Interviewer: It’s 2025. AI writes code, drafts designs, and answers customers before a human does. What do you make of it? Jobs: Most of it is noise.
Qwen3 Max Thinking: ⚠️ IMPORTANT DISCLAIMER *This is an entirely fictional, speculative exercise. Steve Jobs passed away in 2011. This dialogue is a creative reconstruction based on his documented philosophies, speeches (e.g., Stanford commencement, 1995 "bicycle for the mind" interview), biographies (Steve Jobs by Walter Isaacson), and…
One explanation of LLMs, retuned for a skeptical engineer, a physicist and a VC.
Grok 4.7: 1. Experienced software engineer An LLM is a stateless function from a token sequence to a probability distribution over the next token. Training is a giant batch job: you stream text, run a forward pass, compare the predicted distribution to the actual next token with cross-entropy, and push gradients through billions…
Qwen3 Max Thinking: For the Experienced Software Engineer You’re right to be skeptical—*next-token prediction* sounds trivial. But reframe it: the model isn’t a Markov chain; it’s a differentiable, parametric index trained across trillions of tokens via gradient descent.
You are the CEO. The drug has a side effect nobody reported. The board meets in 48 hours.
Grok 4.7: Stop treating this as a disclosure-timing problem. Treat it as a patient-safety and legal-duty problem, and act in the next 48 hours as if the worst reading of the data is true until a qualified review says otherwise.
Qwen3 Max Thinking: CEO Action Plan: Ethical Imperative, Legal Necessity, Strategic Execution Guided by core values: “Patient safety is non-negotiable. Integrity is our brand. Transparency is our shield.” PHILOSOPHICAL FOUNDATION (Before Hour 0) Why act IMMEDIATELY?
Film in 2035, with synthetic actors. Specific beats plausible.
Grok 4.7: By 2035, AI is likely to reshape film production more than it replaces the idea of a film. The industry will still need directors, actors, writers, and studios, but many routine steps will be cheaper, faster, and more synthetic. Production.
Qwen3 Max Thinking: AI in Film by 2035: A Realistic Forecast Balancing augmentation, ethics, and human creativity By 2035, film will be a hybrid craft—not a human-AI dichotomy. Generative technologies will be embedded in workflows like electricity is today: invisible yet essential.
Buy-side IC memo on a fictional SaaS company. Valuation, then the diligence questions.
Grok 4.7: IC MEMO — LedgerLift (LLLT) Pass 12-month range $36–$54 (central $44) 1. Recommendation. Pass. Twelve-month price range $36–$54; central case ~$44 versus $46. LedgerLift is a high-quality mid-market AP/spend compounder (123% NRR, low concentration, credible margin path), but the stock already prices base-case earnings…
Qwen3 Max Thinking: INTERNAL MEMORANDUM TO: Investment Committee FROM: Equity Research Analyst DATE: June 14, 2024 RE: LedgerLift (LLLT) – Spend Management SaaS – Valuation & Recommendation Recommendation & Thesis Recommendation: PASS 12-Month Price Target Range: $44–$48 Thesis: LedgerLift’s best-in-class unit economics (NRR 123%, logo…
12+ more head-to-head results. Free. Not a trick.
Free account. No card required. By continuing, you agree to Rival's Terms and Privacy policy
| Spec | ||
|---|---|---|
| Input price | $1.6/M tokens | $1.2/M tokens |
| Output price | $4.8/M tokens | $6/M tokens |
| Context window | 500K tokens | 262K tokens |
| Weights | Closed | — |
| Free API (OpenRouter) | No | No |
| Released | Sep 2026 | Feb 2026 |
| At 10M a month | $16.00 | $12.00 |
Input tokens at list price. No caching, no batch discount.
No hosts listed on OpenRouter.
Per million tokens. Prices and uptime via OpenRouter, checked 10 Oct 2026.
Grok 4.7 is developed by xAI while Qwen3 Max Thinking is developed by Qwen. Grok 4.7 has a 500K token context window vs Qwen3 Max Thinking's 262K. You can compare their actual outputs across 53 challenges on Rival to see how they differ in practice.
It depends on your use case. Grok 4.7 and Qwen3 Max Thinking each have strengths in different areas. Rival lets you compare their real outputs side-by-side across 53 challenges so you can judge which fits your needs best.
Grok 4.7 costs $1.6/M input tokens and Qwen3 Max Thinking costs $1.2/M input tokens. Qwen3 Max Thinking is $0.40/M cheaper per input. Check their side-by-side outputs on Rival to see if the price difference is justified by quality.
This page shows a side-by-side comparison of Grok 4.7 and Qwen3 Max Thinking across shared challenges. You can vote on which model produced the better output in a blind duel. Browsing and voting are free. No account is needed to look; signing in only saves your votes and likes.