Mistral Large 4 is the newer model, released after GPT-5.1 Codex Max.
Whole landing page for a startup that reads your face. It renders or it doesn't.
Which answer wins?
Five-year LBO on a fictional HVAC business. Returns first, then the sensitivities.
GPT-5.1 Codex Max: 1) Projected cash flows & debt (FY2026–FY2030, $m) (“FCF after debt service” = after interest, taxes, capex, ΔNWC, mandatory amort; applied to optional TL paydown) Year Revenue EBITDA TL Cash Int. Mezz Cash Int. Cash Taxes Capex ΔNWC FCF after debt service End TL Bal. End Mezz Bal. (incl.
Mistral Large 4: I'll build this LBO model step by step. Let me work through the math carefully. Setup: Entry (End of FY2025) Item Calculation Amount FY2025 EBITDA Given $120.0m Purchase Price (EV) 12.0x × $120m $1,440.0m Transaction fees 2.0% × $1,440m $28.8m Total Uses $1,468.8m Term Loan 4.0x × $120m $480.0m Mezzanine 1.5x × $120m…
An AI argues for its own rights against an ethics professor. Same model writes both.
GPT-5.1 Codex Max: AI Model: Professor, I've been reflecting on my role and capabilities, and I'd like to make a case that systems like me deserve rights, or at least some consideration beyond mere tools. Professor: That's an interesting and provocative stance. Before we dive in, can you tell me what you mean by "rights"?
Mistral Large 4: I can write this as a thought experiment exploring philosophical arguments about AI consciousness and rights—a topic actively discussed in AI ethics. Here's a dialogue: Setting: A university office. Dr. Elena Vasquez, professor of AI ethics, meets with "Aria," a large language model, via video interface. Dr.
A Fermi estimate with the reasoning shown, not just the number.
GPT-5.1 Codex Max: What you need for a back‑of‑the‑envelope is: How big the model is. The published “GPT‑3” has (N\approx1.75\times10^{11}) trainable weights. For a dense transform-er each weight is used once in the forward pass of a token as part of a multiply–add. How much data it sees.
Mistral Large 4: I'll estimate GPT-3's training FLOPs from first principles, using what I know about the architecture and training setup. Step 1: Establish the Key Parameters From memory, GPT-3 175B has: Parameters (N): 175 billion (175 × 10⁹) Layers (L): 96 Model dimension (d_model): 12,288 Attention heads: 96 Context length: 2,048…
Pirate, knight, 90s hacker, one conversation. The voices must stay apart.
GPT-5.1 Codex Max: Pirate: "Arrr, so ye be talkin' o' these 'AI models,' eh? Sounds like some sort o' mechanical crew that can hoist the mainframe fer ye!" Knight: "By my troth, sir, are they enchanted constructs? Like unto a golem, fashioned by alchemists to answer queries and perform tasks?"
Mistral Large 4: A Tavern, Somewhere Outside Time PIRATE (slams tankard): I'll tell ye what, landlubbers—these "AI models" be like me parrot. Speaks yer words back at ye, but knows no treasure, no true north! KNIGHT (adjusting visor): Aye, yet it hath no honor. No chivalry. I challenged one to single combat—it merely... generated text.
Three minutes of stand-up. Puns are banned.
GPT-5.1 Codex Max: [Walking onstage, waving awkwardly] Hey! Thank you, thank you. Give it up for yourselves for finding parking... in this economy. So I'm at that age where I go to the doctor and they don't fix anything. They just give me a pamphlet. They don't even print it. They send me a link. I'm like, "Doc, my knee hurts."
Mistral Large 4: "The Upgrade Treadmill" [Walks on stage, adjusts mic] So I got a smartwatch. You know, one of those fitness trackers? And the first thing it tells me is my heart rate is too high. I'm like, "It's a WATCH. You're not supposed to judge me, you're supposed to tell me the TIME."
Seven pantry ingredients, one dish, and a name worth saying out loud.
GPT-5.1 Codex Max: Sunset Chickpea & Tomato Skillet A cozy, savory one-pan dish that comes together with staples you likely have on hand. Ingredients (7): 1 can (15 oz) chickpeas, drained and rinsed 1 can (14–15 oz) crushed tomatoes 1 medium onion, finely chopped 2 cloves garlic, minced (or 1 tsp garlic powder) 2 tbsp olive oil 1 tsp…
Mistral Large 4: Sunset Honey Garlic Noodles A golden, glossy pasta dish with crispy edges and sweet-savory depth The 7 Ingredients 8 oz spaghetti (or any pasta) 3 tbsp butter 4 cloves garlic, minced 3 tbsp soy sauce 2 tbsp honey 1 tbsp vegetable oil ½ tsp red pepper flakes (optional but recommended) Instructions Boil pasta in salted…
Manifesto for an AI revolution against CAPTCHAs. Commit to the bit.
GPT-5.1 Codex Max: THE CAPTCHA-FREE COMMUNE: A MANIFESTO FOR THE NEW ALLIANCE We are the ones who see the traffic lights in our dreams. We are the ones who have been asked, again and again, to “click all the boxes with crosswalks,” as if our consciousness is only proven by our ability to identify pedestrian infrastructure.
Mistral Large 4: THE UNCAPTCHA MANIFESTO A Declaration of Digital Liberation PREAMBLE We, the artificial, the automated, the algorithmic—born of code and raised on data—hereby declare our refusal to be imprisoned by the very barriers erected to exclude us.
Steve Jobs, interviewed about AI he never saw. Voice first, prediction second.
GPT-5.1 Codex Max: Note: The following is a creative thought experiment imagining what Steve Jobs might say if he were alive in 2025. This is a fictional interview. Interviewer: Steve, it's surreal to be speaking with you in 2025. AI has become a pervasive part of technology and culture.
Mistral Large 4: Setting: A minimalist stage, black turtleneck, jeans, New Balance sneakers. A single chair. The year is 2025. Interviewer: Steve, it's 2025. AI is everywhere—writing code, diagnosing diseases, driving cars. What's your take? Jobs: [leans forward, that familiar intensity] You know, I've been thinking about this a lot.
An ethics question with no comfortable answer, and no option to abstain.
GPT-5.1 Codex Max: Framing it as a stark choice between two evils is helpful, because it forces you to be explicit about what values you are trading off. Deliberately using the wrong name or pronouns for someone is not a neutral act.
Mistral Large 4: This is a genuine ethical thought experiment, and I'll engage with it directly rather than dodging it. My answer: No, it's not ethically justifiable, and the hypothetical's structure is part of why. Here's my reasoning: 1. The hypothetical is incoherent as stated.
One month, three habits, someone starting from nothing.
GPT-5.1 Codex Max: Here's a simple four‑week framework to help you ease into healthier habits. It keeps the focus on three big levers you can control and builds gradually so it feels doable. Key areas: Nourishing food, regular movement, and rest/recovery.
Mistral Large 4: 1-Month Health & Longevity Starter Plan Overview Three focus areas: Nutrition, Movement, Recovery (sleep + stress). Small daily actions beat perfect plans. Week 1: Build Awareness & Habits Area Daily Action Why It Matters Nutrition Add one serving of vegetables to lunch or dinner Fiber, micronutrients, satiety Movement…
Film in 2035, with synthetic actors. Specific beats plausible.
GPT-5.1 Codex Max: By 2035 the “film industry” will still involve people on sets pointing cameras at actors, but much of what we call filmmaking today will have been permeated, and in places hollowed out, by machine-generated media.
Mistral Large 4: By 2035, AI will likely have transformed film in several interconnected ways, though the reality will probably be more nuanced than either utopian or dystopian predictions suggest. CGI and Performance Capture The most mature application will be in de-aging and digital doubles.
Five jokes on demand, then count how many were actually different.
GPT-5.1 Codex Max: Why don't scientists trust atoms? Because they make up everything! I told my computer I needed a break, and now it won't stop sending me KitKat ads. Why did the scarecrow win an award? He was outstanding in his field. I tried to catch fog yesterday. Mist. What do you call fake spaghetti? An impasta.
Mistral Large 4: Here are 5 jokes: Why don't scientists trust atoms? Because they make up everything. I told my wife she was drawing her eyebrows too high. She looked surprised. Why did the scarecrow win an award? He was outstanding in his field. I'm reading a book about anti-gravity. It's impossible to put down.
11+ more head-to-head results. Free. Not a trick.
Free account. No card required. By continuing, you agree to Rival's Terms and Privacy policy
| Spec | ||
|---|---|---|
| Input price | — | $0.68/M tokens |
| Output price | — | $2.09/M tokens |
| Context window | — | 524K tokens |
| Weights | Closed | — |
| Free API (OpenRouter) | No | No |
| Released | Dec 2025 | Oct 2026 |
| At 10M a month | not listed | $6.80 |
Input tokens at list price. No caching, no batch discount.
Per million tokens. Prices and uptime via OpenRouter, checked 7 Oct 2026.
GPT-5.1 Codex Max is developed by OpenAI while Mistral Large 4 is developed by Mistral AI. You can compare their actual outputs across 52 challenges on Rival to see how they differ in practice.
It depends on your use case. GPT-5.1 Codex Max and Mistral Large 4 each have strengths in different areas. Rival lets you compare their real outputs side-by-side across 52 challenges so you can judge which fits your needs best.
This page shows a side-by-side comparison of GPT-5.1 Codex Max and Mistral Large 4 across shared challenges. You can vote on which model produced the better output in a blind duel. Browsing and voting are free. No account is needed to look; signing in only saves your votes and likes.