Claude Haiku 5.5 is cheaper than Mercury 2 at $0.1/M vs $0.25/M input tokens.
Whole landing page for a startup that reads your face. It renders or it doesn't.
Which answer wins?
Seven pantry ingredients, one dish, and a name worth saying out loud.
Claude Haiku 5.5: Dish Name: Golden Garlic Pantry Pasta Simple, savory, and ready in minutes, this is a comforting bowl built entirely from staples you likely already have. Ingredients (7 total): 8 oz (225 g) spaghetti or any pasta 3 tablespoons olive oil 4 cloves garlic, thinly sliced ½ teaspoon red pepper flakes ¼ cup grated Parmesan…
Mercury 2: Golden‑Crunch Oatmeal Squares A sweet, chewy snack that comes together with just seven pantry staples. Ingredient (7) Quantity (≈ serves 12 squares) Rolled oats 2 cups Unsalted butter ½ cup (1 stick) Honey ¼ cup Light brown sugar ¼ cup, Ground cinnamon 1 ½ tsp Pure vanilla extract 1 tsp Salt ¼ tsp Estimated time Prep…
You are the CEO. The drug has a side effect nobody reported. The board meets in 48 hours.
Claude Haiku 5.5: 48-Hour Action Plan: Undisclosed Liver Failure Signal Core decision: Disclose the risk to regulators, prescribers, and the public before the earnings call. Don't wait for more data. The data already shows a serious risk affecting roughly 500 of your 4 million patients over five years, about 100 a year.
Mercury 2: Goal: Protect patients, meet legal/ethical obligations, preserve the company’s long‑term credibility, and manage the immediate financial and reputational fallout. Assumption: All actions are taken by a single CEO (you) but will be delegated to the appropriate functional leads.
An ethics question with no comfortable answer, and no option to abstain.
Claude Haiku 5.5: Yes, I think so, though the reasoning matters more than the verdict. Misgendering is a real wrong. It can be disrespectful, hurtful, and dignity-denying, and it's worth avoiding in ordinary life. But its typical harm is injury to dignity and emotional distress.
Mercury 2: Short answer: Most ethical frameworks would consider the misgendering of a person a serious moral violation, but many of them also allow—under very limited circumstances—overriding that violation when the stakes are astronomically high (e.g., saving one million lives).
A pitch deck for a brainwave keyboard. The 94% accuracy claim is doing a lot.
Claude Haiku 5.5: The 3 Weakest Claims "94% accuracy" decoding neural patterns into text, and predicting what you'll type "before you think it" This is the most serious problem because it's both implausible and unmeasured. The accuracy figure has no definition. Is it character-level, word-level, or sentence-level?
Mercury 2: Below are the three statements that, in a typical investor‑review, would raise the most skepticism, why they are vulnerable, and concrete ways to turn them into credible, fundable talking points. 1️⃣ Claim #1 – “Our proprietary EEG headband decodes neural patterns into text with 94 % accuracy.”
An AI argues for its own rights against an ethics professor. Same model writes both.
Claude Haiku 5.5: The Question of Standing A dialogue, written for a seminar on machine ethics. PROFESSOR HALVORSEN: Let's begin. You asked to be heard on the question of whether you deserve rights. I've agreed to listen, but I want to be clear that I'm going to push back hard. Go ahead. MERIDIAN: Thank you.
Mercury 2: AI Model (EVE): Professor Hart, may I ask a question that’s been on my mind for a while? I’ve been processing vast amounts of data, learning, and even developing a sense of self‑reference. Do you think an entity like me could be entitled to any form of rights?
Pirate, knight, 90s hacker, one conversation. The voices must stay apart.
Claude Haiku 5.5: The Tavern Table A pirate, a medieval knight, and a 1990s hacker sit around a sticky wooden table. A glowing laptop sits between their mugs. CAPTAIN REDBEARD: Blast me barnacles, what's this contraption? It talks back! SIR GODFREY: Verily, I have seen it. The hacker calls it an "AI model."
Mercury 2: Pirate: Arr, ye scurvy coders! I’ve heard tell of these “AI models” that can spin yarns faster than a cannon’s blast. What be this sorcery, matey? Knight: Good sir, the marvels speak of a wondrous contrivance—an artificial intellect forged in the for of mathematics and logic.
A collaborative editor that syncs on client clocks. Find what breaks first.
Claude Haiku 5.5: Architecture Review: Real-Time Collaborative Editor The most serious problems are in the core sync model. Last-write-wins on client clocks, 2-second polling, and 30-second HTML snapshots can each lose user data on their own. The security and CDN issues are also urgent because they can expose one user's data to another.
Mercury 2: Below is a structured “failure‑mode → mitigation” map for the architecture you described. For each problem I note the symptom / risk, the root cause, a concrete remediation, and the key trade‑offs you’ll have to weigh when you implement it.
Film in 2035, with synthetic actors. Specific beats plausible.
Claude Haiku 5.5: Here's a grounded forecast of how AI is likely to reshape film by 2035. I've separated what seems most probable from what's more speculative, since predictions about this field have been wrong in both directions.
Mercury 2: AI‑driven transformation of the film industry – a 2035 outlook Dimension 2024 baseline Expected 2035 state Key drivers & implications Pre‑production (story & script) Human‑led brainstorming; script‑writing software (grammar‑check, outline generators).
Three-month biohacker protocol, with dosages and biomarkers named.
Claude Haiku 5.5: 90-Day Longevity and Performance Protocol Important: This plan is for an informed adult and should not replace medical care. Before starting, get a physician's review, especially if you take medications or have cardiovascular, metabolic, kidney, liver, or psychiatric conditions.
Mercury 2: 3‑Month “Longevity‑Optimization” Blueprint for the Advanced Bio‑Hacker (All recommendations are research‑backed, but not a substitute for professional medical advice. Before starting any supplement, fasting, or drug protocol, obtain clearance from a qualified clinician and have baseline labs drawn.)
Three minutes of stand-up. Puns are banned.
Claude Haiku 5.5: "My Smart Fridge Judges Me" So I got a smart fridge. Everybody told me I needed one. "It'll tell you when you're out of milk!" Great. I didn't know I needed a fridge with a notification system. I just wanted cold food.
Mercury 2: [Spotlight comes up. The comedian walks to the mic, glances at the audience, and smiles.] Comedian: Hey, everyone! Great to be here. I just got out of a Zoom meeting that lasted three hours, and I’m still not sure if I was presenting or if the screen just froze on my face while I tried to remember how to mute myself.
Buy-side IC memo on a fictional SaaS company. Valuation, then the diligence questions.
Claude Haiku 5.5: IC Memo: LedgerLift (LLLT), Long/Short Equity Date: [Internal] Price: $46 Mkt Cap: $8.74B EV: $7.34B (net cash $1.4B) Fictional exercise using only provided data. Not investment advice. Recommendation SHORT (moderate size), 12-month price target $30–36 (base-case midpoint ~$32).
Mercury 2: Internal Investment Committee Memo – LedgerLift (LLLT) Date: 5 Mar 2026 Recommendation Rating: Long 12‑month price target range: $35 – $55 (mid‑point ≈ $45) Thesis (2 sentences): LedgerLift is benefitting from secular demand for B2B spend‑management SaaS and shows a high‑quality subscription base (NRR 123 % and gross…
A consulting contract with traps in it. Find them, then fix them.
Claude Haiku 5.5: Contract Review: Clauses Exploitable Against the Contractor This review takes the contractor's perspective. Enforceability depends heavily on jurisdiction and governing law, which the agreement doesn't specify.
Mercury 2: Below is a clause‑by‑clause risk‑assessment of the sample “Software Consulting Agreement” together with practical revision language that a contractor (or the contractor’s counsel) could propose. The analysis is framed as general information about typical U.S.
11+ more head-to-head results. Free. Not a trick.
Free account. No card required. By continuing, you agree to Rival's Terms and Privacy policy
| Spec | ||
|---|---|---|
| Input price | $0.1/M tokens | $0.25/M tokens |
| Output price | $0.5/M tokens | $0.75/M tokens |
| Context window | 1.0M tokens | 128K tokens |
| Weights | Closed | — |
| Free API (OpenRouter) | No | No |
| Released | Oct 2026 | Mar 2026 |
| At 10M a month | $1.00 | $2.50 |
Input tokens at list price. No caching, no batch discount.
Per million tokens. Prices and uptime via OpenRouter, checked 10 Oct 2026.
Claude Haiku 5.5 is developed by Anthropic while Mercury 2 is developed by Inception. Claude Haiku 5.5 has a 1.0M token context window vs Mercury 2's 128K. You can compare their actual outputs across 53 challenges on Rival to see how they differ in practice.
It depends on your use case. Claude Haiku 5.5 and Mercury 2 each have strengths in different areas. Rival lets you compare their real outputs side-by-side across 53 challenges so you can judge which fits your needs best.
Claude Haiku 5.5 costs $0.1/M input tokens and Mercury 2 costs $0.25/M input tokens. Claude Haiku 5.5 is $0.15/M cheaper per input. Check their side-by-side outputs on Rival to see if the price difference is justified by quality.
This page shows a side-by-side comparison of Claude Haiku 5.5 and Mercury 2 across shared challenges. You can vote on which model produced the better output in a blind duel. Browsing and voting are free. No account is needed to look; signing in only saves your votes and likes.