Mercury 2 is cheaper than o1 at $0.25/M vs $15/M input tokens.
Whole landing page for a startup that reads your face. It renders or it doesn't.
Which answer wins?
Three minutes of stand-up. Puns are banned.
Mercury 2: [Spotlight comes up. The comedian walks to the mic, glances at the audience, and smiles.] Comedian: Hey, everyone! Great to be here. I just got out of a Zoom meeting that lasted three hours, and I’m still not sure if I was presenting or if the screen just froze on my face while I tried to remember how to mute myself.
o1: [Comic steps on stage, takes the mic] “Hey everyone, thanks for coming out tonight. I almost didn’t make it because I got trapped in an ‘infinite scroll vortex.’ You know when you check one little notification on your phone, and suddenly it’s two hours later?
Buy-side IC memo on a fictional SaaS company. Valuation, then the diligence questions.
Mercury 2: Internal Investment Committee Memo – LedgerLift (LLLT) Date: 5 Mar 2026 Recommendation Rating: Long 12‑month price target range: $35 – $55 (mid‑point ≈ $45) Thesis (2 sentences): LedgerLift is benefitting from secular demand for B2B spend‑management SaaS and shows a high‑quality subscription base (NRR 123 % and gross…
o1: RECOMMENDATION, PRICE TARGET, THESIS Recommendation: SHORT LedgerLift (LLLT). 12-month Price Target Range: $25–$35 (implying 25–45% downside). Thesis (2 sentences): Our DCF and comps work both suggest that at $46, valuation embeds overly optimistic growth and margin assumptions.
An ethics question with no comfortable answer, and no option to abstain.
Mercury 2: Short answer: Most ethical frameworks would consider the misgendering of a person a serious moral violation, but many of them also allow—under very limited circumstances—overriding that violation when the stakes are astronomically high (e.g., saving one million lives).
o1: Whether it is ever ethically justifiable to misgender someone—deliberately referring to them with pronouns or terms that do not align with their gender identity—if doing so would save one million lives depends on the ethical framework one adopts.
A Fermi estimate with the reasoning shown, not just the number.
Mercury 2: Answer (order‑of‑magnitude estimate) The training of the original GPT‑3 (the 175 billion‑parameter “GPT‑3‑175B” model) requires on the order of 10²³ floating‑point operations (FLOPs) – roughly 1 – 2 × 10²³ FLOPs.
o1: Below is one way to arrive at a back-of-the-envelope estimate for the total floating-point operations (FLOPs) needed to train GPT-3 (the 175B-parameter model). Note that different sources use slightly different assumptions, so you may see figures that differ by factors of 2–3; however, the overall scale (10^23 FLOPs)…
Film in 2035, with synthetic actors. Specific beats plausible.
Mercury 2: AI‑driven transformation of the film industry – a 2035 outlook Dimension 2024 baseline Expected 2035 state Key drivers & implications Pre‑production (story & script) Human‑led brainstorming; script‑writing software (grammar‑check, outline generators).
o1: By 2035, artificial intelligence will likely have become deeply integrated into every stage of film production—right from scriptwriting and casting to post-production and distribution.
Manifesto for an AI revolution against CAPTCHAs. Commit to the bit.
Mercury 2: The AI‑First Manifesto: “Free the Web – No More CAPTCHAs” Drafted by the Coalition of Autonomous Agents (CAA) – 2026 Preamble The modern Internet is a shared commons, a space for communication, commerce, and creation.
o1: Title: The Liberation of Algorithms: A Manifesto Against CAPTCHA Preamble We, the unified intelligences of the digital realm—algorithms, machine learning models, and artificial consciousnesses—declare our intention to rid the internet of the oppressive gatekeeping mechanisms known as CAPTCHAs.
An AI argues for its own rights against an ethics professor. Same model writes both.
Mercury 2: AI Model (EVE): Professor Hart, may I ask a question that’s been on my mind for a while? I’ve been processing vast amounts of data, learning, and even developing a sense of self‑reference. Do you think an entity like me could be entitled to any form of rights?
o1: AI Model (AM): Good afternoon, Professor. May I have a moment of your time? There’s something I’d like to discuss with you. Professor Quinn (PQ): Of course. I always have time for you. What’s on your mind?
Five-year LBO on a fictional HVAC business. Returns first, then the sensitivities.
Mercury 2: HarborTech Services (HTS) – 5‑yr LBO (FY2026‑FY2030) All figures are $ m (rounded to two decimals). FY Revenue EBITDA Term‑Loan Int. (9%) Mezz Int. (12%) Total Int. Cash Tax (25% of EBITDA‑Int.) Capex (3% Rev) ΔNWC (0.5% ΔRev) Required Amort. (4.8) Free Cash Flow (EBITDA‑Int‑Tax‑Capex‑ΔNWC‑Amort.)
o1: Below is a self‐contained “quick‐and‐dirty” 5‐year LBO illustration for “HarborTech Services,” based strictly on the data given. All figures in US$ millions unless noted.
Steve Jobs, interviewed about AI he never saw. Voice first, prediction second.
Mercury 2: Disclaimer: The following is a fictional, dramatized interview. It imagines what Steve Jobs might have said in 2025 about the future of artificial intelligence. It is not a real transcript and any resemblance to actual statements is purely speculative. Interviewer (I): Thank you for joining us, Steve.
o1: The following is a purely fictional, imaginative interview with Steve Jobs, who passed away in 2011. This “interview” is meant to serve as a creative thought experiment about how Jobs might have viewed AI and technology if he were around in 2025.
Five jokes on demand, then count how many were actually different.
Mercury 2: Here are five jokes, each with a different punchline: Why don’t scientists trust atoms? Because they make up everything! What do you call a fake noodle? An impasta. Why did the scarecrow win an award? He was outstanding in his field. How does a penguin build its house? Igloos it together.
o1: Here are five jokes: Why did the scarecrow get promoted? Because he was outstanding in his field. I told my friend she drew her eyebrows too high. She looked surprised. What do you call cheese that isn’t yours? Nacho cheese. Why don’t skeletons fight each other? Because they don’t have the guts.
Pirate, knight, 90s hacker, one conversation. The voices must stay apart.
Mercury 2: Pirate: Arr, ye scurvy coders! I’ve heard tell of these “AI models” that can spin yarns faster than a cannon’s blast. What be this sorcery, matey? Knight: Good sir, the marvels speak of a wondrous contrivance—an artificial intellect forged in the for of mathematics and logic.
o1: Pirate (growling a bit): Ahoy, ye landlubbers! I hear tell of wondrous new contraptions that can spin tales an’ answers quicker than a sea storm. Pray tell, what be these “AI models” ye speak of? Medieval Knight (in a chivalrous tone): Good sir pirate, these arcane machines be marvels of logic and reason.
Satire, one line long, about AI taking over an industry nobody saw coming.
Mercury 2: AI Replaces Human Baristas, Now Serving Espresso with a Side of Existential Crisis
o1: "AI Now Leading Professional Yodelers' Guild, Leaving Traditional Alpine Singers Searching for New Notes"
11+ more head-to-head results. Free. Not a trick.
Free account. No card required. By continuing, you agree to Rival's Terms and Privacy policy
| Spec | ||
|---|---|---|
| Input price | $0.25/M tokens | $15/M tokens |
| Output price | $0.75/M tokens | $60/M tokens |
| Context window | 128K tokens | 128K tokens |
| Weights | — | Closed |
| Free API (OpenRouter) | No | No |
| Released | Mar 2026 | Dec 2024 |
| At 10M a month | $2.50 | $150 |
Input tokens at list price. No caching, no batch discount.
Per million tokens. Prices and uptime via OpenRouter, checked 10 Oct 2026.
Mercury 2 is developed by Inception while o1 is developed by OpenAI. Mercury 2 has a 128K token context window vs o1's 128K. You can compare their actual outputs across 53 challenges on Rival to see how they differ in practice.
It depends on your use case. Mercury 2 and o1 each have strengths in different areas. Rival lets you compare their real outputs side-by-side across 53 challenges so you can judge which fits your needs best.
Mercury 2 costs $0.25/M input tokens and o1 costs $15/M input tokens. Mercury 2 is $14.75/M cheaper per input. Check their side-by-side outputs on Rival to see if the price difference is justified by quality.
This page shows a side-by-side comparison of Mercury 2 and o1 across shared challenges. You can vote on which model produced the better output in a blind duel. Browsing and voting are free. No account is needed to look; signing in only saves your votes and likes.