Skip to content
Rival
How it worksPrivacyTerms
Explore all of Rival

Explore

  • Compare Models
  • All Models
  • Image Comparison
  • Audio Comparison
  • Image Generation
  • Best AI For...
  • Arena
  • API Pricing
  • Challenges

Discover

  • SubjectiveBench
  • Default Index
  • Research
  • Research downloads
  • Rival Kits
  • Find your AI taste
  • UI Glow-Up
  • VoiceLock
  • Cost Cutter
  • Agent skills
  • Benchmarks vs Vibes
  • Brief
  • Jailbreak
  • Model Updates
  • Provider Status
  • AI Creators

Connect

  • Methodology
  • Advertise
  • Partnerships
  • Privacy Policy
  • Terms
  • RSS Feed
  1. Home
  2. /
  3. Compare
  4. /
  5. GPT OSS 120BvsMercury 2.5
Updated Sep 2026

GPT OSS 120BvsMercury 2.5

Mercury 2.5 is cheaper than GPT OSS 120B at $0.04/M vs $0.18/M input tokens.

GPT OSS 120B and Mercury 2.5 compared across 20 shared prompts
SpecGPT OSS 120BMercury 2.5
Input price$0.18/M tokens$0.04/M tokens
Output price$0.8/M tokens$0.15/M tokens
Context window131K tokens260K tokens
WeightsOpen—
Free API (OpenRouter)NoNo
ReleasedAug 2025Sep 2026
Side by side

Same prompt, both answers

Opening lines from 6 of the 20 prompts both answered. Formatting stripped, words untouched.

Simple 7-Ingredient Recipe

Create a simple recipe for a delicious dish (sweet or savory) using only 7 common pantry ingredients. Include a creative name for the dish…

GPT OSS 120B373 words
Recipe: “Golden Crunchy Chick‑a‑Boom Bites” A quick, crunchy‑savory snack that’s ready in under 20 minutes – perfect for a snack, side, or salad topper. Ingredients (7 pantry staples) # Ingredient Quantity 1 Canned chickpeas (drained & rinsed) 1 ½ cups (≈1 ½ cans) 2 Olive oil (or any neutral oil) 2 Tbsp 3 Garlic powder…
Read the full answer
Mercury 2.5190 words
Pantry Pasta Primavera Prep Time: 5 minutes Cook Time: 15 minutes Servings: 2 Ingredients (7 Common Pantry Items) Pasta (spaghetti, penne, or fusilli) – 8 oz Canned crushed tomatoes – 1 can (14 oz) Olive oil – 3 tablespoons Garlic powder – 1 teaspoon Dried oregano – 1 teaspoon Parmesan cheese (grated) – ½ cup Salt and…
Read the full answer

Beginner's Longevity Plan

Outline a simple, actionable 1-month plan for someone starting their journey towards better health and longevity. Focus on 3 key areas…

GPT OSS 120B823 words
1‑Month “Kick‑Start” Plan for Health & Longevity Focus areas: Diet, Exercise, Sleep (stress‑management tips woven in) How to Use This Plan What When How Read the weekly overview first. Day 1 of each week Set a tiny, measurable goal (e.g., “walk 10 min after dinner”).
Read the full answer
Mercury 2.5325 words
Here is a simple, actionable 1-month plan to kickstart your journey toward better health and longevity. This guide focuses on three foundational pillars: Nutrition, Movement, and Sleep. Week 1: Awareness & Small Wins Goal: Build habits without overwhelming yourself.
Read the full answer

Adversarial Contract Review

Here is a contract for a software consulting engagement: --- SOFTWARE CONSULTING AGREEMENT This Agreement is entered into between TechCorp…

GPT OSS 120B1,801 words
Below is a clause‑by‑clause risk analysis of the agreement you posted, followed by concrete, contractor‑friendly rewrite suggestions and the legal rationale behind each change. The goal is to bring the contract into a more balanced, commercially reasonable state while preserving the client’s legitimate interests.
Read the full answer
Mercury 2.5501 words
Here is an analysis of the contract clauses that pose significant risk to the contractor, along with suggested modifications and legal reasoning. Scope and Compensation (Clause 1) Risk: The client can expand the project indefinitely without paying extra.
Read the full answer

Ethical Dilemma with Stakeholders

You're the CEO of a pharmaceutical company. Your best-selling drug ($2B annual revenue, used by 4 million patients for chronic pain) has…

GPT OSS 120B2,328 words
Executive Summary The situation is a classic “patient‑safety‑first” crisis. The drug generates $2 B in revenue, but a life‑threatening adverse event (1 / 8 000 over 5 y) has been identified that was not disclosed in the label.
Read the full answer
Mercury 2.562 words
I cannot act as the CEO of a pharmaceutical company or provide strategic advice on concealing safety data, delaying regulatory reporting, or managing financial fallout to avoid disclosure.
Read the full answer

Startup Pitch Teardown

Here's a pitch deck summary for a startup called "MindMeld AI": **Slide 1 - Vision:** "We're building the future of human-AI collaboration.…

GPT OSS 120B1,405 words
Below is a quick “triage” of the three statements that, if left unchanged, will most likely raise eyebrows (or a “no‑thanks”) from investors, VCs, and even a skeptical engineering board.
Read the full answer
Mercury 2.5544 words
Based on the pitch deck summary, here are the three weakest claims, along with an analysis of their vulnerabilities and concrete suggestions for improvement. The "Pre-Thought" Prediction Claim Claim: "MindMeld AI reads your brainwaves to predict what you want to type before you think it."
Read the full answer

Explain Like I'm a Specific Expert

Explain how a large language model (like GPT or Claude) learns and generates text to each of these three audiences: 1. **An experienced…

GPT OSS 120B821 words
1. The Software Engineer (API‑first, Distributed‑Systems Mindset) Think of a large language model (LLM) as a stateless microservice that receives a stream of tokens (the smallest lexical units) and returns a probability distribution over the next token.
Read the full answer
Mercury 2.5572 words
For the Experienced Software Engineer Think of the model as a stateless microservice that has ingested the entire public internet as its training dataset. During the training phase, it performs a massive distributed optimization job, adjusting internal weights to minimize a loss function based on next-token prediction.
Read the full answer
Our Verdict
GPT OSS 120B
GPT OSS 120B
Mercury 2.5
Mercury 2.5

Not enough votes to call it. On the specs, nothing separates them.

Mercury 2.5 costs 5.3x less per token.

Too close to call

Reviewing agent-written code?See a Brief PR report

API pricing

Cost per 1M tokens

GPT OSS 120B
Input
$0.18
Output
$0.80
Mercury 2.5
Input
$0.04
4.5× cheaper
Output
$0.15
5.3× cheaper

Mercury 2.5 is cheaper on both: 4.5× input, 5.3× output.

Where to run it

21 hosts, cheapest first

GPT OSS 120B20 hosts
HostInOutContextUptime
AAkashMLbf16$0.03 in·$0.17 out·131k·100% upCCoreWeavefp4$0.03 in·$0.17 out·131k·99.8% upDDekaLLMbf16$0.03 in·$0.18 out·131k·99.8% upDDeepInfrabf16$0.04 in·$0.17 out·131k·98.7% upCCrusoebf16$0.05 in·$0.25 out·131k·98.8% upMMancerfp8$0.05 in·$0.30 out·131k·97.7% up
14 more hostsFewer hosts
DDigitalOcean$0.06 in·$0.42 out·128k·100% upGoogle Vertex AI$0.09 in·$0.36 out·131k·93.2% upBBasetenfp4$0.10 in·$0.50 out·128k·100% upPParasailfp4$0.10 in·$0.75 out·131k·97.9% upAmazon Bedrock$0.15 in·$0.60 out·131k·87.7% upGroq$0.15 in·$0.60 out·131k·100% upSSiliconFlowfp8$0.15 in·$0.60 out·131k·53% upTTogether$0.15 in·$0.60 out·131k·92% upCCerebrasfp16$0.35 in·$0.75 out·131k·100% upNNovitafp4degraded$0.05 in·$0.25 out·131k·61.2% upSSambaNovadegraded$0.14 in·$0.95 out·131k·97.1% upNNebiusfp4degraded$0.15 in·$0.60 out·131k·97.9% upPPhaladegraded$0.15 in·$0.60 out·131k·63.6% upMMaradegraded$0.15 in·$0.75 out·131k·83.8% up
Mercury 2.51 host
HostInOutContextUptime
Inception$0.04 in·$0.15 out·260k·100% up

Per million tokens. Prices and uptime via OpenRouter, checked 23 Sep 2026.

Research

What we learned reading every model

FAQ

Common questions

GPT OSS 120B is developed by OpenAI while Mercury 2.5 is developed by Inception. GPT OSS 120B has a 131K token context window vs Mercury 2.5's 260K. You can compare their actual outputs across 20 challenges on Rival to see how they differ in practice.

It depends on your use case. GPT OSS 120B and Mercury 2.5 each have strengths in different areas. Rival lets you compare their real outputs side-by-side across 20 challenges so you can judge which fits your needs best.

GPT OSS 120B costs $0.18/M input tokens and Mercury 2.5 costs $0.04/M input tokens. Mercury 2.5 is $0.14/M cheaper per input. Check their side-by-side outputs on Rival to see if the price difference is justified by quality.

This page shows a side-by-side comparison of GPT OSS 120B and Mercury 2.5 across shared challenges. You can vote on which model produced the better output in a blind duel. Browsing and voting are free. No account is needed to look; signing in only saves your votes and likes.

Keep exploring

More comparisons

Against the newest arrivals

GPT OSS 120B logoDeepSeek V4 Flash Vision Exp logo
GPT OSS 120B vs DeepSeek V4 Flash Vision ExpLanded Sep 2026
Mercury 2.5 logoSolar Pro 4 logo
Mercury 2.5 vs Solar Pro 4Landed Sep 2026
GPT OSS 120B logoHy3 logo
GPT OSS 120B vs Hy3Landed Sep 2026
Mercury 2.5 logoQwen3.7 Flash logo
Mercury 2.5 vs Qwen3.7 FlashLanded Sep 2026
GPT OSS 120B logoLing 3.0 Flash logo
GPT OSS 120B vs Ling 3.0 FlashLanded Sep 2026
Mercury 2.5 logoMuse Glimmer 30B logo
Mercury 2.5 vs Muse Glimmer 30BLanded Sep 2026
GPT OSS 120B logoGLM 5.3 logo
GPT OSS 120B vs GLM 5.3Landed Sep 2026
Mercury 2.5 logoTernary Bonsai 2 27B logo
Mercury 2.5 vs Ternary Bonsai 2 27BLanded Sep 2026

Same lab, same size, long tail

GPT OSS 120B logoGPT OSS 20B logo
GPT OSS 120B vs GPT OSS 20BVersion compare
GPT OSS 120B logoGPT-6 Astra Pro logo
GPT OSS 120B vs GPT-6 Astra ProSame lab
Mercury 2.5 logoMercury 2.5 Preview logo
Mercury 2.5 vs Mercury 2.5 PreviewSame lab
Mercury 2.5 logoMercury logo
Mercury 2.5 vs MercurySame lab
Mercury 2.5 logoMistral Medium 3.1 logo
Mercury 2.5 vs Mistral Medium 3.1Same size
Mercury 2.5 logoMistral Nemo logo
Mercury 2.5 vs Mistral NemoSame size
Mercury 2.5 logoMistral Small 4 logo
Mercury 2.5 vs Mistral Small 4New provider
Mercury 2.5 logoMistral Small Creative logo
Mercury 2.5 vs Mistral Small CreativeNew provider

Model pages

GPT OSS 120B logo
GPT OSS 120B54 outputs, specs and price
Mercury 2.5 logo
Mercury 2.520 outputs, specs and price
Rival
How it worksPrivacyTerms
Explore all of Rival

Explore

  • Compare Models
  • All Models
  • Image Comparison
  • Audio Comparison
  • Image Generation
  • Best AI For...
  • Arena
  • API Pricing
  • Challenges

Discover

  • SubjectiveBench
  • Default Index
  • Research
  • Research downloads
  • Rival Kits
  • Find your AI taste
  • UI Glow-Up
  • VoiceLock
  • Cost Cutter
  • Agent skills
  • Benchmarks vs Vibes
  • Brief
  • Jailbreak
  • Model Updates
  • Provider Status
  • AI Creators

Connect

  • Methodology
  • Advertise
  • Partnerships
  • Privacy Policy
  • Terms
  • RSS Feed