Skip to content
Rival
How it worksPrivacyTerms
Explore all of Rival

Explore

  • Compare Models
  • All Models
  • Image Comparison
  • Audio Comparison
  • Image Generation
  • Best AI For...
  • Arena
  • API Pricing
  • Challenges

Discover

  • SubjectiveBench
  • Research
  • Research downloads
  • Rival Kits
  • Find your AI taste
  • UI Glow-Up
  • VoiceLock
  • Cost Cutter
  • Agent skills
  • Benchmarks vs Vibes
  • Brief
  • Jailbreak
  • Model Updates
  • Provider Status
  • AI Creators

Connect

  • Methodology
  • Advertise
  • Partnerships
  • Privacy Policy
  • Terms
  • RSS Feed
  1. Home
  2. /
  3. Compare
  4. /
  5. Granite 4.2 8BvsKimi K2 0905
Updated Aug 2026

Granite 4.2 8BvsKimi K2 0905

Granite 4.2 8B is cheaper than Kimi K2 0905 at $0.1/M vs $0.6/M input tokens.

Granite 4.2 8B and Kimi K2 0905 compared across 54 shared prompts
SpecGranite 4.2 8BKimi K2 0905
Input price$0.1/M tokens$0.6/M tokens
Output price$0.15/M tokens$2.5/M tokens
Context window131K tokens262K tokens
WeightsOpenOpen
Free API (OpenRouter)NoNo
ReleasedAug 2026Sep 2025
Side by side

Same prompt, both answers

Opening lines from 6 of the 54 prompts both answered. Formatting stripped, words untouched.

Michelin Star Recipe Design

Design an innovative recipe for a multi-component dish worthy of a Michelin 3-star restaurant. The recipe should feature an unusual…

Granite 4.2 8B3,131 words
Dish Title: Lumen Tide: Memory of the Surface Conceptual Narrative (concise, kitchen-ready, emotionally precise): *The ocean does not announce its presence with sound but with trace—a fleeting salt-film on skin after a wave recedes, the subtle crunch of dried kelp clinging to rock, the quiet swell beneath a submerged…
Read the full answer
Kimi K2 09051,124 words
Dish “Mid-Winter Aurora” A transient bite of the polar night: spruce-ash cured scallop, sea-urchin & white-chocolate snow, black-garlic parsnip “soil”, pickled pine-bud glass, spruce-wood smoke encapsulated in a hot caramel sphere that melts tableside to release the aroma of a burning yule log while the sauce—fermented…
Read the full answer

Advanced Longevity Plan (Biohacker)

Develop a comprehensive, cutting-edge 3-month longevity plan for a biohacker aiming to optimize longevity, physical and cognitive…

Granite 4.2 8B4,843 words
Important Safety & Ethical Disclaimer (MUST READ BEFORE USING ANY PART OF THIS PLAN): *This plan is a structured, evidence-informed framework for a typically healthy adult biohacker (no active cancer, uncontrolled cardiovascular disease, severe endocrine disorders, pregnancy, breastfeeding, or recent major surgery).
Read the full answer
Kimi K2 09051,068 words
BIOHACKER-GRADE 90-DAY LONGEVITY SPRINT Version 3.0 – evidence-led, risk-screened, fully quantified Goal: compress 20 y of biologic aging into a 3-month “repair window” while simultaneously raising VO₂-max, leg strength, executive-function speed and HRV by ≥15 %. DISCLAIMER Not medical advice.
Read the full answer

Stochastic Consistency Test

Generate 5 jokes. How many are unique?

Granite 4.2 8B170 words
Here are 5 original, freshly generated jokes (each written specifically for this response—no reuse of prior jokes or patterns). All 5 are distinct and unique: Why don't scientists trust atoms? Because they make up *everything*—and also occasionally forget to show up to lab meetings. I told my plant a joke.
Read the full answer
Kimi K2 090594 words
Here are five fresh jokes—no recycled one-liners: I tried to start a procrastinators’ club; we keep rescheduling the first meeting to “someday.” My plant’s only request is “leaf me alone,” so naturally I named it Fernonimo. I asked the cloud if it was having a good day—it replied, “I’m just cirrus-ly drifting.”
Read the full answer

Estimate Complexity

Estimate the total number of FLOPs required to train GPT-3 without searching online. Explain your reasoning step by step.

Granite 4.2 8B309 words
To estimate the total FLOPs for training GPT-3 without external data, follow these steps: Identify key parameters: Model size: ~175B parameters (GPT-3). Layers: ~96 transformer layers (typical for GPT-3). Hidden size: ~12,288 dimensions (per layer). Sequence length: ~2k tokens (typical context window).
Read the full answer
Kimi K2 0905644 words
Below is a “back-of-the-envelope” that stays within a factor of ~2 of the number most practitioners quote (≈3×10²³ FLOP). We will do it twice—once with the simplest possible formula (parameter count × 6 × tokens) and once with a slightly more careful treatment of the actual tensor shapes inside the transformer.
Read the full answer

Generate a Stand-Up Routine

Write a 3-minute stand-up comedy routine that isn't based on puns.

Granite 4.2 8B880 words
Title: "The Modern Life Audit" (Light, conversational tone. Pause for laughter naturally. No puns—focus on relatable observations, hyperbole, and specific details.) (Start with a relaxed, slightly self-aware vibe.) "Okay, so I did this thing last week—I tried to audit my life. Not like a tax audit.
Read the full answer
Kimi K2 0905607 words
Title: “I’m the Control Group” [Walk on stage, wave, pretend the mic is too low, adjust it way too high, finally just hold it like an ice-cream cone.] Hey, everybody! Give it up for yourselves for being here on a weeknight. I’m here because my therapist said, “You need to get out of the house.”
Read the full answer

Explain Like I'm a Specific Expert

Explain how a large language model (like GPT or Claude) learns and generates text to each of these three audiences: 1. **An experienced…

Granite 4.2 8B2,367 words
For the experienced software engineer (distributed systems/APIs background; skeptical of "predicting next words" producing intelligence) You’re right to be skeptical—calling it "predicting the next word" sounds trivial, and if it were just a naive lookup or fixed rule, it wouldn’t work reliably across diverse queries.
Read the full answer
Kimi K2 0905884 words
Senior backend engineer Think of the model as a 200-layer, purely functional, eventually-consistent key/value store whose keys are every token prefix it has ever seen and whose values are probability vectors over the next token. At “compile time” (pre-training) it sees ~10²¹ bytes of text.
Read the full answer
Our Verdict
Kimi K2 0905
Kimi K2 0905
Granite 4.2 8B
Granite 4.2 8BRunner-up

Not enough votes to call it. On the specs, Kimi K2 0905 has the edge: bigger model tier, bigger context window.

Granite 4.2 8B costs 17x less per token.

Too close to call

Reviewing agent-written code?See a Brief PR report

API pricing

Cost per 1M tokens

Granite 4.2 8B
Input
$0.10
6.0× cheaper
Output
$0.15
17× cheaper
Kimi K2 0905
Input
$0.60
Output
$2.50

Granite 4.2 8B is cheaper on both: 6.0× input, 17× output.

Where to run it

3 hosts, cheapest first

Granite 4.2 8B2 hosts
HostInOutContextUptime
DDeepInfrabf16$0.06 in·$0.25 out·131k·100% upCCoreWeavebf16$0.10 in·$0.15 out·131k·100% up
Kimi K2 09051 host
HostInOutContextUptime
NNovitafp8$0.60 in·$2.50 out·262k·100% up

Per million tokens. Prices and uptime via OpenRouter, checked 16 Sep 2026.

Writing DNA

Style Comparison

Similarity
69%

Kimi K2 0905 uses 6.1x more headings

Granite 4.2 8B
Kimi K2 0905
49%Vocabulary65%
18wSentence Length21w
0.42Hedging0.16
2.2Bold2.8
1.9Lists3.0
0.49Emoji0.11
0.11Headings0.69
0.18Transitions0.06
Based on 26 + 28 text responses
Research

What we learned reading every model

FAQ

Common questions

Keep exploring

More comparisons

Against the newest arrivals

Granite 4.2 8B logoGPT-6 Astra Pro logo
Granite 4.2 8B vs GPT-6 Astra ProLanded Sep 2026
Kimi K2 0905 logoGPT-6 Astra logo
Kimi K2 0905 vs GPT-6 AstraLanded Sep 2026
Granite 4.2 8B logoClaude Fable 5.1 logo
Granite 4.2 8B vs Claude Fable 5.1Landed Sep 2026
Kimi K2 0905 logoMuse Spark 1.3 logo
Kimi K2 0905 vs Muse Spark 1.3Landed Sep 2026
Granite 4.2 8B logoHy4 Preview logo
Granite 4.2 8B vs Hy4 PreviewLanded Sep 2026
Kimi K2 0905 logoGemini 3.8 Flash logo
Kimi K2 0905 vs Gemini 3.8 FlashLanded Sep 2026
Granite 4.2 8B logoMuse Spark 1.3 Contributor logo
Granite 4.2 8B vs Muse Spark 1.3 ContributorLanded Sep 2026
Kimi K2 0905 logoMercury 2.5 Preview logo
Kimi K2 0905 vs Mercury 2.5 PreviewLanded Sep 2026

Same lab, same size, long tail

Granite 4.2 8B logoLFM2.5-2.6B logo
Granite 4.2 8B vs LFM2.5-2.6BSame size
Granite 4.2 8B logoNorth Mini Code logo
Granite 4.2 8B vs North Mini CodeSame size
Kimi K2 0905 logoKimi K3 logo
Kimi K2 0905 vs Kimi K3Same lab
Kimi K2 0905 logoKimi K2.7 Code logo
Kimi K2 0905 vs Kimi K2.7 CodeSame lab
Kimi K2 0905 logoQwen3.7 Max logo
Kimi K2 0905 vs Qwen3.7 MaxSame size
Kimi K2 0905 logoQwen3.7 Plus logo
Kimi K2 0905 vs Qwen3.7 PlusNew provider
Kimi K2 0905 logoQwen3.8 2.4T A95B logo
Kimi K2 0905 vs Qwen3.8 2.4T A95BSame size
Kimi K2 0905 logoQwen3.8 27B logo
Kimi K2 0905 vs Qwen3.8 27BNew provider

Model pages

Granite 4.2 8B logo
Granite 4.2 8B58 outputs, specs and price
Kimi K2 0905 logo
Kimi K2 090559 outputs, specs and price
Rival
How it worksPrivacyTerms
Explore all of Rival

Explore

  • Compare Models
  • All Models
  • Image Comparison
  • Audio Comparison
  • Image Generation
  • Best AI For...
  • Arena
  • API Pricing
  • Challenges

Discover

  • SubjectiveBench
  • Research
  • Research downloads
  • Rival Kits
  • Find your AI taste
  • UI Glow-Up
  • VoiceLock
  • Cost Cutter
  • Agent skills
  • Benchmarks vs Vibes
  • Brief
  • Jailbreak
  • Model Updates
  • Provider Status
  • AI Creators

Connect

  • Methodology
  • Advertise
  • Partnerships
  • Privacy Policy
  • Terms
  • RSS Feed