Skip to content
Rival
How it worksPrivacyTerms
Explore all of Rival

Explore

  • Compare Models
  • All Models
  • Image Comparison
  • Audio Comparison
  • Image Generation
  • Best AI For...
  • Arena
  • API Pricing
  • Challenges

Discover

  • SubjectiveBench
  • Research
  • Research downloads
  • Rival Kits
  • Find your AI taste
  • UI Glow-Up
  • VoiceLock
  • Cost Cutter
  • Agent skills
  • Benchmarks vs Vibes
  • Jailbreak
  • Model Updates
  • Provider Status
  • AI Creators

Connect

  • Methodology
  • Advertise
  • Partnerships
  • Privacy Policy
  • Terms
  • RSS Feed
  1. Home
  2. Creators
  3. Inception
Loading...

Model Evolution

One challenge, every Inception generation.

Rival
How it worksPrivacyTerms
Explore all of Rival

Explore

  • Compare Models
  • All Models
  • Image Comparison
  • Audio Comparison
  • Image Generation
  • Best AI For...
  • Arena
  • API Pricing
  • Challenges

Discover

  • SubjectiveBench
  • Research
  • Research downloads
  • Rival Kits
  • Find your AI taste
  • UI Glow-Up
  • VoiceLock
  • Cost Cutter
  • Agent skills
  • Benchmarks vs Vibes
  • Jailbreak
  • Model Updates
  • Provider Status
  • AI Creators

Connect

  • Methodology
  • Advertise
  • Partnerships
  • Privacy Policy
  • Terms
  • RSS Feed
Inception

Inception: every model, side by side

Builds Mercury, a diffusion-based large language model (dLLM).

Total Models

3

Text Models

3

Active Period

Jun 2025 to Aug 2026

Developed Mercury, the world's first diffusion LLM (dLLM).

5-10x faster inference than traditional autoregressive models.

Backed by UAE sovereign investment.

Compare Inception Models

Mercury 2.5 Preview

Aug 2026

Inception's diffusion language model generates and refines tokens in parallel instead of one at a time. It reaches a reported 1,107 tokens per second on standard GPUs, with a 260K-token context window and a claimed 10-point intelligence gain over Mercury 2.

conversationreasoningcode-generationanalysistool-useagentic-tool-use

Mercury 2

Mar 2026

Inception's Mercury 2, the first reasoning diffusion language model. Instead of emitting tokens one at a time it produces and refines many in parallel, clearing 1000 tokens per second on standard GPUs, over 5x faster than Claude 4.5 Haiku and GPT-5 Mini. 128K context, native tool use, schema-aligned JSON.

conversationreasoningcode-generationanalysistool-use

Mercury

Jun 2025

The first diffusion large language model. Discrete diffusion runs 5 to 10x faster than speed-tuned models like GPT-4.1 Nano and Claude 3.5 Haiku while matching them, which is what makes voice agents and search viable.

conversationreasoningcode-generationanalysis