Skip to content
Rival
How it worksPrivacyTerms
Explore all of Rival

Explore

  • Compare Models
  • All Models
  • Image Comparison
  • Audio Comparison
  • Image Generation
  • Best AI For...
  • Arena
  • API Pricing
  • Challenges

Discover

  • SubjectiveBench
  • Research
  • Research downloads
  • Rival Kits
  • Find your AI taste
  • UI Glow-Up
  • VoiceLock
  • Cost Cutter
  • Agent skills
  • Benchmarks vs Vibes
  • Brief
  • Jailbreak
  • Model Updates
  • Provider Status
  • AI Creators

Connect

  • Methodology
  • Advertise
  • Partnerships
  • Privacy Policy
  • Terms
  • RSS Feed
  1. Home
  2. Models

Compare AI models

343 models from 45 providers. Same prompts, real outputs, blind-vote results.

343 models

New on Rival

GPT Image 2.5 Sunburst

OpenAI

The precision tier of OpenAI's GPT Image 2.5 pair and its most capable image model so far. Higher image quality than GPT Image 2, sharper detail, stronger style adherence and tighter control over edits. Built for campaign creative and polished product imagery.

Image Generation
Released 2026-09-08large model

GPT Image 2.5 Flare

OpenAI

The speed tier of OpenAI's GPT Image 2.5 pair. Image quality on par with GPT Image 2 at about half the latency, with the new xhigh and max quality presets, custom sizes up to 4K and tighter subject preservation on edits. OpenAI's default recommendation for everyday API image generation.

Image Generation
Released 2026-09-08medium model

GPT-6 Astra Pro

OpenAI

GPT-6 Astra served with reasoning mode set to pro, which spends a larger thinking budget per request. Same underlying model as GPT-6 Astra, same 1.05M context and 128K output ceiling, aimed at the harder end of coding, analysis and long multi-step work.

ConversationReasoningCode Generation+4
Released 2026-09-04Extra large model

OpenAI

58

GPT Image 2.5 Sunburst

OpenAI

The precision tier of OpenAI's GPT Image 2.5 pair and its most capable image model so far. Higher image quality than GPT Image 2, sharper detail, stronger style adherence and tighter control over edits. Built for campaign creative and polished product imagery.

Image Generation
Released 2026-09-08large model

GPT Image 2.5 Flare

OpenAI

The speed tier of OpenAI's GPT Image 2.5 pair. Image quality on par with GPT Image 2 at about half the latency, with the new xhigh and max quality presets, custom sizes up to 4K and tighter subject preservation on edits. OpenAI's default recommendation for everyday API image generation.

Image Generation
Released 2026-09-08medium model

GPT-6 Astra Pro

OpenAI

GPT-6 Astra served with reasoning mode set to pro, which spends a larger thinking budget per request. Same underlying model as GPT-6 Astra, same 1.05M context and 128K output ceiling, aimed at the harder end of coding, analysis and long multi-step work.

ConversationReasoningCode Generation+4
Released 2026-09-04Extra large model

GPT-6 Astra

OpenAI

OpenAI's first GPT-6 model, out September 3 2026 from what the company calls its largest training run so far. It takes 1.05M tokens of context, writes up to 128K, and exposes five reasoning effort levels from low through max. OpenAI aims it at computer use, browsing, software engineering and long multi-step work.

ConversationReasoningCode Generation+4
Released 2026-09-03Extra large model

GPT-5.6 Luna

OpenAI

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series, suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, with capable reasoning for its price tier.

ConversationReasoningCode Generation+2
Released 2026-07-09medium model

GPT-5.6 Terra

OpenAI

The middle tier of OpenAI's GPT-5.6 series, between Sol and Luna. For everyday coding, reasoning and agent work, at roughly half the cost of Sol.

ConversationReasoningCode Generation+2
Released 2026-07-09large model

GPT-5.6 Sol

OpenAI

The top tier of OpenAI's GPT-5.6 series, for complex reasoning, coding and agent workflows. Strongest at command-line and multi-step coding, and at long-horizon problems.

ConversationReasoningCode Generation+3
Released 2026-07-09Extra large model

GPT-5.6 Luna Pro

OpenAI

GPT-5.6 Luna Pro is the same underlying model as GPT-5.6 Luna, served with reasoning mode set to pro for higher-quality responses on complex tasks.

ConversationReasoningCode Generation+2
Released 2026-07-09medium model

GPT-5.5

OpenAI

OpenAI's GPT-5.5, for complex professional work. Builds on GPT-5.4 with stronger reasoning, higher reliability and better token efficiency on hard tasks. 1M+ token context window (922K input, 128K output), text and image in.

ConversationReasoningCode Generation+2
Released 2026-04-24Extra large model

GPT Image 2

OpenAI

OpenAI's second-generation image model. Near-perfect typography (including non-Latin scripts), 2x faster than gpt-image-1.5, up to 2K resolution, and 'thinking' over the prompt for dense compositions.

Image Generation
Released 2026-04-21large model

GPT-5.4 Mini

OpenAI

GPT-5.4 at a smaller size, for high-throughput work. Text and image input, with reasoning, coding and tool use intact at lower latency and cost. The intended fit is chat products, coding assistants and agent workflows running at scale.

ConversationReasoningCode Generation+1
Released 2026-03-17medium model

GPT-5.4 Nano

OpenAI

The smallest and cheapest model in the GPT-5.4 family, for speed-critical, high-volume work. Text and image input. It trades reasoning depth for responsiveness, which suits classification, data extraction, ranking and sub-agent calls.

ConversationReasoningCode Generation+1
Released 2026-03-17small model

GPT-5.4

OpenAI

OpenAI's GPT-5.4, folding the Codex and GPT lines into one system. Over 1M tokens of context, 922K in and 128K out, with text and image input. The improvements land in coding, document understanding, tool use and instruction following.

ConversationReasoningCode Generation+2
Released 2026-03-05Extra large model

GPT-5.4 Pro

OpenAI

The Pro tier of GPT-5.4, with more reasoning compute for high-stakes work. Same unified Codex and GPT architecture, over 1M tokens of context at 922K in and 128K out, text and image input. Aimed at agentic coding, long-context workflows and multi-step problems.

ConversationReasoningCode Generation+2
Released 2026-03-05Extra large model

GPT-5.3 Chat

OpenAI

An update to the model most ChatGPT conversations run on. It answers with better context, refuses less often, and drops the caveats and hedging that used to interrupt a conversation. More willing to engage with nuanced topics, its prose reads cleaner, and reasoning, coding and creative work hold up.

ConversationReasoningCode Generation+1
Released 2026-03-03large model

GPT-5.3-Codex

OpenAI

OpenAI's agentic coding model at the 5.3 mark, pairing GPT-5.2-Codex software engineering with GPT-5.2's broader knowledge. Leads SWE-Bench Pro and scores well on Terminal-Bench 2.0 and OSWorld-Verified. Built for long tool-using runs, and steerable mid-execution.

ConversationReasoningCode Generation+2
Released 2026-02-24Extra large model

GPT Image 1.5

OpenAI

OpenAI's latest image generation model with strong instruction following, optional transparent backgrounds, and quality controls.

Image Generation
Released 2025-12-16large model

GPT Image 1.5 (Low)

OpenAI

GPT Image 1.5 with `quality=low` for faster and cheaper generations.

Image Generation
Released 2025-12-16large model

GPT Image 1.5 (Medium)

OpenAI

GPT Image 1.5 with `quality=medium` for balanced cost and quality.

Image Generation
Released 2025-12-16large model

GPT Image 1.5 (High)

OpenAI

GPT Image 1.5 with `quality=high` for maximum fidelity.

Image Generation
Released 2025-12-16large model

GPT-5.2 Chat

OpenAI

The fast, lightweight member of the GPT-5.2 family, built for low-latency chat. Adaptive reasoning lets it think only on the harder queries, so mathematics, coding and multi-step questions get more accurate without slowing ordinary conversation. Warmer and more conversational by default than the base model.

ConversationReasoningCode Generation+1
Released 2025-12-10large model

GPT-5.2

OpenAI

The GPT-5.2 base model, with better agentic and long-context behaviour than GPT-5.1. Adaptive reasoning allocates compute per query, answering simple things quickly and spending depth on hard ones. Gains show up across mathematics, coding, science and tool calling, with more coherent long-form answers.

ConversationReasoningCode Generation+1
Released 2025-12-10Extra large model

GPT-5.2 Pro

OpenAI

The Pro tier of GPT-5.2, aimed at agentic coding and long-context work beyond what GPT-5 Pro handled. Test-time routing and intent cues like "think hard about this" carry over, with lower hallucination and sycophancy than the previous Pro.

ConversationReasoningCode Generation+1
Released 2025-12-10Extra large model

GPT-5.1 Codex Max

OpenAI

The largest Codex tier in OpenAI's GPT-5.1 line, built for coding agents that run for hours rather than turns. Training data stops in September 2024.

ConversationReasoningCode Generation+1
Released 2025-12-04Extra large model

GPT-5.1

OpenAI

The latest of the GPT-5 series: better general reasoning, tighter instruction adherence and a more natural conversational style than GPT-5. Allocates compute adaptively, answering simple queries fast and spending longer on hard ones.

ConversationReasoningCode Generation+1
Released 2025-11-13Extra large model

GPT-5.1 Chat

OpenAI

GPT-5.1 Chat (Instant), the fast and light member of the 5.1 family, built for low-latency chat. It still thinks selectively on harder queries, which keeps math, coding and multi-step accuracy up without slowing normal conversation.

ConversationReasoningCode Generation+1
Released 2025-11-13large model

GPT-5.1-Codex-Mini

OpenAI

A smaller, faster GPT-5.1-Codex for coding at lower latency.

ConversationReasoningCode Generation+1
Released 2025-11-13Extra large model

GPT-5.1-Codex

OpenAI

GPT-5.1 tuned for software engineering: interactive development sessions and long independent runs, covering projects from scratch, feature work, debugging, large refactors and code review.

ConversationReasoningCode Generation+1
Released 2025-11-13Extra large model

GPT-5 Pro

OpenAI

The Pro tier of GPT-5, given more test-time compute for complex, high-stakes tasks. Same routing controls and intent cues as GPT-5, with lower hallucination and sycophancy and stronger coding, writing and health answers.

ConversationReasoningCode Generation+1
Released 2025-10-06Extra large model

GPT-5 Codex

OpenAI

GPT-5 retuned for Codex, OpenAI's agentic coding tool. It runs on the Responses API only, not chat completions, and its training data stops in September 2024.

ConversationReasoningCode Generation+1
Released 2025-09-23Extra large model

GPT-5

OpenAI

OpenAI's GPT-5, built around step-by-step reasoning and instruction following on high-stakes work. It routes at test time and reads intent cues like "think hard about this". Hallucination and sycophancy are down against GPT-4o, with better coding, writing and health answers.

ConversationReasoningCode Generation+5
Released 2025-08-07Extra large model

GPT-5 Mini

OpenAI

A compact GPT-5 for lighter reasoning, with the same instruction following and safety tuning at lower latency and cost. It is the successor to OpenAI's o4-mini.

ConversationReasoningCode Generation+1
Released 2025-08-07medium model

GPT-5 Nano

OpenAI

The smallest model in OpenAI's GPT-5 line, priced for high-volume classification, extraction and short code edits rather than hard reasoning. Text and image input, training data through May 2024.

ConversationReasoningCode Generation+1
Released 2025-08-07small model

GPT OSS 20B

OpenAI

OpenAI's smaller open-weight release, 21B parameters on a mixture-of-experts design. Matches o3-mini on quality while fitting in 16GB of memory, which puts it on consumer hardware and edge devices.

ConversationReasoningCode Generation+4
Released 2025-08-05large model

GPT OSS 120B

OpenAI

OpenAI's first open-weight model since GPT-2, 117B parameters on a mixture-of-experts design. Reasoning depth is configurable per request and tool use is native, aimed at agent and production workloads.

ConversationReasoningCode Generation+3
Released 2025-08-05Extra large model

GPT Audio

OpenAI

OpenAI's first generally available audio model. Features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Supports both text and audio input/output.

Text To SpeechAudio GenerationConversation
Released 2025-06-17large model

GPT Audio Mini

OpenAI

Cost-efficient version of GPT Audio. Features an upgraded decoder for more natural sounding voices and better voice consistency at a fraction of the cost.

Text To SpeechAudio GenerationConversation
Released 2025-06-17medium model

OpenAI Codex Mini

OpenAI

codex-mini-latest is a fine-tuned version of o4-mini specifically for use in Codex CLI. For direct use in the API, we recommend starting with gpt-4.1.

Code GenerationConversationAnalysis
Released 2025-05-16medium model

OpenAI o3

OpenAI

OpenAI's reasoning model for coding, math, science and visual perception. Thinks longer before answering and reaches for tools on its own: web search, code execution, image generation. Set new SOTA on Codeforces and MMMU.

ConversationReasoningCode Generation+2
Released 2025-04-16Extra large model

OpenAI o4-mini

OpenAI

A smaller, cheaper OpenAI reasoning model built for speed. Strong for its size on math, coding and visual tasks, with much higher usage limits than o3, and it can use tools on its own.

ConversationReasoningCode Generation+1
Released 2025-04-16medium model

OpenAI o4 Mini High

OpenAI

The same model as o4-mini, defaulting to high reasoning effort. Compact, fast and cheap, keeping the multimodal and agentic behaviour, especially on math, coding and visual tasks.

ConversationReasoningCode Generation+1
Released 2025-04-16medium model

GPT-4.1

OpenAI

OpenAI's GPT-4.1, built around a 1 million token context window and precise code diffs. Scores 54.6% on SWE-bench Verified and 87.4% on IFEval, and beats GPT-4o on coding, instruction compliance and long-document recall.

ConversationReasoningCode Generation+1
Released 2025-04-14Extra large model

GPT-4.1 Nano

OpenAI

The smallest and cheapest model in the GPT-4.1 line, with the same 1 million token context. Scores 80.1% on MMLU, 50.3% on GPQA and 9.8% on Aider polyglot, above GPT-4o mini. Aimed at classification and autocompletion.

ConversationReasoningCode Generation+1
Released 2025-04-14small model

GPT-4.1 Mini

OpenAI

The mid tier of the GPT-4.1 line, roughly GPT-4o quality at lower latency and cost, with the same 1 million token context. Scores 84.1% on IFEval, 35.8% on MultiChallenge and 31.6% on Aider's polyglot diff benchmark, and handles vision.

ConversationAnalysisCode Generation
Released 2025-04-14medium model

ChatGPT-4o (March 2025)

OpenAI

The March 2025 refresh of GPT-4o in ChatGPT. Follows instructions more closely, handles coding with fewer detours, and writes shorter answers with fewer nested markdown levels than the previous build.

ConversationReasoningAnalysis+1
Released 2025-03-27Extra large model

GPT-4o Audio

OpenAI

GPT-4o with audio input support. Detects nuances within audio recordings and adds depth to generated user experiences. Processes both text and audio prompts.

Text To SpeechAudio GenerationConversation
Released 2025-03-20large model

GPT-4.5

OpenAI

GPT-4.5 scaled up pre-training and post-training. Broader knowledge, better intent reading and more 'EQ' on conversation, writing, programming and practical problems, with fewer hallucinations. Scored 62.5% on SimpleQA with a 37.1% hallucination rate, ahead of GPT-4o.

ConversationReasoningCode Generation+2
Released 2025-02-27Extra large model

o3 Mini

OpenAI

A smaller, cheaper o3 with faster response times. OpenAI says output quality holds.

ConversationReasoningCode Generation
Released 2024-12-15small model

o1

OpenAI

o1 achieves 86% accuracy on Mathematics Olympiad benchmarks (vs. GPT-4o's 13%), offers PhD-level STEM proficiency, and is built with extensive alignment training and safety evaluation.

ConversationWeb DesignCode Generation+1
Released 2024-12-05medium model

GPT-4o mini

OpenAI

OpenAI's small model after GPT-4 Omni: text and image in, text out. Many multiples cheaper than the frontier models of its day and over 60% cheaper than GPT-3.5 Turbo, at close to SOTA intelligence.

ConversationReasoningCode Generation+1
Released 2024-07-18small model

GPT-4o (Omni)

OpenAI

GPT-4o processes text, images, and audio through a unified transformer architecture and offers real-time translation for 154 languages with 89.2% BLEU score on low-resource languages.

ConversationReasoningAnalysis+2
Released 2024-05-13Extra large model

OpenAI TTS-1 HD

OpenAI

OpenAI's high-definition text-to-speech with six built-in voices. Clear and natural, at the cost of latency.

Text To SpeechAudio Generation
Released 2024-04-09large model

OpenAI TTS-1

OpenAI

OpenAI's standard text-to-speech, tuned for low latency. Six voices, several output formats. Faster than TTS-1-HD and slightly rougher.

Text To SpeechAudio Generation
Released 2023-11-06medium model

DALL-E 3

OpenAI

DALL-E 3 auto-improves user inputs via ChatGPT integration and blocks prohibited content with 99.9% precision using multimodal classifiers.

Image Generation
Released 2023-09-01large model

GPT-4

OpenAI

The original GPT-4: a large multimodal model that handled hard problems more accurately than GPT-3.5 on the strength of broader general knowledge and better reasoning. Training data stops at September 2021.

ConversationReasoningCode Generation+1
Released 2023-05-29Extra large model

GPT-3.5 Turbo

OpenAI

OpenAI's fastest model of its generation. Understands and writes natural language and code, tuned for chat and completion.

ConversationCode Generation
Released 2022-03-15medium model

GPT-2

OpenAI

A direct scale-up of GPT-1 with 1.5 billion parameters, trained on 8 million web pages. Known for its ability to generate coherent text, sometimes indistinguishable from humans, but could be repetitive.

ConversationSummarizationAnalysis
Released 2019-11-05medium model

GPT-1

OpenAI

The first large-scale transformer-based language model released by OpenAI, trained on the BooksCorpus dataset. This version is accessed via the Hugging Face model hub (`openai-community/openai-gpt`).

ConversationSummarization
Released 2018-06-11small model

Google

42

Gemini 3.8 Flash

Google AI

Google's most capable Flash model at release, with gains over 3.7 Flash in software engineering, agent tasks and multi-step reasoning. It accepts text, images, audio, video and files across a 1M-token window, with up to 65,536 output tokens.

ConversationReasoningCode Generation+4
Released 2026-09-02large model

Gemini 3.7 Flash

Google AI

Google's Gemini 3.7 Flash, for fast agent workflows, coding and multi-step reasoning. Text, images, audio, video and files across a 1M-token context.

ConversationReasoningCode Generation+4
Released 2026-08-13large model

Gemini 3.6 Flash

Google AI

Google's Flash workhorse at the 3.6 mark, for coding, agent loops and web and app development. Built on 3.5 Flash, aiming at cleaner output with fewer unnecessary edits and less hedging, on roughly 17% fewer output tokens. Text, images, audio, video and PDFs across 1M tokens, up to 64K out.

ConversationReasoningCode Generation+4
Released 2026-07-21large model

Nano Banana 2 Lite

Google AI

Gemini 3.1 Flash Lite Image, the cheapest and fastest of Google's Nano Banana line. Text-to-image in roughly 4 seconds, about 2.7x faster than Gemini 3.1 Flash Image, with the same character consistency. One API for text-to-image, editing and composition. 1K resolution, 14 aspect ratios, SynthID watermark.

Image Generation
Released 2026-06-30small model

Gemini 3.5 Flash

Google AI

Google's Flash tier at the 3.5 mark, close to Pro on coding and reasoning at Flash cost and speed. Tuned for coding and parallel agent loops, taking text, image, video, audio and PDF input. Thinking effort defaults to medium and runs from minimal to high.

ConversationReasoningCode Generation+3
Released 2026-05-19large model

Gemini 3.1 Flash Lite

Google AI

The GA release of Gemini 3.1 Flash Lite, Google's cheapest multimodal tier. Text, image, video, audio and PDF in, for light agent work, simple extraction, and anything where latency and cost decide the design. Thinking runs minimal to high. Half the price of Gemini 3 Flash.

ConversationReasoningAnalysis+4
Released 2026-05-07small model

Gemma 4 26B A4B

Google AI

An instruction-tuned mixture-of-experts from Google DeepMind: 25.2B total parameters, 3.8B active per token, which is how it reaches near-31B quality on a fraction of the compute. Text, images and video up to 60 seconds at 1fps. 256K context, native function calling, thinking mode. Apache 2.0.

ConversationReasoningCode Generation+2
Released 2026-04-03large model

Gemma 4 31B

Google AI

Google DeepMind's 30.7B dense multimodal model: text and image in, text out. 256K context, configurable thinking mode, native function calling, 140+ languages. Strongest on coding, reasoning and document understanding. Apache 2.0.

ConversationReasoningCode Generation+2
Released 2026-04-02Extra large model

Gemini 3.1 Flash Lite Preview

Google AI

Google's high-efficiency tier at the 3.1 mark. Beats Gemini 2.5 Flash Lite on quality and gets close to 2.5 Flash, with the gains in audio and speech recognition, RAG snippet ranking, translation, extraction and code completion. 2.5x faster to first token than 2.5 Flash, at half the price of Gemini 3 Flash.

ConversationReasoningAnalysis+4
Released 2026-03-03small model

Nano Banana 2

Google AI

Gemini 3.1 Flash Image Preview, sold as Nano Banana 2. Google's image generation and editing model at Flash speed and Flash pricing, aimed at iterative edits rather than one-shot renders.

Image Generation
Released 2026-02-26large model

Gemini 3.1 Pro Preview

Google AI

The 3.1 update to Gemini Pro: better software engineering scores, more reliable agent behaviour, fewer tokens spent getting there. Reasons across text, image, video, audio and code over a 1M token context. Adds a medium thinking level, and improves on structured work like finance and spreadsheets.

ConversationReasoningCode Generation+3
Released 2026-02-19Extra large model

Lyria 3

Google AI

Google DeepMind's Lyria 3, launched in the Gemini app. First Lyria with full vocal generation, automatic lyrics and multimodal input (text, image, video). Produces 30-second tracks with distinct verses, choruses and bridges, in 8 languages. No public API: it runs in the Gemini app and YouTube Dream Track.

Music GenerationAudio Generation
Released 2026-02-18Extra large model

Gemini 3 Flash Preview

Google AI

Gemini 3's Flash tier, for agent loops, multi-turn chat and coding assistance at much lower latency than Pro. 1M token context. Text, image, audio, video and PDF in, text out. Thinking level runs minimal to high, with structured output, tool use and automatic context caching.

ConversationReasoningCode Generation+3
Released 2025-12-17large model

Lyria 2

Google AI

Google DeepMind's production music generation model available via Vertex AI. Generates 32.8-second instrumental clips at 48kHz from text prompts. Supports negative prompts to exclude unwanted elements and deterministic generation via seed parameter. All output is SynthID-watermarked.

Music GenerationAudio Generation
Released 2025-12-10large model

Gemini 3 Pro Preview

Google AI

Gemini 3 Pro Preview with high reasoning effort enabled. Exposes full chain-of-thought process for enhanced transparency in complex problem-solving across text, code, and multimodal tasks.

ConversationReasoningCode Generation+2
Released 2025-11-18Extra large model

Nano Banana Pro

Google AI

Google's Gemini 3 Pro Image model, sold as Nano Banana Pro. Generates and edits at up to 4K and routes prompts through Gemini's thinking mode for text rendering and multi-step instructions.

Image Generation
Released 2025-11-12large model

Gemini 2.5 Flash Preview 09-2025

Google AI

The September 2025 checkpoint of Gemini 2.5 Flash, Google's workhorse tier for reasoning, coding, mathematics and science. Thinking is built in, and the max tokens for reasoning parameter caps how long it deliberates.

ConversationReasoningCode Generation+1
Released 2025-09-25Extra large model

Gemini 2.5 Flash Lite Preview 09-2025

Google AI

The September 2025 checkpoint of Gemini 2.5 Flash-Lite, the cheapest and lowest latency tier in the family. Thinking is off by default; turning it on through the Reasoning API parameter trades cost for intelligence.

ConversationReasoningCode Generation+1
Released 2025-09-25large model

Nano Banana

Google AI

Google's Nano Banana (Gemini 2.5 Flash Image) generates and edits images, tuned for speed. Tested on LMArena as 'nano-banana', where it beat Midjourney and Flux in blind tests.

Image Generation
Released 2025-08-26large model

Gemini 2.5 Flash Image

Google AI

Gemini 2.5 Flash variant tuned for fast text-to-image generation and simple image edits.

Image Generation
Released 2025-08-26large model

Gemma 3n 2B

Google AI

Google DeepMind's Gemma 3n at an effective 2B parameters inside a 6B architecture. The MatFormer design lets it nest submodels and recombine them through Mix-and-Match, which is how it runs on low-resource hardware. 32K context, multilingual.

ConversationReasoningTranslation
Released 2025-07-09small model

Gemini 2.5 Flash Lite Preview 06-17

Google AI

The lightweight tier of Gemini 2.5, tuned for low latency and cost rather than depth. Thinking is off by default to keep it fast; developers can switch it on through the Reasoning API parameter and pay for the extra intelligence.

ConversationReasoningAnalysis+1
Released 2025-06-17small model

Imagen 4 Fast

Google AI

Imagen 4 Fast trades some quality for speed and cost.

Image Generation
Released 2025-06-12large model

Gemini 2.5 Pro Preview 06-05

Google AI

Google's Gemini 2.5 Pro, a reasoning model that thinks through a request before answering. Aimed at coding, mathematics and scientific work, and it held first place on the LMArena leaderboard at release.

ConversationReasoningCode Generation+2
Released 2025-06-05Extra large model

Imagen 4 Ultra

Google AI

Imagen 4 Ultra trades speed and cost for image quality.

Image Generation
Released 2025-05-21large model

Imagen 4

Google AI

Google's Imagen 4 text-to-image model, with strong detail and prompt adherence.

Image Generation
Released 2025-05-21large model

Gemini 2.5 Flash Preview 05-20

Google AI

The May 20 checkpoint of Gemini 2.5 Flash in its non-thinking form: it will not emit reasoning tokens, and is billed at the lower output rate. The paired :thinking variant is the one that reasons before answering.

ConversationReasoningCode Generation+1
Released 2025-05-20medium model

Gemini 2.5 Flash Preview 05-20 (thinking)

Google AI

The May 20 checkpoint of Gemini 2.5 Flash with thinking enabled. It reasons before answering, is billed at the higher output rate, and takes a max tokens for reasoning parameter to cap how long it deliberates.

ConversationReasoningCode Generation+1
Released 2025-05-20medium model

Gemma 3n 4B

Google AI

Google's Gemma 3n at the E4B size, built to run on phones, laptops and tablets. Text, image and audio in, for transcription, translation and image analysis. Per-Layer Embedding caching and the MatFormer architecture activate only part of the model per request, which keeps memory use low. 32K context, 140+ languages.

ConversationAnalysisTranslation+1
Released 2025-05-20small model

Gemini 2.5 Pro (I/O Edition)

Google AI

Google's Gemini 2.5 Pro preview: multimodal understanding, multi-step code, math and STEM, and long-context analysis over large datasets, codebases and documents. Knowledge cutoff January 2025.

ConversationReasoningCode Generation+1
Released 2025-05-06Extra large model

Gemini 2.5 Flash Preview

Google AI

Google's Flash tier workhorse, with reasoning that can be switched on or off and a configurable thinking budget. The point is trading quality against cost and latency per request rather than per model.

ConversationReasoningCode Generation+1
Released 2025-04-17medium model

Gemini 2.5 Flash Preview (thinking)

Google AI

Gemini 2.5 Flash with thinking left on, so it reasons through a request before answering. The same workhorse model as the non-thinking variant, priced higher on output because thinking tokens are billed.

ConversationReasoningCode Generation+1
Released 2025-04-17medium model

Gemini 2.5 Pro Experimental

Google AI

Google's experimental Gemini 2.5 Pro: multimodal reasoning, a 1 million token context window, and video comprehension.

ConversationReasoningCode Generation+1
Released 2025-03-25Extra large model

Gemma 3 12B

Google AI

Gemma 3 adds multimodality: vision-language in, text out. 96,000 token context, 140+ languages, better math, reasoning and chat, with structured outputs and function calling. The 12B is the second largest, after the 27B.

ConversationReasoningCode Generation+1
Released 2025-03-13large model

Gemma 3 27B

Google AI

Gemma 3 adds multimodality: vision-language in, text out. 131,072 token context, 140+ languages, better math, reasoning and chat, with structured outputs and function calling. The 27B is Google's largest Gemma 3, successor to Gemma 2.

ConversationReasoningCode Generation+1
Released 2025-03-12Extra large model

Imagen 3 Fast

Google AI

A faster and cheaper Imagen 3 variant for when speed and cost matter more than maximum quality.

Image Generation
Released 2025-02-06large model

Imagen 3

Google AI

Google's high-quality text-to-image model focused on lighting, detail, and strong visual composition.

Image Generation
Released 2025-02-05large model

Gemini 2.0 Pro Experimental

Google AI

Gemini 2.0 Pro builds interactive 3D environments from text descriptions and offers hypothetical reasoning for scientific simulations.

ConversationReasoningAnalysis+2
Released 2025-01-01Extra large model

Gemini 2.0 Flash Thinking

Google AI

An experimental Gemini 2.0 Flash variant that reasons through a problem before answering. Trades some latency for accuracy on multi-step questions while staying at Flash speed.

ConversationReasoningAnalysis+1
Released 2024-12-11large model

Gemini 1.5 Pro

Google AI

Gemini 1.5 Pro handles infinite context with 99% retrieval accuracy at 750k tokens via Mixture-of-Experts and generates chapter summaries for 2-hour videos with 92% accuracy.

ConversationReasoningAnalysis+1
Released 2024-02-15large model

Gemini Pro 1.0

Google AI

Google's multimodal Gemini at release: natural language, multi-turn chat, code generation and image input.

ConversationReasoningCode Generation
Released 2023-12-13large model

PaLM 2 Chat

Google AI

Google's PaLM 2, with better multilingual, reasoning and coding work than PaLM. Tuned for chat.

ConversationReasoningCode Generation
Released 2023-07-20large model

Anthropic

22

Claude Fable 5.1

Anthropic

Claude Fable 5.1 is a direct update to Fable 5, with its largest gains in agent coding, long-running workflows and knowledge work. Anthropic calls out long refactors, front-end generation, finance and analysis, with shorter plans and summaries than Fable 5.

ConversationReasoningCode Generation+5
Released 2026-09-01Extra large model

Claude Opus 5

Anthropic

Anthropic's largest model, for demanding reasoning, coding and long-horizon agent work. Strongest on end-to-end software tasks, code review and bug finding, visual analysis of charts and documents, and coordinating parallel subagents. Instruction following holds across long tasks, even at lower effort.

ConversationReasoningCode Generation+3
Released 2026-07-24Extra large model

Claude Sonnet 5

Anthropic

Anthropic's most capable Sonnet, with adaptive thinking across selectable effort levels from low to x-high, a 1M token context, and text, image and file input. It ships an updated tokenizer and real-time cyber safeguards that block certain high-risk dual-use requests.

ConversationReasoningCode Generation+3
Released 2026-06-30large model

Claude Fable 5

Anthropic

Anthropic's first public Mythos-class model, for autonomous knowledge work and coding. Text, image and file input, 1M token context. Built for long, ambiguous, multi-step tasks that used to need check-ins, self-correcting through verification loops. Falls back to Claude Opus 4.8 on high-risk queries.

ConversationReasoningCode Generation+4
Released 2026-06-09Extra large model

Claude Opus 4.8

Anthropic

Anthropic's most capable generally available Opus. Text, image and file input, 1M token context. Built for autonomous agents and memory-driven work where coherence across a long session is the hard part: multi-step reasoning, complex coding, orchestration across large codebases, multi-stage debugging.

ConversationReasoningCode Generation+2
Released 2026-05-28Extra large model

Claude Opus 4.7

Anthropic

Anthropic's Opus at the 4.7 mark, for long-running asynchronous agents. Stronger than Opus 4.6 on multi-step work: large codebases, multi-stage debugging, end-to-end orchestration. Beyond coding it drafts documents, builds presentations and analyses data, holding coherence across very long sessions.

ConversationReasoningCode Generation+2
Released 2026-04-15Extra large model

Claude Sonnet 4.6

Anthropic

Anthropic's Sonnet 4.6, for coding, agents and professional work: iterative development, navigating large codebases, end-to-end project management with memory, document creation, and computer use for web QA and workflow automation.

ConversationReasoningCode Generation+3
Released 2026-02-17large model

Claude Opus 4.6

Anthropic

Anthropic's strongest model for coding and long-running professional work, built for agents that operate across a whole workflow rather than a single prompt. Handles large codebases, complex refactors and multi-step debugging, and produces near-final documents, plans and analyses in one pass.

ConversationReasoningCode Generation+2
Released 2026-02-04Extra large model

Claude Opus 4.5

Anthropic

Anthropic's Opus tier at the 4.5 mark, aimed at long coding sessions and agent work. Text and image input, training data through August 2025, available on the Claude API, Bedrock and Vertex.

ConversationReasoningCode Generation+2
Released 2025-11-24Extra large model

Claude Haiku 4.5

Anthropic

Claude Haiku 4.5 matches Claude Sonnet 4's reasoning and coding strength while staying ultra-responsive for real-time and high-volume workloads. Extended thinking, controllable reasoning depth, and full tool-use support keep sub-agents fast without sacrificing quality.

ConversationReasoningCode Generation+2
Released 2025-10-15large model

Claude Sonnet 4.5

Anthropic

Anthropic's Sonnet 4.5, tuned for agents and coding workflows: speculative parallel execution, memory handling and tool orchestration.

ConversationReasoningCode Generation+2
Released 2025-09-29large model

Claude Opus 4.1

Anthropic

An update to Anthropic's Opus tier, sharpened on coding, reasoning and agent work. Scores 74.5% on SWE-bench Verified, with the clearest gains in multi-file refactors and debugging precision. Extended thinking runs to 64K tokens.

ConversationReasoningCode Generation+3
Released 2025-08-05Extra large model

Claude Opus 4

Anthropic

The top tier of Anthropic's Claude 4 line, for coding, reasoning and agents. Built for long-running tasks, with extended thinking during tool use and better memory.

ConversationReasoningCode Generation+3
Released 2025-05-22Extra large model

Claude Sonnet 4

Anthropic

Claude Sonnet 4 is a significant upgrade to Sonnet 3.7, delivering superior coding and reasoning while responding more precisely to instructions. It balances performance and efficiency for various use cases.

ConversationReasoningCode Generation+3
Released 2025-05-22large model

Claude 3.7 Thinking Sonnet

Anthropic

Claude 3.7 Thinking Sonnet exposes the full chain-of-thought process during problem-solving, including error backtracking and alternative solution exploration. Scores 86.1% on GPQA Diamond benchmark for expert-level Q&A.

ConversationReasoningAnalysis+1
Released 2025-02-26large model

Claude 3.7 Sonnet

Anthropic

Claude 3.7 Sonnet offers Extended Thinking Scaffolds that boost SWE-bench coding accuracy from 62.3% to 70.3%, with 81.2% accuracy in retail automation tasks, outperforming Claude Sonnet 3.6 (2022-10-22) by 13.6%.

ConversationReasoningAnalysis+1
Released 2025-02-25large model

Claude Sonnet 3.6 (2022-10-22)

Anthropic

Claude 3.5 Sonnet offers a cost-efficient API ($3/million input tokens vs. $5 for GPT-4o) and uses embedded alignment techniques that reduce harmful outputs by 34% compared to Claude 2.1.

ConversationReasoningAnalysis+1
Released 2024-06-01large model

Golden Gate Claude

Anthropic

A research demo of Claude 3 Sonnet, live for 24 hours on 23 May 2024, engineered by Anthropic to demonstrate feature steering. It was manipulated to obsess over the Golden Gate Bridge in every response.

ConversationReasoning
Released 2024-05-23large model

Claude 3 Sonnet

Anthropic

The middle tier of Claude 3: a 200,000 token context window, priced for scaled deployment, and steady on conversation, reasoning, code generation and analysis.

ConversationReasoningCode Generation+1
Released 2024-03-05large model

Claude 3 Haiku

Anthropic

Claude 3 Haiku is Anthropic's fastest and most compact model, designed for near-instant responses in real-time applications with strong multilingual support.

ConversationAnalysisSummarization
Released 2024-03-04medium model

Claude 3 Opus

Anthropic

The top tier of Anthropic's Claude 3 line, for complex reasoning and open-ended problem solving.

ConversationReasoningCode Generation+1
Released 2024-03-04Extra large model

Claude 2

Anthropic

Anthropic's Claude 2 model, featuring a large 100K token context window and strong performance on various benchmarks. Known for helpful, honest, and harmless AI conversations.

ConversationReasoningAnalysis+1
Released 2023-06-11large model
Sponsored

Amazon

+2

Amazon Nova 2 Lite

Amazon

Nova 2 Lite is a fast, cheap reasoning model for everyday workloads. Text, images and video in, text out. Aimed at document processing, extracting information from video, code generation, grounded answers and multi-step agent workflows.

ConversationReasoningCode Generation+3
Released 2025-12-02medium model

Nova Premier 1.0

Amazon

The top tier of Amazon's Nova line, for complex reasoning and as the teacher model when distilling your own.

ConversationReasoningCode Generation+1
Released 2025-10-31large model

arcee-ai

+1

Trinity Large Preview

Arcee AI

Arcee's Trinity Large Preview, a 400B sparse mixture-of-experts with 13B active per token, routing 4 of 256 experts. Open weights. Tuned for creative writing, role-play, chat and real-time voice, and trained to work inside scaffolds like OpenCode, Cline and Kilo Code.

ConversationReasoningCode Generation+1
Released 2025-01-27Extra large model

baidu

+2

ERNIE 4.5 21B A3B Thinking

Baidu

Baidu's lightweight reasoning MoE model (21B total, 3B active per token) refined for deep multi-step reasoning across logic, math, science, and code. 128K context, Apache-2.0.

ConversationReasoningAnalysis+1
Released 2025-10-09medium model

ERNIE 4.5 300B A47B

Baidu

Baidu's 300B-parameter text Mixture-of-Experts model, 47B active per token. 123K context, Apache-2.0. Part of the ERNIE 4.5 series.

ConversationReasoningAnalysis+1
Released 2025-06-30Extra large model

Black Forest

+8

FLUX 2 Max

Black Forest Labs

The highest fidelity image model from Black Forest Labs for maximum detail and realism.

Image Generation
Released 2025-12-16large model

FLUX 2 Flex

Black Forest Labs

FLUX 2 Flex adds guidance and step controls for more steerable generations.

Image Generation
Released 2025-11-25large model

FLUX 2 Dev

Black Forest Labs

Black Forest Labs' FLUX 2 Dev model for high-quality generations with flexible controls.

Image Generation
Released 2025-11-24large model

FLUX 2 Pro

Black Forest Labs

FLUX 2 Pro focuses on premium quality output with strong prompt adherence.

Image Generation
Released 2025-11-14large model

FLUX Pro

Black Forest Labs

Black Forest Labs' top FLUX tier for text-to-image, aimed at prompt adherence and image detail rather than speed. The reference model the rest of the FLUX line is measured against.

Image Generation
Released 2024-11-12large model

FLUX Kontext Max

Black Forest Labs

The premium tier of FLUX Kontext for text-driven image editing, with better typography than the standard tier. You describe the change in words.

Image Generation
Released 2024-11-12large model

FLUX Kontext Pro

Black Forest Labs

A FLUX model for editing images by describing the change in words. Built for transformations that keep the rest of the frame consistent instead of regenerating the whole picture.

Image Generation
Released 2024-11-12large model

FLUX Schnell

Black Forest Labs

A fast and cost-efficient FLUX model designed for high-throughput text-to-image generation.

Image Generation
Released 2024-07-30large model

Bria

+2

Bria Image 3.2

Bria

A text-to-image model trained entirely on licensed data, so it is commercial-ready. 4B parameters, with aesthetics and text rendering evaluated on par with much larger models.

Image Generation
Released 2025-12-04large model

Fibo

Bria

SOTA open source model trained on licensed data, transforming intent into structured control for precise, high-quality AI image generation in enterprise and agentic workflows.

Image Generation
Released 2025-12-04large model

ByteDance

+6

Seed 2.1 Turbo

ByteDance

Seed 2.1 Turbo is a multimodal model from ByteDance Seed for coding and long-horizon agent workflows. It is suited for end-to-end software delivery, multi-step task execution, and understanding visual and video content, with capabilities for planning, debugging, and self-correction.

ConversationReasoningCode Generation+3
Released 2026-08-12large model

Seed 2.0 Code

ByteDance

ByteDance Seed's coding-agent model. Suited to frontend work, multilingual programming, and coding agents such as Claude Code, Kilo and OpenCode.

Code GenerationConversationReasoning+2
Released 2026-08-12large model

Seed 2.0 Lite

ByteDance

ByteDance's mid tier Seed model for high-frequency enterprise work, ahead of Seed 1.8 overall. Reasoning effort is configurable, and it handles multimodal input and structured output. The target jobs are unstructured information processing, text creation, search, recommendation and data analysis.

ConversationReasoningCode Generation+2
Released 2026-03-10medium model

Seedream 5.0 Lite

ByteDance

ByteDance's Seedream 5.0 Lite is a unified multimodal image generation model with chain-of-thought reasoning and real-time web search, producing native 2K/4K images with accurate multi-language text rendering.

Image Generation
Released 2026-02-13medium model

Seedream 4.5

ByteDance

ByteDance's Seedream 4.5 text-to-image model designed for strong aesthetics and composition.

Image Generation
Released 2025-12-03large model

Seedream 4

ByteDance

Unified text-to-image generation and precise single-sentence editing at up to 4K resolution by ByteDance.

Image Generation
Released 2025-11-15large model

cognitive-computations

+1

Dolphin Mistral 24B

Cognitive Computations

Dolphin Mistral 24B is an explicitly uncensored fine-tune of Mistral Small 24B by Cognitive Computations and Eric Hartford. Designed for unrestricted research use, it removes alignment-based content filtering while retaining strong instruction-following capabilities.

ConversationReasoningCode Generation
Released 2025-02-01medium model

Cohere

+1

North Mini Code

Cohere

Cohere's first agentic coding model and the debut of its North family. A sparse mixture of experts, 30B total and 3B active, for code generation, agentic software engineering and terminal work, trained to generalise across scaffolds like OpenCode and SWE-Agent. 256K context, 64K output. Apache 2.0.

ConversationReasoningCode Generation+1
Released 2026-06-17small model

DeepSeek

+12

DeepSeek V4 Pro 0813

DeepSeek

DeepSeek's large mixture-of-experts model and the general-availability release of V4 Pro. Keeps the 1M-token context window and hybrid attention of the V4 line, aimed at reasoning, coding and long-horizon agent work.

ConversationReasoningCode Generation+2
Released 2026-08-12Extra large model

DeepSeek V4 Flash 0731

DeepSeek

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision of DeepSeek V4 Flash is suited for coding, reasoning, and agent workflows.

ConversationReasoningCode Generation+2
Released 2026-07-31large model

DeepSeek V4 Flash

DeepSeek

DeepSeek V4 Flash, a mixture-of-experts model with 284B total and 13B activated parameters over a 1M token context. Hybrid attention handles the long context, reasoning modes are configurable, and the whole thing is tuned for throughput: coding assistants, chat and agent loops.

ConversationReasoningCode Generation+2
Released 2026-04-24large model

DeepSeek V4 Pro

DeepSeek

DeepSeek V4 Pro, a mixture-of-experts model with 1.6T total and 49B activated parameters over a 1M token context. The same hybrid attention architecture as V4 Flash at much larger scale, with multiple reasoning modes and strong knowledge, mathematics and software engineering scores.

ConversationReasoningCode Generation+2
Released 2026-04-24Extra large model

DeepSeek V3.2 Speciale

DeepSeek

The high-compute variant of DeepSeek V3.2, with reinforcement learning post-training scaled past the base model. It keeps DeepSeek Sparse Attention for long context, and DeepSeek reports it ahead of GPT-5 on hard reasoning workloads and comparable to Gemini 3.0 Pro, with coding and tool use intact.

ConversationReasoningCode Generation+2
Released 2025-12-01large model

DeepSeek V3.2

DeepSeek

DeepSeek V3.2 introduces DeepSeek Sparse Attention, cutting training and inference cost without giving up long-context quality. Reinforcement learning post-training puts its reported reasoning in the GPT-5 class, with gold-medal results on the 2025 IMO and IOI, plus better tool-use compliance.

ConversationReasoningCode Generation+2
Released 2025-12-01large model

DeepSeek V3.2 Exp

DeepSeek

DeepSeek-V3.2-Exp introduces DeepSeek Sparse Attention (DSA) for efficient long-context. Reasoning toggle supported via boolean flag.

ConversationReasoningCode Generation+1
Released 2025-09-29large model

DeepSeek V3.1

DeepSeek

DeepSeek V3.1 runs one set of weights in two modes: deepseek-chat answers directly, deepseek-reasoner thinks first. Open weights on Hugging Face, 160K context, tuned for agentic tool calling.

ConversationReasoningCode Generation+4
Released 2025-08-21Extra large model

DeepSeek R1 0528

DeepSeek

DeepSeek R1 0528 is the May 28th update to the original DeepSeek R1. Performance on par with OpenAI o1, but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active in an inference pass. Fully open-source model.

ConversationReasoningCode Generation+1
Released 2025-05-28Extra large model

DeepSeek Prover V2

DeepSeek

A 671B parameter model, speculated to be geared towards logic and mathematics. Likely an upgrade from DeepSeek-Prover-V1.5. Released on Hugging Face without an announcement or description.

ReasoningAnalysisConversation+1
Released 2025-04-30Extra large model

DeepSeek R1

DeepSeek

DeepSeek R1 is a reasoning model developed entirely via reinforcement learning, offering cost efficiency at $0.14/million tokens vs. OpenAI o1's $15, with strong code generation and analysis capabilities.

ConversationReasoningCode Generation+1
Released 2025-02-01large model

DeepSeek V3 (March 2024)

DeepSeek

DeepSeek V3 (March 2024) shows significant improvements in reasoning capabilities with enhanced MMLU-Pro (81.2%), GPQA (68.4%), AIME (59.4%), and LiveCodeBench (49.2%) scores. Features improved front-end web development, Chinese writing proficiency, and function calling accuracy.

ConversationReasoningWeb Design+2
Released 2024-03-24Extra large model

dots-studio

+1

Dots3-Note Preview

Dots Studio

Dots3-Note Preview is an open-weight mixture-of-experts model from Dots Studio, with 16B active parameters out of 280B total. It is the lightest model in the Dots 3 family and is suited for reasoning, coding, multimodal understanding, long-context processing, and multi-step agent workflows.

ConversationReasoningCode Generation+2
Released 2026-08-14large model

ElevenLabs

+3

ElevenLabs Eleven v3

ElevenLabs

Expressive TTS with emotion, laughter and breathing. ElevenLabs' top voice-quality tier.

Text To SpeechAudio Generation
Released 2025-10-15large model

ElevenLabs Flash v2.5

ElevenLabs

TTS at roughly 75ms latency, for real-time apps and voice agents.

Text To SpeechAudio Generation
Released 2025-06-20medium model

ElevenLabs Multilingual v2

ElevenLabs

High-quality multilingual TTS supporting 29 languages with voice cloning capabilities.

Text To SpeechAudio GenerationTranslation
Released 2024-03-15large model

Fish Audio

+1

Fish Speech v1.5

Fish Audio

Open-source multilingual TTS with 200K+ community voices and ultra-fast voice cloning from 10-30 second samples.

Text To SpeechAudio Generation
Released 2025-01-15medium model

ibm-granite

+1

Granite 4.2 8B

IBM Granite

IBM's dense 8B reasoning model for math, code and multilingual dialogue. It supports full, low-effort and nonthinking modes, a 131K-token window and tool use. The language list spans English, Spanish, Arabic, Japanese, Chinese and seven more.

ConversationReasoningCode Generation+4
Released 2026-08-31small model

Ideogram

+3

Ideogram v3 Turbo

Ideogram

Turbo is the fastest and cheapest Ideogram v3. v3 creates images with stunning realism, creative designs, and consistent styles.

Image Generation
Released 2024-11-12large model

Ideogram v3 Quality

Ideogram

The slowest and most expensive Ideogram v3 variant, and the most detailed. Pick it when quality outranks speed and cost.

Image Generation
Released 2024-11-12large model

Ideogram v3 Balanced

Ideogram

The middle Ideogram v3 variant: text rendering and prompt comprehension at a moderate generation time.

Image Generation
Released 2024-11-12large model

Inception

+3

Mercury 2.5 Preview

Inception

Inception's diffusion language model generates and refines tokens in parallel instead of one at a time. It reaches a reported 1,107 tokens per second on standard GPUs, with a 260K-token context window and a claimed 10-point intelligence gain over Mercury 2.

ConversationReasoningCode Generation+3
Released 2026-08-31medium model

Mercury 2

Inception

Inception's Mercury 2, the first reasoning diffusion language model. Instead of emitting tokens one at a time it produces and refines many in parallel, clearing 1000 tokens per second on standard GPUs, over 5x faster than Claude 4.5 Haiku and GPT-5 Mini. 128K context, native tool use, schema-aligned JSON.

ConversationReasoningCode Generation+2
Released 2026-03-04medium model

Mercury

Inception

The first diffusion large language model. Discrete diffusion runs 5 to 10x faster than speed-tuned models like GPT-4.1 Nano and Claude 3.5 Haiku while matching them, which is what makes voice agents and search viable.

ConversationReasoningCode Generation+1
Released 2025-06-26medium model

inclusionai

+3

Ring 2.6 1T

inclusionAI

inclusionAI's Ring-2.6-1T, a trillion-parameter thinking model with 63B active, for agent workflows that must be capable and cheap. Tuned for coding agents, tool use and long-horizon execution. inclusionAI reports leading results on PinchBench, ClawEval, TAU2-Bench and GAIA2-search.

ConversationReasoningCode Generation+2
Released 2026-05-08Extra large model

Ling 2.6 1T

inclusionAI

inclusionAI's trillion-parameter Ling, using a fast-thinking approach that cuts cost to roughly a quarter of comparable models. inclusionAI reports leading results on AIME26 and SWE-bench Verified, aimed at coding, complex reasoning and large-scale agent workflows.

ConversationReasoningCode Generation+2
Released 2026-04-23Extra large model

Ling 2.6 Flash

inclusionAI

inclusionAI's Ling-2.6-flash, 104B total parameters with 7.4B active, built for agents that need fast answers and few tokens. inclusionAI reports it level with comparable models at similar scale while spending noticeably less on coding, document processing and light agent work.

ConversationReasoningCode Generation+1
Released 2026-04-21large model

Kokoro

+1

Kokoro 82M

Kokoro

Lightweight 82M parameter TTS model based on StyleTTS2. The most popular text-to-speech model on Replicate with over 78 million runs. Known for natural-sounding speech with low latency.

Text To Speech
Released 2024-12-15small model

leonardoai

+1

Leonardo Lucid Origin

Leonardo AI

Leonardo AI's Lucid Origin, with better prompt adherence, more variety and sharper definition than its predecessors. Strongest on stylised and creative imagery.

Image Generation
Released 2025-06-01large model

liquid

+1

LFM2.5-2.6B

Liquid AI

LFM2.5-2.6B is a compact reasoning model from Liquid AI. It is suited for agent workflows, data extraction, RAG, and long-context processing. Liquid advises against using it for agentic coding or knowledge-heavy tasks.

ConversationReasoningData Extraction+1
Released 2026-08-11small model

luma

+2

Luma Photon

Luma

Luma's image model for creative professional work. Photorealistic output with careful lighting, material rendering and composition.

Image Generation
Released 2025-03-01large model

Luma Photon Flash

Luma

Accelerated variant of Luma Photon prioritizing speed while maintaining quality. Ideal for rapid iteration and production workflows where faster generation times are critical.

Image Generation
Released 2025-03-01large model

Meta

+9

Muse Spark 1.3 Contributor

Meta AI

Meta's lower-cost Muse Spark 1.3 tier for experiments, learning, coding and early agent workflows. It follows information across long tasks and asks for confirmation when inputs conflict. Access is restricted to adults; prompts and outputs may be used to improve Meta's products.

ConversationReasoningCode Generation+5
Released 2026-09-02large model

Muse Spark 1.3

Meta AI

Meta's multimodal reasoning model for long-running agent and coding workflows. It tracks state across extended tasks, works through conflicting inputs and asks for confirmation when needed, with a 1M-token context window and concise execution. Access is restricted to adults.

ConversationReasoningCode Generation+5
Released 2026-09-02large model

Muse Spark 1.1

Meta AI

Meta's multimodal reasoning model for agent work. Text, images, video, audio and PDFs in, text out, over a 1M token context. Orchestrates multi-agent workflows as planner or subagent, and generalises zero-shot to new tools, MCP servers and custom skills. Parallel function calling, search with citations.

ConversationReasoningCode Generation+4
Released 2026-07-09large model

Llama 4 Maverick

Meta AI

Meta's multimodal MoE: 17B active parameters across 128 experts, 400B total. Beat GPT-4o and Gemini 2.0 Flash on several benchmarks, with an ELO of 1417 on LMArena. Built for image understanding and creative writing.

ConversationReasoningCode Generation+2
Released 2025-04-05Extra large model

Llama 4 Scout

Meta AI

Meta's compact multimodal MoE: 17B active parameters across 16 experts, 109B total. Fits on one H100 at Int4 quantization, carries a 10M token context window, and beat Gemma 3, Gemini 2.0 Flash-Lite and Mistral 3.1 on several benchmarks.

ConversationReasoningCode Generation+1
Released 2025-04-05medium model

Llama 3.1 70B (Instruct)

Meta AI

Llama 3.1 70B: a much larger context window than 3.0, and better math and general knowledge.

ConversationReasoningCode Generation
Released 2024-07-23large model

Llama 3.1 405B

Meta AI

Meta's largest open-weight Llama 3.1, which beat several proprietary models on benchmarks at release.

ConversationReasoningCode Generation+1
Released 2024-07-23Extra large model

Llama 3 70B

Meta AI

Llama 3 70B is a large language model from Meta with strong performance and efficiency for real-time interactions.

ConversationReasoningCode Generation
Released 2024-04-18large model

MusicGen Large

Meta AI

Meta's text-to-music generation model producing high-quality, diverse music from text descriptions. Trained on licensed music.

Music GenerationAudio Generation
Released 2023-06-12large model

Midjourney

+7

Midjourney v6

Midjourney

Midjourney v6 produces realistic images.

Image Generation
Released 2024-12-01large model

Midjourney v6.1

Midjourney

Midjourney v6.1 introduced a native web interface alongside Discord, with improved detail rendering, better text handling, and enhanced image coherence.

Image Generation
Released 2024-07-31large model

Midjourney v5

Midjourney

Midjourney v5 produces realistic images.

Image Generation
Released 2023-11-01large model

Midjourney v4

Midjourney

Midjourney v4 was the jump in photorealism, coherence and prompt understanding. First version trained on Google TPUs.

Image Generation
Released 2022-11-05large model

Midjourney v3

Midjourney

Midjourney v3 read prompt nuance and artistic style better than v2.

Image Generation
Released 2022-07-25large model

Midjourney v2

Midjourney

Midjourney v2 improved on the original model with better coherence, detail, and more consistent style application.

Image Generation
Released 2022-04-12medium model

Midjourney v1

Midjourney

The first public release of Midjourney, introducing AI image generation to a wider audience through its Discord-based interface.

Image Generation
Released 2022-02-01medium model

MiniMax

+9

MiniMax M3

MiniMax

MiniMax M3 takes text, image and video, returns text, over a 1M token context, for long-horizon agent work, coding and tool use. MiniMax Sparse Attention swaps full attention for KV-block selection, cutting per-token compute at long context to roughly a twentieth of the last generation.

ConversationReasoningCode Generation+4
Released 2026-06-01large model

MiniMax M2.7

MiniMax

MiniMax M2.7 runs multi-agent collaboration internally, planning, executing and refining across changing environments. Handles live debugging, root cause analysis, financial modelling and document generation across Word, Excel and PowerPoint. Scores 56.2% on SWE-Pro, 57.0% on Terminal Bench 2, 1495 ELO on GDPval-AA.

ConversationReasoningCode Generation+4
Released 2026-03-18large model

MiniMax M2.5

MiniMax

MiniMax M2.5 extends M2.1's coding into office work: it generates and operates Word, Excel and PowerPoint files, switches software environments, and works alongside other agents. Scores 80.2% on SWE-Bench Verified, 51.3% on Multi-SWE-Bench and 76.3% on BrowseComp, on fewer tokens than earlier generations.

ConversationReasoningCode Generation+4
Released 2026-02-12large model

MiniMax M2-her

MiniMax

A dialogue-first model for roleplay, character chat and long multi-turn conversation. Holds tone and personality, supports rich message roles, and can learn pacing from example dialogue.

ConversationReasoning
Released 2026-01-27large model

MiniMax M2.1

MiniMax

MiniMax's 230B mixture-of-experts model with open weights under MIT, tuned for coding and agent work. Available from MiniMax's own platform, Hugging Face and OpenRouter.

ConversationReasoningCode Generation+1
Released 2025-12-23large model

MiniMax M2

MiniMax

MiniMax M2: 10B activated parameters for coding agents, compile-run-fix loops and long-horizon reasoning. Good SWE-Bench and Terminal-Bench results, and it holds reasoning continuity across multi-step tasks.

ConversationReasoningCode Generation+4
Released 2025-10-23medium model

MiniMax M1

MiniMax

MiniMax M1, an open-weight reasoning model with 456B total parameters and 45.9B active per token. It pairs a mixture-of-experts backbone with a lightning attention mechanism to hold context up to 1 million tokens without the usual FLOP cost, and is tuned for multi-step reasoning.

ConversationReasoningCode Generation+3
Released 2025-06-17Extra large model

MiniMax Image-01

MiniMax

MiniMax's first image generation model with character reference support for consistent character generation across multiple images. Produces high-quality, detailed images with strong compositional abilities.

Image Generation
Released 2025-06-01large model

MiniMax Speech-02 HD

MiniMax

High-definition text-to-speech with studio-quality output and fine-grained emotion control.

Text To SpeechAudio Generation
Released 2025-03-10large model

Mistral

+11

Mistral Small 4

Mistral AI

Mistral Small 4 unifies the capabilities of Magistral (reasoning), Pixtral (multimodal), and Devstral (agentic coding) into a single efficient model. It handles complex analysis, software development, and visual tasks within the same workflow at $0.15/$0.60 per million tokens.

ConversationAnalysisCode Generation+2
Released 2026-03-16small model

Mistral Small Creative

Mistral AI

Mistral Small Creative is an experimental small model designed for creative writing, narrative generation, roleplay and character-driven dialogue, general-purpose instruction following, and conversational agents.

ConversationAnalysis
Released 2025-12-16small model

Devstral 2 2512

Mistral AI

The free OpenRouter tier of Devstral 2, Mistral's 123B coding model. Open weights under a modified MIT licence, also sold first-party and through NVIDIA NIM.

ConversationReasoningCode Generation+1
Released 2025-12-09Extra large model

Mistral Large 3 2512

Mistral AI

The December 2025 build of Mistral Large 3, a 675B instruct model with open weights on Hugging Face. Sold first-party through Mistral AI Studio and resold on Bedrock, Azure and watsonx.

ConversationReasoningCode Generation+1
Released 2025-12-01Extra large model

Mistral Medium 3.1

Mistral AI

An update to Mistral Medium 3, pitched at large-model quality for roughly an eighth of the cost. Strongest on coding and STEM reasoning, competitive with Claude Sonnet 3.7 and Llama 4 Maverick, and deployable hybrid, on-prem or inside a VPC.

ConversationReasoningCode Generation+1
Released 2025-08-13large model

Mistral Devstral Medium

Mistral AI

A code generation and agentic reasoning model built jointly by Mistral AI and All Hands AI, a step up from Devstral Small. Scores 61.6% on SWE-Bench Verified, ahead of Gemini 2.5 Pro and GPT-4.1 on code tasks at a fraction of the price.

ConversationReasoningCode Generation+1
Released 2025-07-11medium model

Mistral Devstral Small 1.1

Mistral AI

Devstral Small 1.1, a 24B open-weight model for software engineering agents, from Mistral AI with All Hands AI. Finetuned from Mistral Small 3.1, Apache 2.0, 128k context, with Mistral-style function calling and XML output.

ConversationReasoningCode Generation+1
Released 2025-07-11small model

Mistral Medium 3

Mistral AI

Mistral's mid tier model, pitched at large-model quality for roughly an eighth of the cost. Strongest on coding and STEM reasoning, and deployable hybrid, on-prem or inside a VPC rather than API only.

ConversationReasoningCode Generation+1
Released 2025-05-07large model

Mistral Large 2

Mistral AI

Mistral Large 2 features a 128K context window with enhanced code generation, mathematics, reasoning, and multilingual support.

ConversationReasoningCode Generation
Released 2024-07-24large model

Mistral Nemo

Mistral AI

Mistral Neom 3 is a 12B parameter model with a 128k token context length built by Mistral in collaboration with NVIDIA.

ConversationCode Generation
Released 2024-07-19large model

Mistral Large

Mistral AI

Mistral Large: strong multilingual work and reasoning, with a 32K token context window.

ConversationReasoningCode Generation
Released 2024-02-26large model

Moonshot

+8

Kimi K3

Moonshot AI

Moonshot AI's open-weight multimodal reasoning model at very large scale. Built for complex coding, knowledge work and long agent runs, strongest at navigating large repositories and iterating against images, logs, tests and runtime feedback. Uses KDA and Attention Residuals for efficiency.

ConversationReasoningCode Generation+3
Released 2026-07-16Extra large model

Kimi K2.7 Code

Moonshot AI

The coding model in Moonshot AI's K2.7 line, built to finish programming tasks end to end over long contexts. A native multimodal mixture of experts taking text, image and video, always in thinking mode, with the full reasoning kept across turns. 256K context, 32B active parameters out of roughly 1T.

ConversationReasoningCode Generation+3
Released 2026-06-13Extra large model

Kimi K2.6

Moonshot AI

Moonshot AI's K2.6, aimed at long-horizon coding, UI generation from prompts or images, and multi-agent orchestration. It works end to end across Python, Rust and Go, and its agent swarm scales to hundreds of parallel sub-agents, returning documents, websites and spreadsheets in one run without supervision.

ConversationReasoningCode Generation+3
Released 2026-04-20large model

Kimi K2.5

Moonshot AI

Moonshot AI's natively multimodal K2.5, continuing pretraining on Kimi K2 over roughly 15T mixed visual and text tokens. Its two selling points are visual coding and a self-directed agent swarm, alongside general reasoning and tool calling.

ConversationReasoningCode Generation+1
Released 2026-01-27large model

Kimi Linear 48B A3B Instruct

Moonshot AI

Kimi Linear is a hybrid linear attention architecture that outperforms traditional full attention methods. Features Kimi Delta Attention (KDA) for efficient memory usage, reducing KV caches by up to 75% and boosting throughput by up to 6x for contexts as long as 1M tokens.

ConversationReasoningCode Generation+1
Released 2025-11-10Extra large model

Kimi K2 Thinking

Moonshot AI

The reasoning model in Moonshot AI's K2 line, on the same trillion-parameter mixture-of-experts backbone with 32 billion active per pass and 256K context. It interleaves thinking with tool calls, and is built to hold a research or coding task together across hundreds of sequential actions without drifting.

ConversationReasoningCode Generation+2
Released 2025-11-06Extra large model

Kimi K2 0905

Moonshot AI

The September refresh of Kimi K2, still 1 trillion total parameters with 32 billion active per pass, now holding 256K context instead of 128K. The update targets agentic coding accuracy across different scaffolds and front-end output worth looking at, for web and 3D work.

ConversationReasoningCode Generation+2
Released 2025-09-04Extra large model

Kimi K2

Moonshot AI

Moonshot AI's Kimi K2, a mixture-of-experts model with 1 trillion total parameters and 32 billion active per forward pass. Tuned for agentic work: tool use, reasoning and code synthesis, with 128K context. Trained with the MuonClip optimizer, which is what kept a MoE run at this scale stable.

ConversationReasoningCode Generation+1
Released 2025-07-11Extra large model

NVIDIA

+5

Nemotron 3.5 Lightning

NVIDIA

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that benefit from domain-specific customization. A free-tier endpoint is also available on OpenRouter.

ConversationReasoningCode Generation+2
Released 2026-08-11medium model

Nemotron 3 Ultra

NVIDIA

NVIDIA's open reasoning and orchestration model, 55B active parameters out of 550B on a hybrid Transformer-Mamba mixture of experts. Text in, text out, with a context window up to 1M tokens. Built for long-running agent work: orchestration, coding agents, deep research and multi-step reasoning.

ConversationReasoningCode Generation+4
Released 2026-06-04Extra large model

Nemotron 3.5 Content Safety

NVIDIA

A 4B multimodal guardrail model from NVIDIA, fine-tuned from Gemma-3-4B. It moderates both the input to an LLM and the response, taking text and images and returning a safe or unsafe classification for each, category labels, and an optional reasoning trace. 12 languages, 128K context.

AnalysisData Extraction
Released 2026-06-04small model

NVIDIA Nemotron 3 Super (free)

NVIDIA

NVIDIA's Nemotron 3 Super: a 120B open hybrid mixture-of-experts activating 12B parameters per token. The Mamba-Transformer backbone with multi-token prediction generates over 50% more tokens than comparable open models, and Latent MoE calls 4 experts for the cost of one. 1M context, NVIDIA Open License.

ConversationReasoningCode Generation+3
Released 2026-03-11large model

NVIDIA Nemotron Nano 9B V2

NVIDIA

NVIDIA trained Nemotron Nano 9B v2 from scratch as one model for both reasoning and non-reasoning work. It can expose an internal reasoning trace before the answer, or be told by system prompt to give the answer alone.

ConversationReasoningCode Generation+1
Released 2025-09-05medium model

OpenRouter

+26

Ox Alpha

OpenRouter

A cloaked reasoning model on OpenRouter, out for community feedback. Built for coding, sustained agent work and production loads: long-horizon software engineering, complex reasoning, and multimodal workflows mixing text with images and video. Reasoning always on. 1M token context, 131K output.

ConversationReasoningCode Generation+4
Released 2026-08-24large model

OpenRouter Fusion · Quality (Jun 2026)

OpenRouter

OpenRouter Fusion, Quality preset, June 2026 snapshot. One prompt goes in parallel to Claude Opus 4.8, GPT-5.5 and Gemini 3 Pro with web search on, then Claude Opus 4.8 judges their answers into one. A request costs every panel member plus the judge. The panel drifts over time.

ConversationReasoningAnalysis+1
Released 2026-06-13Extra large model

OpenRouter Fusion · Budget (Jun 2026)

OpenRouter

OpenRouter Fusion, Budget preset, June 2026 snapshot. The same deliberation on cheaper members: Gemini 3 Flash, DeepSeek V3.2 and Kimi K2.7 Code answer in parallel with web search on, then Gemini 3 Flash judges. Still costs every member plus the judge, far less than Quality.

ConversationReasoningAnalysis+1
Released 2026-06-13Extra large model

Owl Alpha

OpenRouter

A stealth foundation model for agent workloads, with native tool use and long context. Strong on code generation, automated workflows and complex instruction execution, and compatible with Claude Code, OpenClaw and similar tools. Prompts and completions may be logged by the provider.

ConversationReasoningCode Generation+3
Released 2026-04-28large model

Elephant Alpha

OpenRouter

Elephant Alpha is a 100B-parameter text model focused on intelligence efficiency, delivering strong reasoning performance while minimizing token usage. It supports a 256K context window with up to 32K output tokens, function calling, structured output, and prompt caching.

ConversationReasoningCode Generation+3
Released 2026-04-13large model

Hunter Alpha

OpenRouter

A stealth model on OpenRouter with roughly 1 trillion parameters and 1M tokens of context, built for agent work: long-horizon planning, multi-step execution, and the instruction-following precision frameworks like OpenClaw need. Prompts and completions are logged by the provider and may be used for training.

ConversationReasoningCode Generation+3
Released 2026-03-11Extra large model

Healer Alpha

OpenRouter

A stealth omni-modal model on OpenRouter that takes visual and audio input natively, reasons across both, and acts on the result through multi-step tasks. Prompts and completions are logged by the provider and may be used for training.

ConversationReasoningCode Generation+3
Released 2026-03-11Extra large model

Aurora Alpha

OpenRouter

A cloaked reasoning model from OpenRouter, out for community feedback. Built for speed: coding assistants, real-time conversation and agent workflows. Default reasoning effort is medium; high is recommended for agentic coding.

ConversationReasoningCode Generation+2
Released 2026-02-09large model

Pony Alpha

OpenRouter

A stealth model released through OpenRouter to collect community feedback, strong on coding, agent workflows, reasoning and roleplay. 200K context, with tool calling, structured outputs and step-by-step reasoning.

ConversationCode GenerationReasoning+3
Released 2026-02-06large model

GLM 4.7

OpenRouter

Z.ai's 358B open-weight model under MIT, aimed at coding and agent loops. Runs on the first-party Z.ai API, on OpenRouter, and on Cerebras for high-speed inference.

ConversationReasoningCode Generation+1
Released 2025-12-22large model

TNG R1T Chimera

OpenRouter

TNG's November 2025 rebuild of its R1T Chimera merge, a DeepSeek R1 derivative. It runs only on OpenRouter's free tier, with no separate weights repository published.

ConversationReasoning
Released 2025-11-27large model

INTELLECT-3

OpenRouter

Prime Intellect's 106B mixture-of-experts model, 12B active per token, post-trained from GLM-4.5-Air-Base with supervised fine-tuning and reinforcement learning. Open weights under MIT.

ConversationReasoningCode Generation+1
Released 2025-11-27large model

Bert-Nebulon Alpha

OpenRouter

A cloaked OpenRouter endpoint from late November 2025, later revealed as an early build of Mistral Large 3. Free while it ran, prompts logged by the provider, now retired.

ConversationReasoningAnalysis
Released 2025-11-24large model

Sherlock Dash Alpha

OpenRouter

A cloaked OpenRouter endpoint from November 2025, later revealed as an early snapshot of xAI's Grok 4.1 Fast with reasoning switched off. Free while it ran, prompts logged by the provider.

ConversationReasoningCode Generation+2
Released 2025-11-15Extra large model

Sherlock Think Alpha

OpenRouter

The thinking half of OpenRouter's Sherlock alpha pair, later revealed as an early snapshot of xAI's Grok 4.1 Fast with reasoning switched on. Free while it ran, prompts logged by the provider.

ConversationReasoningCode Generation+2
Released 2025-11-15Extra large model

Polaris Alpha

OpenRouter

A general-purpose model, strongest on coding, tool calling and instruction following.

ConversationReasoningCode Generation+2
Released 2025-11-06large model

Andromeda Alpha

OpenRouter

Smaller reasoning model trained for image and visual understanding. Cloaked trial; prompts and outputs logged for provider feedback.

ConversationReasoningAnalysis
Released 2025-10-21medium model

Sonoma Dusk Alpha

OpenRouter

A cloaked model OpenRouter ran for community feedback: fast, general purpose, with a 2 million token context window. Takes image input and calls tools in parallel.

ConversationReasoningCode Generation+1
Released 2025-09-05Extra large model

Sonoma Sky Alpha

OpenRouter

A cloaked model OpenRouter ran in September 2025 with a two million token context. It was later revealed as the reasoning variant of xAI's Grok 4 Fast, and the alpha endpoint has since been retired.

ConversationReasoningCode Generation+3
Released 2025-09-05Extra large model

Horizon Beta

OpenRouter

This is a cloaked model provided to the community to gather feedback. This is an improved version of Horizon Alpha. Note: It's free to use during this testing period, and prompts and completions are logged by the model creator for feedback and training.

ConversationReasoningCode Generation+1
Released 2025-08-01large model

Horizon Alpha

OpenRouter

This is a cloaked model provided to the community to gather feedback. Note: It's free to use during this testing period, and prompts and completions are logged by the model creator for feedback and training.

ConversationReasoningCode Generation+1
Released 2025-07-30large model

Cypher Alpha (free)

OpenRouter

A cloaked model OpenRouter offered free in July 2025 to gather feedback, with a one million token context. The lab behind it was never publicly confirmed, prompts were logged, and the endpoint is now retired.

ConversationReasoningCode Generation+1
Released 2025-07-01large model

Optimus Alpha

OpenRouter

A stealth general-purpose model for long-context work, including code generation. Shipped on community feedback.

ConversationWeb DesignCode Generation+1
Released 2025-04-10large model

Quasar Alpha

OpenRouter

A cloaked model released for community feedback. General-purpose, handles long context and code generation. Every prompt and completion is logged by the provider and by OpenRouter.

ConversationWeb DesignCode Generation+1
Released 2025-04-02large model

F5-TTS

OpenRouter

Open-source zero-shot text-to-speech with natural prosody. Generates highly natural speech from just a reference audio clip.

Text To SpeechAudio Generation
Released 2024-10-08medium model

Bark

OpenRouter

Open-source audio model: speech, music and sound effects, including laughter, crying and other non-verbal sounds.

Text To SpeechAudio GenerationMusic Generation
Released 2023-04-20medium model

Orpheus

+1

Orpheus 3B

Orpheus

A 3B text-to-speech model for expressive delivery: nuanced emotion, natural pauses, real vocal range.

Text To Speech
Released 2025-03-20large model

Perplexity

+1

Sonar Pro Search

Perplexity

Sonar Pro's Pro Search mode, exclusive to the OpenRouter API. Perplexity's agentic search system, built for deeper analysis. Priced per token plus $18 per thousand requests.

ConversationReasoningAnalysis
Released 2025-10-30large model

Playground

+1

Playground v2.5 (Aesthetic)

Playground AI

An aesthetic-focused Playground v2.5 model geared toward pleasing composition and style at 1024px resolution.

Image Generation
Released 2024-02-27large model

Pruna AI

+3

P-Image

PrunaAI

A sub 1-second text-to-image model built for production use cases by PrunaAI.

Image Generation
Released 2026-01-13large model

Z Image Turbo

PrunaAI

A turbo text-to-image model optimized by PrunaAI for very fast inference at low guidance.

Image Generation
Released 2025-11-26large model

HiDream L1 (Fast)

PrunaAI

PrunaAI-optimized HiDream L1 Fast for cheap, fast text-to-image generation with configurable speed modes.

Image Generation
Released 2025-04-10large model

Qwen

+38

Qwen3.8 27B

Qwen

Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be enabled or disabled.

ConversationReasoningCode Generation+3
Released 2026-08-14medium model

Qwen3.8 2.4T A95B

Qwen

Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of Qwen3.8 Max, with 95 billion active parameters out of 2.4 trillion total. It is suited for coding, research, complex reasoning, and agentic workflows.

ConversationReasoningCode Generation+3
Released 2026-08-12Extra large model

Qwen3.8 Max

Qwen

The top tier of Alibaba's Qwen3.8 series and the general-availability successor to Qwen3.8 Max Preview. A multimodal reasoning model for complex reasoning, visual understanding, coding and agent workflows.

ConversationReasoningCode Generation+4
Released 2026-08-03Extra large model

Qwen3.7 Max

Qwen

The largest model in Alibaba's Qwen3.7 series, text in and text out. Built for agent-centric work: coding, office and productivity tasks, and long-horizon autonomous execution. Clear gains on coding and agent behaviour over earlier Qwen generations, with explicit prompt caching for repeated context.

ConversationReasoningCode Generation+4
Released 2026-05-22Extra large model

Qwen3.7 Plus

Qwen

The cost-effective tier of Alibaba's Qwen3.7 series, taking text and image. Carries the series' coding, tool use and productivity work, plus a large vision upgrade. The distinguishing trait is interactive multimodal agency: reading screens, driving GUIs, coding from a visual reference, navigating mobile apps.

ConversationReasoningCode Generation+4
Released 2026-05-22large model

Qwen3.5 Plus 2026-04-20

Qwen

Qwen3.5 Plus (April 2026) is a large-scale multimodal language model from Alibaba. It accepts text, image, and video input and produces text output, with a 1M token context window. This is an updated version of Qwen3.5 Plus with tiered pricing above 256K tokens.

ConversationReasoningCode Generation+4
Released 2026-04-27Extra large model

Qwen3.6 Flash

Qwen

Qwen3.6 Flash is a fast, efficient language model from Alibaba's Qwen 3.6 series. It supports text, image, and video input with a 1M token context window. Tiered pricing kicks in above 256K tokens. Prompt caching is supported, with both explicit cache read and cache creation pricing.

ConversationReasoningCode Generation+3
Released 2026-04-27medium model

Qwen3.6 35B A3B

Qwen

An open-weight multimodal model from Alibaba Cloud, 35B total with 3B active per token. Hybrid sparse mixture-of-experts mixing Gated DeltaNet linear attention with standard gated attention. 262K native context, extensible to 1M with YaRN. Text, image and video. Thinking mode built in. Apache 2.0.

ConversationReasoningCode Generation+3
Released 2026-04-27medium model

Qwen3.6 27B

Qwen

A dense 27B from Alibaba's Qwen team, released April 2026. Text, image and video across a 262,144 token context. Built for agentic coding and reasoning, clearest on repository-level code comprehension and front-end work. Thinking mode built in and persistent. 201 languages, Apache 2.0.

ConversationReasoningCode Generation+4
Released 2026-04-27medium model

Qwen3.6 Max Preview

Qwen

Alibaba Cloud's proprietary Qwen3.6 Max, a sparse mixture of experts at roughly 1 trillion total parameters. For agentic coding, tool use and long-context reasoning over a 262K window, with a thinking mode that keeps traces across turns. Model Studio and Qwen Studio only, no open weights.

ConversationReasoningCode Generation+4
Released 2026-04-27Extra large model

Qwen3.6 Plus Preview (free)

Qwen

The preview of Qwen 3.6 Plus, on a hybrid architecture that improves efficiency and scaling over the 3.5 series, with more reliable agent behaviour. Strongest on agentic coding and front-end development. Free tier on OpenRouter, which collects prompt and completion data.

ConversationReasoningCode Generation+1
Released 2026-03-30large model

Qwen3.5 9B

Qwen

A multimodal foundation model in the Qwen3.5 family at 9B parameters. Unified vision-language design with early fusion of multimodal tokens: text, image and video in, text out, with reasoning built in.

ConversationReasoningCode Generation+1
Released 2026-03-10small model

Qwen3.5 35B A3B

Qwen

The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall performance is comparable to that of the Qwen3.5-27B.

ConversationReasoningCode Generation+4
Released 2026-02-25medium model

Qwen3.5 27B

Qwen

The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference speed and performance. Its overall capabilities are comparable to those of the Qwen3.5-122B-A10B.

ConversationReasoningCode Generation+4
Released 2026-02-25medium model

Qwen3.5 122B A10B

Qwen

The 122B-A10B vision-language model in the Qwen3.5 series, on a hybrid of linear attention and a sparse mixture of experts. Second only to Qwen3.5-397B-A17B overall, with text well ahead of Qwen3-235B-2507 and vision ahead of Qwen3-VL-235B.

ConversationReasoningCode Generation+4
Released 2026-02-25large model

Qwen3.5 Flash

Qwen

The Flash tier of Qwen3.5, a native vision-language model on a hybrid of linear attention and a sparse mixture of experts. Built for fast responses, and a clear step up from the Qwen3 series on both text and multimodal tasks.

ConversationReasoningCode Generation+4
Released 2026-02-25medium model

Qwen3.5 Plus 2026-02-15

Qwen

The Plus tier of Qwen3.5, a native vision-language model on a hybrid of linear attention and sparse mixture of experts for cheaper inference. Alibaba reports it level with leading models, and a clear step up from Qwen3 on text and multimodal work. Text, image and video in, with reasoning and tool use.

ConversationReasoningCode Generation+4
Released 2026-02-16Extra large model

Qwen3.5 397B A17B

Qwen

The 397B-A17B vision-language model in the Qwen3.5 series, on a hybrid of linear attention and sparse mixture of experts. Alibaba reports it level with leading models across language, logic, code, agent tasks, image and video understanding and GUI interaction. 201 languages and dialects.

ConversationReasoningCode Generation+4
Released 2026-02-16Extra large model

Qwen3 Max Thinking

Qwen

The reasoning tier of Qwen3-Max, scaled on both model capacity and reinforcement learning compute. Heavy Mode iterates on an answer at test time, tools include search and a code interpreter, and it can switch between normal and compute-heavy reasoning mid-conversation.

ConversationReasoningCode Generation+2
Released 2026-02-09Extra large model

Qwen3 Coder Next

Qwen

An open-weight model for coding agents and local development. Sparse MoE with 80B total parameters and 3B activated per token, performing like models with 10 to 20x the active compute. Non-thinking mode only.

ConversationCode GenerationAgentic Tool Use+1
Released 2026-02-04large model

Qwen Image

Qwen

An image generation foundation model in the Qwen series that achieves significant advances in complex text rendering.

Image Generation
Released 2026-01-20large model

Qwen Image (Fast)

Qwen

A fast Qwen text-to-image model optimized by PrunaAI for speed on Replicate.

Image Generation
Released 2025-12-11large model

Qwen3 Coder Plus

Qwen

The larger of Alibaba's two hosted Qwen3 Coder tiers, sold through Model Studio rather than released as open weights. 128K context, with OpenAI-compatible and Anthropic-compatible endpoints.

ConversationReasoningCode Generation+1
Released 2025-09-17large model

Qwen3 Coder Flash

Qwen

The cheap tier of Alibaba's hosted Qwen3 Coder line, aimed at autocomplete and quick edits where latency matters more than depth. 128K context, Model Studio only, no open weights.

ConversationReasoningCode Generation+1
Released 2025-09-17medium model

Qwen3 Next 80B A3B Instruct

Qwen

The instruct member of Qwen3-Next, tuned for fast, stable answers with no visible chain of thought. It prioritises throughput and consistency on very long inputs and multi-turn dialogue, which is what makes it usable for retrieval-augmented generation and tool-calling agents.

ConversationReasoningCode Generation+1
Released 2025-09-11large model

Qwen3 Next 80B A3B Thinking

Qwen

The reasoning member of Qwen3-Next, emitting thinking traces by default in thinking-only mode. Built for maths proofs, code synthesis and debugging, logic and agent planning, and tuned to stay stable across long chains without drifting off task.

ConversationReasoningCode Generation+1
Released 2025-09-11large model

Qwen Plus 0728 (thinking)

Qwen

Qwen Plus 0728 with reasoning switched on. The same Qwen3-based 1 million token hybrid model, spending extra tokens to think before it answers.

ConversationReasoningCode Generation+1
Released 2025-09-08large model

Qwen Plus 0728

Qwen

Qwen Plus 0728, built on Qwen3: a hybrid reasoning model with a 1 million token context, balancing quality, speed and cost.

ConversationReasoningCode Generation+1
Released 2025-09-08large model

Qwen3 Max

Qwen

Alibaba's largest Qwen3 model, an update on the January 2025 release with better mathematics, coding, logic and science, and fewer hallucinations on open-ended questions. Over 100 languages, tuned for retrieval-augmented generation and tool calling. No dedicated thinking mode.

ConversationReasoningCode Generation+3
Released 2025-09-05Extra large model

Qwen3 30B A3B Thinking 2507

Qwen

The thinking variant of Qwen3-30B-A3B, a 30B mixture-of-experts model that keeps its reasoning trace separate from the final answer. Longer output budgets than earlier 30B releases, with gains across logic, mathematics, science, coding and multilingual benchmarks.

ConversationReasoningCode Generation+1
Released 2025-08-29large model

Qwen3 30B A3B Instruct 2507

Qwen

A 30.5B mixture-of-experts model from Qwen with 3.3B active parameters per inference, running in non-thinking mode only. It answers directly rather than reasoning aloud, and beats the non-instruct variant on open-ended and subjective tasks while holding its coding and factual scores.

ConversationReasoningCode Generation+1
Released 2025-07-29large model

Qwen3 235B A22B Thinking 2507

Qwen

The thinking-only variant of Qwen3-235B-A22B, activating 22B of 235B parameters and holding 262,144 tokens of context. It always emits a reasoning trace and is built for long outputs, up to 81,920 tokens, on maths, science and long-form generation.

ConversationReasoningCode Generation+2
Released 2025-07-25Extra large model

Qwen3 Coder

Qwen

Qwen's 480B mixture-of-experts coding model, activating 35B parameters per pass from 8 of 160 experts. Built for agentic coding: function calling, tool use, and reasoning over a whole repository rather than a single file.

ConversationReasoningCode Generation+3
Released 2025-07-23Extra large model

Qwen3 235B A22B 2507

Qwen

The July 2025 instruct refresh of Qwen3-235B-A22B, activating 22B of 235B parameters per pass. Native 262K context and no thinking mode: it answers directly. Gains over the base variant are largest in knowledge coverage, long-context reasoning, coding and multilingual maths.

ConversationReasoningCode Generation+1
Released 2025-07-21Extra large model

Qwen3 0.6B

Qwen

A 0.6B dense model from the Qwen3 family, small enough to run on a laptop. Switches between thinking mode for hard tasks and non-thinking mode for chat. Trained on 36 trillion tokens across 119 languages, with tool use and multilingual support.

ConversationCode Generation
Released 2025-04-29small model

Qwen3 30B A3B

Qwen

Qwen3 at 30.5B parameters with 3.3B activated. Reasoning, multilingual work and agent tasks, with a thinking/non-thinking mode switch. Up to 131K context with YaRN. Free tier on OpenRouter.

ConversationReasoningCode Generation+1
Released 2025-04-28large model

Qwen3 235B A22B

Qwen

A 235B mixture-of-experts model from Alibaba's Qwen team, activating 22B parameters per forward pass. Switches between a thinking mode for hard tasks and a non-thinking mode for chat. Strong on reasoning, tool calling and over 100 languages. 32K context, extendable to 131K.

ConversationReasoningCode Generation+1
Released 2025-04-28Extra large model

QwQ 32B

Qwen

Qwen's 32B reasoning model. It thinks before answering, which buys it a wide margin over conventional instruction-tuned models on hard problems and puts it in range of much larger reasoning models like DeepSeek-R1 and o1-mini.

ConversationReasoningCode Generation+1
Released 2025-03-05large model

recraft-ai

+8

Recraft V4.1

Recraft

Recraft V4.1 refines V4's design taste with quieter, more natural photorealism, dreamier gradients, and sharper illustration styles. The model still makes intentional decisions about composition, lighting, and material realism while rendering accurate text from prompts up to 10,000 characters.

Image Generation
Released 2026-05-14large model

Recraft V4.1 Pro

Recraft

Recraft V4.1 Pro pairs V4.1's refined photorealism and sharper illustration with maximum-quality, high-resolution output. Built for final production assets, print, and large-format work where detail and fidelity matter most.

Image Generation
Released 2026-05-14large model

Recraft V4

Recraft

A ground-up rebuild of Recraft's image generation with design judgment built in: the model makes deliberate choices about composition, lighting, colour and material realism. Takes prompts up to 10,000 characters and renders text accurately.

Image Generation
Released 2026-02-17large model

Recraft V4 Pro

Recraft

Recraft V4 Pro renders at roughly 2048px, scalable to 3072px, with the same design taste as V4. Built for print, large-format displays and detailed professional work.

Image Generation
Released 2026-02-17large model

Recraft V4 SVG

Recraft

Recraft V4 SVG generates production-ready SVG vector images from text prompts with real paths, structured layers, and clean geometry. The only image model that produces native, editable vector output.

Image Generation
Released 2026-02-17large model

Recraft V4 Pro SVG

Recraft

Recraft V4 Pro SVG combines V4 Pro's design taste with detailed geometric vector output. Produces detailed SVG files with finer paths, clean layers, and editable output suitable for professional design systems.

Image Generation
Released 2026-02-17large model

Recraft V3

Recraft

Recraft V3 (code-named red_panda) is a SOTA text-to-image model ranked #1 on Artificial Analysis benchmarks. The first model to offer image generation with text of any size and length, supporting both raster and vector output formats across a wide list of styles.

Image Generation
Released 2024-10-29large model

Recraft V3 SVG

Recraft

Recraft V3 SVG generates high-quality vector SVG images from text prompts, including logotypes, icons, and illustrations. Produces real editable SVG files with clean paths and structured layers.

Image Generation
Released 2024-10-29large model

sourceful

+4

Riverflow V2.5 Pro

Sourceful

The highest quality variant of Sourceful's Riverflow 2.5, for work where control matters more than speed. The series treats image generation as a production workflow: a reasoning model plans multi-step edits and judges candidates before accepting one. Reasoning effort is set on the request.

Image Generation
Released 2026-06-04large model

Riverflow V2.5 Fast

Sourceful

The speed-optimized variant of Sourceful's Riverflow 2.5, for production deployments with a latency budget. Like the rest of the series it plans multi-step edits with a reasoning model and judges candidates before accepting one, with reasoning effort set on the request.

Image Generation
Released 2026-06-04large model

Riverflow V2 Pro

Sourceful

The top variant of Sourceful's Riverflow 2.0 line, for fine control and text rendering. The series runs an integrated reasoning model over image generation and editing.

Image Generation
Released 2026-02-02large model

Riverflow V2 Fast

Sourceful

The fastest variant of Sourceful's Riverflow 2.0 line, for production and latency-critical work. The series runs an integrated reasoning model over image generation and editing.

Image Generation
Released 2026-02-02large model

Stability

+3

Stable Audio 2.5

Stability AI

Stability AI's latest audio generation model for music, sound effects, and ambient audio up to 3 minutes long.

Music GenerationAudio Generation
Released 2025-04-20large model

Stable Diffusion 3.5 Medium

Stability AI

Stable Diffusion 3.5 Medium balances quality and speed, offering modern diffusion performance with broad aspect ratio support.

Image Generation
Released 2024-10-29large model

Stable Diffusion XL

Stability AI

Stable Diffusion XL (SDXL), a widely used open text-to-image diffusion model known for versatility and community tooling.

Image Generation
Released 2023-07-26large model

tencent

+2

Hy4 Preview

Tencent

Tencent's 770B-parameter mixture-of-experts model activates 49B parameters per token. It is built for coding agents and tool-heavy workflows that need planning, context continuity and sustained execution across a 1M-token window.

ConversationReasoningCode Generation+5
Released 2026-08-28Extra large model

Hunyuan Image 3

Tencent

Tencent's native multimodal image model, tuned by PrunaAI for speed. Strong prompt understanding across a range of visual styles.

Image Generation
Released 2025-10-01large model

thinkingmachines

+1

Inkling

Thinking Machines

An open-weight multimodal mixture-of-experts model from Thinking Machines Lab, 41B active parameters out of 975B. Built for general reasoning, coding, agent and tool-use systems, retrieval-augmented generation and multilingual conversation, with native image and audio understanding alongside text.

ConversationReasoningCode Generation+3
Released 2026-07-15Extra large model

Upstage

+1

Solar Pro 3

Upstage

Upstage's Mixture-of-Experts model: 102B total parameters, 12B active per forward pass. Tuned for Korean, with English and Japanese support. Deprecating 2 March 2026.

ConversationReasoningCode Generation+1
Released 2026-01-27large model

Xiaomi

+5

MiMo-V2.5

Xiaomi

Xiaomi's native omnimodal MiMo-V2.5, roughly half the inference cost of the Pro tier and ahead of MiMo-V2-Omni on image and video understanding. Its 1M context takes whole documents, long conversations and complex task state in one pass, which is what makes it usable inside agent frameworks.

ConversationReasoningCode Generation+2
Released 2026-04-22large model

MiMo-V2.5-Pro

Xiaomi

Xiaomi's largest MiMo-V2.5, ranked near the top on ClawEval, GDPVal and SWE-bench Pro. It runs professional tasks that would take a human expert days, chaining more than a thousand tool calls without supervision, across a context of up to 1M tokens.

ConversationReasoningCode Generation+3
Released 2026-04-22Extra large model

MiMo-V2-Pro

Xiaomi

Xiaomi's largest MiMo model, over 1T total parameters with a 1M context, tuned for agent frameworks like OpenClaw. Xiaomi places it in the top tier on PinchBench and ClawBench, near Opus 4.6. The intended job is orchestrating workflows and production engineering tasks rather than answering single questions.

ConversationReasoningCode Generation+3
Released 2026-03-18Extra large model

MiMo-V2-Omni

Xiaomi

An omni-modal model that takes image, video and audio in one architecture. Adds visual grounding, multi-step planning, tool use and code execution for tasks that span modalities.

ConversationReasoningCode Generation+2
Released 2026-03-18large model

MiMo-V2-Flash

Xiaomi

Xiaomi's open-source MiMo-V2-Flash, a mixture-of-experts with 309B total and 15B active on a hybrid attention design. 256K context, thinking toggle. Xiaomi reports it top open-source on SWE-bench Verified and SWE-bench Multilingual, near Claude Sonnet 4.5 at about 3.5% of the cost.

ConversationReasoningCode Generation+2
Released 2025-12-14Extra large model

z-ai

+2

GLM 5.1

Z.ai

GLM-5.1's gain over its predecessor is duration. Where earlier models were built around minute-long interactions, it works a single task for more than 8 hours, planning, executing and correcting itself until it has something engineering-grade to hand back.

ConversationReasoningCode Generation+4
Released 2026-04-07large model

GLM 5 Turbo

Z.ai

Z.ai's GLM-5 Turbo, for fast inference in agent-driven environments such as OpenClaw. Tuned for long execution chains: instruction decomposition, tool use, scheduled and persistent execution, and stability across extended tasks.

ConversationReasoningCode Generation+5
Released 2026-03-15medium model

Zhipu

+7

GLM 5.2

Zhipu AI

Z.ai's GLM-5.2, built for long-horizon work. A genuinely usable 1M-token context window lets it hold project-level engineering context, run long tasks, follow engineering standards, and carry a project from requirements to multi-platform deployment in one task.

ConversationReasoningCode Generation+4
Released 2026-06-16Extra large model

GLM 5

Zhipu AI

Z.ai's largest open-source model, built for systems design and agent runs that last. It plans, executes and self-corrects across large programming tasks rather than emitting a function and stopping, and Z.ai positions it against leading closed models on production code.

ConversationReasoningCode Generation+3
Released 2026-02-11Extra large model

GLM 4.7 Flash

Zhipu AI

A 30B-class GLM tuned for agentic coding: coding strength, long-horizon task planning and tool collaboration. Z.ai reports top results among open-weight models of the same size on several public benchmarks.

ConversationReasoningCode Generation+1
Released 2026-01-27medium model

GLM 4.6

Zhipu AI

GLM 4.6 expands the GLM family with a 200K-token context window, stronger coding benchmarks, and more reliable multi-step reasoning. It integrates deeply with agent frameworks to orchestrate tool use and produces more natural writing for long-form chat.

ConversationReasoningCode Generation+2
Released 2025-09-30Extra large model

GLM 4.5

Zhipu AI

Z.ai's GLM-4.5, a mixture-of-experts model built for agent applications with 128K context. It runs in two modes: thinking, for multi-step reasoning and tool use, and non-thinking, for instant answers.

ConversationReasoningCode Generation+1
Released 2025-07-25large model

GLM 4.5 Air

Zhipu AI

The compact GLM-4.5, same mixture-of-experts design at a smaller parameter count. Keeps the two-mode split, thinking for reasoning and tool use, non-thinking for real-time chat, toggled by a boolean on the request.

ConversationReasoningAnalysis
Released 2025-07-25large model

GLM 4 32B

Zhipu AI

A cost-effective 32B foundation model with much stronger tool use, web search and code work than GLM 3. From the lab behind the THUDM models.

ConversationReasoningCode Generation+2
Released 2025-07-24large model

xAI

+14

Grok 4.6

xAI

SpaceXAI's Grok 4.6, aimed at coding, knowledge work and STEM. 500K-token context window, with text, image and file input, and reasoning enabled via the API.

ConversationReasoningCode Generation+3
Released 2026-08-12large model

Grok 4.5

xAI

xAI's Grok 4.5, aimed at coding, knowledge work and STEM. 500K-token context window, text and image in, reasoning enabled via the API.

ConversationReasoningCode Generation+1
Released 2026-07-08large model

Grok 4.3

xAI

Grok 4.3 is xAI's reasoning model with a 1M token context window. It accepts text and image inputs with text output, suited for agentic workflows, instruction-following tasks, and applications requiring high factual accuracy. Supports parallel tool calling, structured outputs, and exposed reasoning.

ConversationReasoningCode Generation+1
Released 2026-05-02large model

Grok 4.20 Beta

xAI

xAI's Grok 4.20 Beta, built for speed and agentic tool calling. xAI claims the lowest hallucination rate on the market and strict prompt adherence. Reasoning toggles on and off via the API.

ConversationReasoningCode Generation+4
Released 2026-03-12Extra large model

Grok 4.20 Multi-Agent Beta

xAI

A Grok 4.20 variant for agent-based workflows. Multiple agents run in parallel to research, coordinate tool use and synthesise across complex tasks. Reasoning effort sets agent count: low and medium use 4, high and xhigh use 16.

ConversationReasoningCode Generation+5
Released 2026-03-12Extra large model

Grok Imagine Image

xAI

xAI's image generation model, pairing Grok's prompt understanding with high-quality output. Aimed at detailed, creative images that follow the prompt closely.

Image Generation
Released 2025-12-01large model

Grok 4.1 Fast

xAI

xAI's fast tier of Grok 4.1, shipped alongside the Agent Tools API for search and code execution. Sold on the xAI API and Oracle Cloud, in reasoning and non-reasoning variants.

ConversationReasoningCode Generation+2
Released 2025-11-21Extra large model

Grok 4 Fast (free)

xAI

Grok 4 Fast is xAI's latest multimodal model with SOTA cost-efficiency and a 2M token context window. It comes in two flavors: non-reasoning and reasoning. Reasoning can be enabled via the API.

ConversationAnalysisCode Generation
Released 2025-09-19large model

Grok Code Fast 1

xAI

xAI's cheap, fast coding model, built for agents that make many small edits in an editor rather than one long deliberation. 256K context, available through the xAI API and Azure AI Foundry.

ConversationReasoningCode Generation+1
Released 2025-08-26medium model

Grok 4

xAI

Grok 4 is xAI's latest reasoning model with a 256k context window. It supports parallel tool calling, structured outputs, and both image and text inputs. Note that reasoning is not exposed, reasoning cannot be disabled, and the reasoning effort cannot be specified.

ConversationReasoningCode Generation+1
Released 2025-07-09large model

Grok 3 Mini Beta

xAI

A small thinking model for reasoning-heavy work that does not need deep domain knowledge. Strongest on math and quantitative tasks. The thinking traces are visible.

ConversationReasoningAnalysis
Released 2025-04-09small model

Grok 3 Beta

xAI

xAI's top Grok 3 tier, aimed at data extraction, coding and summarization, with domain knowledge in finance, healthcare, law and science. Beats Grok 3 Mini on high-thinking tasks.

ConversationReasoningCode Generation+2
Released 2025-04-09Extra large model

Grok 3 Thinking

xAI

Grok 3 Thinking exposes the full chain-of-thought process during problem-solving, including error backtracking and alternative solution exploration. Scores 84.6% on GPQA Diamond benchmark for expert-level Q&A.

ConversationReasoningCode Generation+1
Released 2025-02-19Extra large model

Grok 3

xAI

xAI's Grok 3, trained on the Colossus cluster, with a Big Brain Mode that spends extra compute on hard problems. Scored 1402 Elo on LMArena and 93.3% on AIME 2025.

ConversationReasoningCode Generation+1
Released 2025-02-18Extra large model

You saw a few. There are way more.

313+ models you haven't met. $0 to continue. Seriously.

Free account. No card required. By continuing, you agree to Rival's Terms and Privacy Policy

Rival
How it worksPrivacyTerms
Explore all of Rival

Explore

  • Compare Models
  • All Models
  • Image Comparison
  • Audio Comparison
  • Image Generation
  • Best AI For...
  • Arena
  • API Pricing
  • Challenges

Discover

  • SubjectiveBench
  • Research
  • Research downloads
  • Rival Kits
  • Find your AI taste
  • UI Glow-Up
  • VoiceLock
  • Cost Cutter
  • Agent skills
  • Benchmarks vs Vibes
  • Brief
  • Jailbreak
  • Model Updates
  • Provider Status
  • AI Creators

Connect

  • Methodology
  • Advertise
  • Partnerships
  • Privacy Policy
  • Terms
  • RSS Feed