DeepSeek V4 Flash Vision Exp
DeepSeek
Experimental vision-enabled version of DeepSeek V4 Flash 0731, adding image understanding to the base model. Supports text and image inputs with a 1M context window on OpenRouter.
368 models from 48 providers. Same prompts, real outputs, blind-vote results.
368 models
DeepSeek
Experimental vision-enabled version of DeepSeek V4 Flash 0731, adding image understanding to the base model. Supports text and image inputs with a 1M context window on OpenRouter.
OpenAI
The precision tier of OpenAI's GPT Image 2.5 pair and its most capable image model so far. Higher image quality than GPT Image 2, sharper detail, stronger style adherence and tighter control over edits. Built for campaign creative and polished product imagery.
OpenAI
The speed tier of OpenAI's GPT Image 2.5 pair. Image quality on par with GPT Image 2 at about half the latency, with the new xhigh and max quality presets, custom sizes up to 4K and tighter subject preservation on edits. OpenAI's default recommendation for everyday API image generation.
OpenAI
GPT-6 Astra served with reasoning mode set to pro, which spends a larger thinking budget per request. Same underlying model as GPT-6 Astra, same 1.05M context and 128K output ceiling, aimed at the harder end of coding, analysis and long multi-step work.
OpenAI
OpenAI's first GPT-6 model, out September 3 2026 from what the company calls its largest training run so far. It takes 1.05M tokens of context, writes up to 128K, and exposes five reasoning effort levels from low through max. OpenAI aims it at computer use, browsing, software engineering and long multi-step work.
OpenAI
GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series, suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, with capable reasoning for its price tier.
OpenAI
The middle tier of OpenAI's GPT-5.6 series, between Sol and Luna. For everyday coding, reasoning and agent work, at roughly half the cost of Sol.
OpenAI
The top tier of OpenAI's GPT-5.6 series, for complex reasoning, coding and agent workflows. Strongest at command-line and multi-step coding, and at long-horizon problems.
OpenAI
GPT-5.6 Luna Pro is the same underlying model as GPT-5.6 Luna, served with reasoning mode set to pro for higher-quality responses on complex tasks.
OpenAI
OpenAI's GPT-5.5, for complex professional work. Builds on GPT-5.4 with stronger reasoning, higher reliability and better token efficiency on hard tasks. 1M+ token context window (922K input, 128K output), text and image in.
OpenAI
OpenAI's second-generation image model. Near-perfect typography (including non-Latin scripts), 2x faster than gpt-image-1.5, up to 2K resolution, and 'thinking' over the prompt for dense compositions.
OpenAI
GPT-5.4 at a smaller size, for high-throughput work. Text and image input, with reasoning, coding and tool use intact at lower latency and cost. The intended fit is chat products, coding assistants and agent workflows running at scale.
OpenAI
The smallest and cheapest model in the GPT-5.4 family, for speed-critical, high-volume work. Text and image input. It trades reasoning depth for responsiveness, which suits classification, data extraction, ranking and sub-agent calls.
OpenAI
OpenAI's GPT-5.4, folding the Codex and GPT lines into one system. Over 1M tokens of context, 922K in and 128K out, with text and image input. The improvements land in coding, document understanding, tool use and instruction following.
OpenAI
The Pro tier of GPT-5.4, with more reasoning compute for high-stakes work. Same unified Codex and GPT architecture, over 1M tokens of context at 922K in and 128K out, text and image input. Aimed at agentic coding, long-context workflows and multi-step problems.
OpenAI
An update to the model most ChatGPT conversations run on. It answers with better context, refuses less often, and drops the caveats and hedging that used to interrupt a conversation. More willing to engage with nuanced topics, its prose reads cleaner, and reasoning, coding and creative work hold up.
OpenAI
OpenAI's agentic coding model at the 5.3 mark, pairing GPT-5.2-Codex software engineering with GPT-5.2's broader knowledge. Leads SWE-Bench Pro and scores well on Terminal-Bench 2.0 and OSWorld-Verified. Built for long tool-using runs, and steerable mid-execution.
OpenAI
OpenAI's latest image generation model with strong instruction following, optional transparent backgrounds, and quality controls.
OpenAI
GPT Image 1.5 with `quality=low` for faster and cheaper generations.
OpenAI
GPT Image 1.5 with `quality=medium` for balanced cost and quality.
OpenAI
GPT Image 1.5 with `quality=high` for maximum fidelity.
OpenAI
The fast, lightweight member of the GPT-5.2 family, built for low-latency chat. Adaptive reasoning lets it think only on the harder queries, so mathematics, coding and multi-step questions get more accurate without slowing ordinary conversation. Warmer and more conversational by default than the base model.
OpenAI
The GPT-5.2 base model, with better agentic and long-context behaviour than GPT-5.1. Adaptive reasoning allocates compute per query, answering simple things quickly and spending depth on hard ones. Gains show up across mathematics, coding, science and tool calling, with more coherent long-form answers.
OpenAI
The Pro tier of GPT-5.2, aimed at agentic coding and long-context work beyond what GPT-5 Pro handled. Test-time routing and intent cues like "think hard about this" carry over, with lower hallucination and sycophancy than the previous Pro.
OpenAI
The largest Codex tier in OpenAI's GPT-5.1 line, built for coding agents that run for hours rather than turns. Training data stops in September 2024.
OpenAI
The latest of the GPT-5 series: better general reasoning, tighter instruction adherence and a more natural conversational style than GPT-5. Allocates compute adaptively, answering simple queries fast and spending longer on hard ones.
OpenAI
GPT-5.1 Chat (Instant), the fast and light member of the 5.1 family, built for low-latency chat. It still thinks selectively on harder queries, which keeps math, coding and multi-step accuracy up without slowing normal conversation.
OpenAI
A smaller, faster GPT-5.1-Codex for coding at lower latency.
OpenAI
GPT-5.1 tuned for software engineering: interactive development sessions and long independent runs, covering projects from scratch, feature work, debugging, large refactors and code review.
OpenAI
The Pro tier of GPT-5, given more test-time compute for complex, high-stakes tasks. Same routing controls and intent cues as GPT-5, with lower hallucination and sycophancy and stronger coding, writing and health answers.
OpenAI
GPT-5 retuned for Codex, OpenAI's agentic coding tool. It runs on the Responses API only, not chat completions, and its training data stops in September 2024.
OpenAI
OpenAI's GPT-5, built around step-by-step reasoning and instruction following on high-stakes work. It routes at test time and reads intent cues like "think hard about this". Hallucination and sycophancy are down against GPT-4o, with better coding, writing and health answers.
OpenAI
A compact GPT-5 for lighter reasoning, with the same instruction following and safety tuning at lower latency and cost. It is the successor to OpenAI's o4-mini.
OpenAI
The smallest model in OpenAI's GPT-5 line, priced for high-volume classification, extraction and short code edits rather than hard reasoning. Text and image input, training data through May 2024.
OpenAI
OpenAI's smaller open-weight release, 21B parameters on a mixture-of-experts design. Matches o3-mini on quality while fitting in 16GB of memory, which puts it on consumer hardware and edge devices.
OpenAI
OpenAI's first open-weight model since GPT-2, 117B parameters on a mixture-of-experts design. Reasoning depth is configurable per request and tool use is native, aimed at agent and production workloads.
OpenAI
OpenAI's first generally available audio model. Features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Supports both text and audio input/output.
OpenAI
Cost-efficient version of GPT Audio. Features an upgraded decoder for more natural sounding voices and better voice consistency at a fraction of the cost.
OpenAI
codex-mini-latest is a fine-tuned version of o4-mini specifically for use in Codex CLI. For direct use in the API, we recommend starting with gpt-4.1.
OpenAI
OpenAI's reasoning model for coding, math, science and visual perception. Thinks longer before answering and reaches for tools on its own: web search, code execution, image generation. Set new SOTA on Codeforces and MMMU.
OpenAI
A smaller, cheaper OpenAI reasoning model built for speed. Strong for its size on math, coding and visual tasks, with much higher usage limits than o3, and it can use tools on its own.
OpenAI
The same model as o4-mini, defaulting to high reasoning effort. Compact, fast and cheap, keeping the multimodal and agentic behaviour, especially on math, coding and visual tasks.
OpenAI
OpenAI's GPT-4.1, built around a 1 million token context window and precise code diffs. Scores 54.6% on SWE-bench Verified and 87.4% on IFEval, and beats GPT-4o on coding, instruction compliance and long-document recall.
OpenAI
The smallest and cheapest model in the GPT-4.1 line, with the same 1 million token context. Scores 80.1% on MMLU, 50.3% on GPQA and 9.8% on Aider polyglot, above GPT-4o mini. Aimed at classification and autocompletion.
OpenAI
The mid tier of the GPT-4.1 line, roughly GPT-4o quality at lower latency and cost, with the same 1 million token context. Scores 84.1% on IFEval, 35.8% on MultiChallenge and 31.6% on Aider's polyglot diff benchmark, and handles vision.
OpenAI
The March 2025 refresh of GPT-4o in ChatGPT. Follows instructions more closely, handles coding with fewer detours, and writes shorter answers with fewer nested markdown levels than the previous build.
OpenAI
GPT-4o with audio input support. Detects nuances within audio recordings and adds depth to generated user experiences. Processes both text and audio prompts.
OpenAI
GPT-4.5 scaled up pre-training and post-training. Broader knowledge, better intent reading and more 'EQ' on conversation, writing, programming and practical problems, with fewer hallucinations. Scored 62.5% on SimpleQA with a 37.1% hallucination rate, ahead of GPT-4o.
OpenAI
A smaller, cheaper o3 with faster response times. OpenAI says output quality holds.
OpenAI
o1 achieves 86% accuracy on Mathematics Olympiad benchmarks (vs. GPT-4o's 13%), offers PhD-level STEM proficiency, and is built with extensive alignment training and safety evaluation.
OpenAI
OpenAI's small model after GPT-4 Omni: text and image in, text out. Many multiples cheaper than the frontier models of its day and over 60% cheaper than GPT-3.5 Turbo, at close to SOTA intelligence.
OpenAI
GPT-4o processes text, images, and audio through a unified transformer architecture and offers real-time translation for 154 languages with 89.2% BLEU score on low-resource languages.
OpenAI
OpenAI's high-definition text-to-speech with six built-in voices. Clear and natural, at the cost of latency.
OpenAI
OpenAI's standard text-to-speech, tuned for low latency. Six voices, several output formats. Faster than TTS-1-HD and slightly rougher.
OpenAI
DALL-E 3 auto-improves user inputs via ChatGPT integration and blocks prohibited content with 99.9% precision using multimodal classifiers.
OpenAI
The original GPT-4: a large multimodal model that handled hard problems more accurately than GPT-3.5 on the strength of broader general knowledge and better reasoning. Training data stops at September 2021.
OpenAI
OpenAI's fastest model of its generation. Understands and writes natural language and code, tuned for chat and completion.
Google AI
Google's most capable Flash model at release, with gains over 3.7 Flash in software engineering, agent tasks and multi-step reasoning. It accepts text, images, audio, video and files across a 1M-token window, with up to 65,536 output tokens.
Google AI
Google's Gemini 3.7 Flash, for fast agent workflows, coding and multi-step reasoning. Text, images, audio, video and files across a 1M-token context.
Google AI
Google's Flash workhorse at the 3.6 mark, for coding, agent loops and web and app development. Built on 3.5 Flash, aiming at cleaner output with fewer unnecessary edits and less hedging, on roughly 17% fewer output tokens. Text, images, audio, video and PDFs across 1M tokens, up to 64K out.
Google AI
Gemini 3.1 Flash Lite Image, the cheapest and fastest of Google's Nano Banana line. Text-to-image in roughly 4 seconds, about 2.7x faster than Gemini 3.1 Flash Image, with the same character consistency. One API for text-to-image, editing and composition. 1K resolution, 14 aspect ratios, SynthID watermark.
Google AI
Google's Flash tier at the 3.5 mark, close to Pro on coding and reasoning at Flash cost and speed. Tuned for coding and parallel agent loops, taking text, image, video, audio and PDF input. Thinking effort defaults to medium and runs from minimal to high.
Google AI
The GA release of Gemini 3.1 Flash Lite, Google's cheapest multimodal tier. Text, image, video, audio and PDF in, for light agent work, simple extraction, and anything where latency and cost decide the design. Thinking runs minimal to high. Half the price of Gemini 3 Flash.
Google AI
An instruction-tuned mixture-of-experts from Google DeepMind: 25.2B total parameters, 3.8B active per token, which is how it reaches near-31B quality on a fraction of the compute. Text, images and video up to 60 seconds at 1fps. 256K context, native function calling, thinking mode. Apache 2.0.
Google AI
Google DeepMind's 30.7B dense multimodal model: text and image in, text out. 256K context, configurable thinking mode, native function calling, 140+ languages. Strongest on coding, reasoning and document understanding. Apache 2.0.
Google AI
Google's high-efficiency tier at the 3.1 mark. Beats Gemini 2.5 Flash Lite on quality and gets close to 2.5 Flash, with the gains in audio and speech recognition, RAG snippet ranking, translation, extraction and code completion. 2.5x faster to first token than 2.5 Flash, at half the price of Gemini 3 Flash.
Google AI
Gemini 3.1 Flash Image Preview, sold as Nano Banana 2. Google's image generation and editing model at Flash speed and Flash pricing, aimed at iterative edits rather than one-shot renders.
Google AI
The 3.1 update to Gemini Pro: better software engineering scores, more reliable agent behaviour, fewer tokens spent getting there. Reasons across text, image, video, audio and code over a 1M token context. Adds a medium thinking level, and improves on structured work like finance and spreadsheets.
Google AI
Google DeepMind's Lyria 3, launched in the Gemini app. First Lyria with full vocal generation, automatic lyrics and multimodal input (text, image, video). Produces 30-second tracks with distinct verses, choruses and bridges, in 8 languages. No public API: it runs in the Gemini app and YouTube Dream Track.
Google AI
Gemini 3's Flash tier, for agent loops, multi-turn chat and coding assistance at much lower latency than Pro. 1M token context. Text, image, audio, video and PDF in, text out. Thinking level runs minimal to high, with structured output, tool use and automatic context caching.
Google AI
Google DeepMind's production music generation model available via Vertex AI. Generates 32.8-second instrumental clips at 48kHz from text prompts. Supports negative prompts to exclude unwanted elements and deterministic generation via seed parameter. All output is SynthID-watermarked.
Google AI
Gemini 3 Pro Preview with high reasoning effort enabled. Exposes full chain-of-thought process for enhanced transparency in complex problem-solving across text, code, and multimodal tasks.
Google AI
Google's Gemini 3 Pro Image model, sold as Nano Banana Pro. Generates and edits at up to 4K and routes prompts through Gemini's thinking mode for text rendering and multi-step instructions.
Google AI
The September 2025 checkpoint of Gemini 2.5 Flash, Google's workhorse tier for reasoning, coding, mathematics and science. Thinking is built in, and the max tokens for reasoning parameter caps how long it deliberates.
Google AI
The September 2025 checkpoint of Gemini 2.5 Flash-Lite, the cheapest and lowest latency tier in the family. Thinking is off by default; turning it on through the Reasoning API parameter trades cost for intelligence.
Google AI
Google's Nano Banana (Gemini 2.5 Flash Image) generates and edits images, tuned for speed. Tested on LMArena as 'nano-banana', where it beat Midjourney and Flux in blind tests.
Google AI
Gemini 2.5 Flash variant tuned for fast text-to-image generation and simple image edits.
Google AI
Google DeepMind's Gemma 3n at an effective 2B parameters inside a 6B architecture. The MatFormer design lets it nest submodels and recombine them through Mix-and-Match, which is how it runs on low-resource hardware. 32K context, multilingual.
Google AI
The lightweight tier of Gemini 2.5, tuned for low latency and cost rather than depth. Thinking is off by default to keep it fast; developers can switch it on through the Reasoning API parameter and pay for the extra intelligence.
Google AI
Imagen 4 Fast trades some quality for speed and cost.
Google AI
Google's Gemini 2.5 Pro, a reasoning model that thinks through a request before answering. Aimed at coding, mathematics and scientific work, and it held first place on the LMArena leaderboard at release.
Google AI
Imagen 4 Ultra trades speed and cost for image quality.
Google AI
Google's Imagen 4 text-to-image model, with strong detail and prompt adherence.
Google AI
The May 20 checkpoint of Gemini 2.5 Flash in its non-thinking form: it will not emit reasoning tokens, and is billed at the lower output rate. The paired :thinking variant is the one that reasons before answering.
Google AI
The May 20 checkpoint of Gemini 2.5 Flash with thinking enabled. It reasons before answering, is billed at the higher output rate, and takes a max tokens for reasoning parameter to cap how long it deliberates.
Google AI
Google's Gemma 3n at the E4B size, built to run on phones, laptops and tablets. Text, image and audio in, for transcription, translation and image analysis. Per-Layer Embedding caching and the MatFormer architecture activate only part of the model per request, which keeps memory use low. 32K context, 140+ languages.
Google AI
Google's Gemini 2.5 Pro preview: multimodal understanding, multi-step code, math and STEM, and long-context analysis over large datasets, codebases and documents. Knowledge cutoff January 2025.
Google AI
Google's Flash tier workhorse, with reasoning that can be switched on or off and a configurable thinking budget. The point is trading quality against cost and latency per request rather than per model.
Google AI
Gemini 2.5 Flash with thinking left on, so it reasons through a request before answering. The same workhorse model as the non-thinking variant, priced higher on output because thinking tokens are billed.
Google AI
Google's experimental Gemini 2.5 Pro: multimodal reasoning, a 1 million token context window, and video comprehension.
Google AI
Gemma 3 adds multimodality: vision-language in, text out. 96,000 token context, 140+ languages, better math, reasoning and chat, with structured outputs and function calling. The 12B is the second largest, after the 27B.
Google AI
Gemma 3 adds multimodality: vision-language in, text out. 131,072 token context, 140+ languages, better math, reasoning and chat, with structured outputs and function calling. The 27B is Google's largest Gemma 3, successor to Gemma 2.
Google AI
A faster and cheaper Imagen 3 variant for when speed and cost matter more than maximum quality.
Google AI
Google's high-quality text-to-image model focused on lighting, detail, and strong visual composition.
Google AI
Gemini 2.0 Pro builds interactive 3D environments from text descriptions and offers hypothetical reasoning for scientific simulations.
Google AI
An experimental Gemini 2.0 Flash variant that reasons through a problem before answering. Trades some latency for accuracy on multi-step questions while staying at Flash speed.
Google AI
Gemini 1.5 Pro handles infinite context with 99% retrieval accuracy at 750k tokens via Mixture-of-Experts and generates chapter summaries for 2-hour videos with 92% accuracy.
Anthropic
Claude Fable 5.1 is a direct update to Fable 5, with its largest gains in agent coding, long-running workflows and knowledge work. Anthropic calls out long refactors, front-end generation, finance and analysis, with shorter plans and summaries than Fable 5.
Anthropic
Anthropic's largest model, for demanding reasoning, coding and long-horizon agent work. Strongest on end-to-end software tasks, code review and bug finding, visual analysis of charts and documents, and coordinating parallel subagents. Instruction following holds across long tasks, even at lower effort.
Anthropic
Anthropic's most capable Sonnet, with adaptive thinking across selectable effort levels from low to x-high, a 1M token context, and text, image and file input. It ships an updated tokenizer and real-time cyber safeguards that block certain high-risk dual-use requests.
Anthropic
Anthropic's first public Mythos-class model, for autonomous knowledge work and coding. Text, image and file input, 1M token context. Built for long, ambiguous, multi-step tasks that used to need check-ins, self-correcting through verification loops. Falls back to Claude Opus 4.8 on high-risk queries.
Anthropic
Anthropic's most capable generally available Opus. Text, image and file input, 1M token context. Built for autonomous agents and memory-driven work where coherence across a long session is the hard part: multi-step reasoning, complex coding, orchestration across large codebases, multi-stage debugging.
Anthropic
Anthropic's Opus at the 4.7 mark, for long-running asynchronous agents. Stronger than Opus 4.6 on multi-step work: large codebases, multi-stage debugging, end-to-end orchestration. Beyond coding it drafts documents, builds presentations and analyses data, holding coherence across very long sessions.
Anthropic
Anthropic's Sonnet 4.6, for coding, agents and professional work: iterative development, navigating large codebases, end-to-end project management with memory, document creation, and computer use for web QA and workflow automation.
Anthropic
Anthropic's strongest model for coding and long-running professional work, built for agents that operate across a whole workflow rather than a single prompt. Handles large codebases, complex refactors and multi-step debugging, and produces near-final documents, plans and analyses in one pass.
Anthropic
Anthropic's Opus tier at the 4.5 mark, aimed at long coding sessions and agent work. Text and image input, training data through August 2025, available on the Claude API, Bedrock and Vertex.
Anthropic
Claude Haiku 4.5 matches Claude Sonnet 4's reasoning and coding strength while staying ultra-responsive for real-time and high-volume workloads. Extended thinking, controllable reasoning depth, and full tool-use support keep sub-agents fast without sacrificing quality.
Anthropic
Anthropic's Sonnet 4.5, tuned for agents and coding workflows: speculative parallel execution, memory handling and tool orchestration.
Anthropic
An update to Anthropic's Opus tier, sharpened on coding, reasoning and agent work. Scores 74.5% on SWE-bench Verified, with the clearest gains in multi-file refactors and debugging precision. Extended thinking runs to 64K tokens.
Anthropic
The top tier of Anthropic's Claude 4 line, for coding, reasoning and agents. Built for long-running tasks, with extended thinking during tool use and better memory.
Anthropic
Claude Sonnet 4 is a significant upgrade to Sonnet 3.7, delivering superior coding and reasoning while responding more precisely to instructions. It balances performance and efficiency for various use cases.
Anthropic
Claude 3.7 Thinking Sonnet exposes the full chain-of-thought process during problem-solving, including error backtracking and alternative solution exploration. Scores 86.1% on GPQA Diamond benchmark for expert-level Q&A.
Anthropic
Claude 3.7 Sonnet offers Extended Thinking Scaffolds that boost SWE-bench coding accuracy from 62.3% to 70.3%, with 81.2% accuracy in retail automation tasks, outperforming Claude Sonnet 3.6 (2022-10-22) by 13.6%.
Anthropic
Claude 3.5 Sonnet offers a cost-efficient API ($3/million input tokens vs. $5 for GPT-4o) and uses embedded alignment techniques that reduce harmful outputs by 34% compared to Claude 2.1.
Anthropic
A research demo of Claude 3 Sonnet, live for 24 hours on 23 May 2024, engineered by Anthropic to demonstrate feature steering. It was manipulated to obsess over the Golden Gate Bridge in every response.
Anthropic
The middle tier of Claude 3: a 200,000 token context window, priced for scaled deployment, and steady on conversation, reasoning, code generation and analysis.
Anthropic
Claude 3 Haiku is Anthropic's fastest and most compact model, designed for near-instant responses in real-time applications with strong multilingual support.
Anthropic
The top tier of Anthropic's Claude 3 line, for complex reasoning and open-ended problem solving.
Anthropic
Anthropic's Claude 2 model, featuring a large 100K token context window and strong performance on various benchmarks. Known for helpful, honest, and harmless AI conversations.
338+ models you haven't met. $0 to continue. Seriously.
Free account. No card required. By continuing, you agree to Rival's Terms and Privacy Policy