Gemini 2.0 Flash Thinking is cheaper than Llama 3.1 405B at $0.25/M vs $2.7/M input tokens.
| Spec | Gemini 2.0 Flash Thinking | Llama 3.1 405B |
|---|---|---|
| Input price | $0.25/M tokens | $2.7/M tokens |
| Output price | $0.5/M tokens | $3.1/M tokens |
| Context window | 500K tokens | 128K tokens |
| Released | Dec 2024 | Jul 2024 |
| MMLU | 82.3% | 88.6% |
Not enough votes to call it. On the specs, nothing separates them.
Gemini 2.0 Flash Thinking costs 6.2x less per token.
Gemini 2.0 Flash Thinking is cheaper on both — 11× input, 6.2× output
Gemini 2.0 Flash Thinking uses 376.7x more bold
Gemini 2.0 Flash Thinking is cheaper than Llama 3.1 405B at $0.25/M vs $2.7/M input tokens.
| Spec | Gemini 2.0 Flash Thinking | Llama 3.1 405B |
|---|---|---|
| Input price | $0.25/M tokens | $2.7/M tokens |
| Output price | $0.5/M tokens | $3.1/M tokens |
| Context window | 500K tokens | 128K tokens |
| Released | Dec 2024 | Jul 2024 |
| MMLU | 82.3% | 88.6% |
Not enough votes to call it. On the specs, nothing separates them.
Gemini 2.0 Flash Thinking costs 6.2x less per token.
Gemini 2.0 Flash Thinking is cheaper on both — 11× input, 6.2× output
Gemini 2.0 Flash Thinking uses 376.7x more bold