QwQ 32B is cheaper than Llama 3.1 405B at $0.5/M vs $2.7/M input tokens.
| Spec | Llama 3.1 405B | QwQ 32B |
|---|---|---|
| Input price | $2.7/M tokens | $0.5/M tokens |
| Output price | $3.1/M tokens | $1.5/M tokens |
| Context window | 128K tokens | 40K tokens |
| Released | Jul 2024 | Mar 2025 |
Not enough votes to call it. On the specs, Llama 3.1 405B has the edge: bigger model tier, bigger context window, major provider backing.
QwQ 32B is cheaper on both — 5.4× input, 2.1× output
QwQ 32B uses 1360.7x more bold
QwQ 32B is cheaper than Llama 3.1 405B at $0.5/M vs $2.7/M input tokens.
| Spec | Llama 3.1 405B | QwQ 32B |
|---|---|---|
| Input price | $2.7/M tokens | $0.5/M tokens |
| Output price | $3.1/M tokens | $1.5/M tokens |
| Context window | 128K tokens | 40K tokens |
| Released | Jul 2024 | Mar 2025 |
Not enough votes to call it. On the specs, Llama 3.1 405B has the edge: bigger model tier, bigger context window, major provider backing.
QwQ 32B is cheaper on both — 5.4× input, 2.1× output
QwQ 32B uses 1360.7x more bold