Complete API price reference for every major AI model: OpenAI GPT-4o and o-series, Anthropic Claude, Google Gemini, Mistral, Meta Llama, xAI Grok, DeepSeek, MiniMax and Moonshot Kimi. All prices are USD per million tokens, with context windows, prompt-cache rates and batch discounts included.
Prices verified August 20, 2026| Model | Input / M | Output / M | Cache read | Context | Batch |
|---|---|---|---|---|---|
| GPT-4o vision audio gpt-4o | $2.50 / M | $10 / M | $1.25 / M | 128K | −50% |
| GPT-4o mini vision audio gpt-4o-mini | $0.15 / M | $0.60 / M | $0.07 / M | 128K | −50% |
| GPT-4 Turbo vision gpt-4-turbo | $10 / M | $30 / M | $5 / M | 128K | −50% |
| GPT-3.5 Turbo gpt-3.5-turbo | $0.50 / M | $1.50 / M | — | 16K | −50% |
| o1 vision reasoning o1 | $15 / M | $60 / M | $7.50 / M | 200K | −50% |
| o1-mini reasoning o1-mini | $1.10 / M | $4.40 / M | $0.55 / M | 128K | −50% |
| o3 vision reasoning o3 | $2 / M | $8 / M | $1 / M | 200K | −50% |
| o3-mini reasoning o3-mini | $1.10 / M | $4.40 / M | $0.55 / M | 200K | −50% |
| GPT-4.1 vision gpt-4-1 | $2 / M | $8 / M | $1 / M | 1048K | −50% |
| GPT-4.1 mini vision gpt-4-1-mini | $0.40 / M | $1.60 / M | $0.20 / M | 1048K | −50% |
| GPT-5.5 vision reasoning audio gpt-5-5 | $5 / M | $30 / M | $0.50 / M | 1050K | −50% |
| GPT-5.6 Sol vision reasoning audio gpt-5-6-sol | $2.50 / M | $15 / M | $0.50 / M | 1050K | −50% |
| GPT-5.6 Terra vision reasoning audio gpt-5-6-terra | $2 / M | $12 / M | $0.20 / M | 1050K | −50% |
| GPT-5.6 Luna vision reasoning audio gpt-5-6-luna | $0.20 / M | $1.20 / M | $0.02 / M | 1050K | −50% |
| Model | Input / M | Output / M | Cache read | Context | Batch |
|---|---|---|---|---|---|
| Claude 3.5 Sonnet vision claude-3-5-sonnet | $3 / M | $15 / M | $0.30 / M | 200K | — |
| Claude 3.5 Haiku vision claude-3-5-haiku | $0.80 / M | $4 / M | $0.08 / M | 200K | — |
| Claude 3 Opus vision claude-3-opus | $15 / M | $75 / M | $1.50 / M | 200K | — |
| Claude 3 Sonnet vision claude-3-sonnet | $3 / M | $15 / M | $0.30 / M | 200K | — |
| Claude Sonnet 4.6 vision claude-sonnet-4-6 | $3 / M | $15 / M | $0.30 / M | 1000K | — |
| Claude Haiku 4.5 vision claude-haiku-4-5 | $1 / M | $5 / M | $0.10 / M | 200K | — |
| Claude Sonnet 5 vision claude-sonnet-5 | $2 / M | $10 / M | $0.20 / M | 1000K | — |
| Claude Opus 5 vision reasoning claude-opus-5 | $5 / M | $25 / M | $0.50 / M | 1000K | — |
| Claude Fable 5 vision reasoning claude-fable-5 | $10 / M | $50 / M | $1 / M | 1000K | — |
| Model | Input / M | Output / M | Cache read | Context | Batch |
|---|---|---|---|---|---|
| Gemini 1.5 Pro vision audio gemini-1-5-pro | $1.25 / M | $5 / M | — | 2000K | — |
| Gemini 1.5 Flash vision audio gemini-1-5-flash | $0.07 / M | $0.30 / M | — | 1000K | — |
| Gemini 2.0 Flash vision audio gemini-2-0-flash | $0.10 / M | $0.40 / M | — | 1000K | — |
| Gemini Ultra vision reasoning audio gemini-ultra | $7.50 / M | $30 / M | — | 1000K | — |
| Gemini 2.5 Flash vision reasoning audio gemini-2-5-flash | $0.30 / M | $2.50 / M | — | 1049K | −50% |
| Gemini 3.5 Flash vision reasoning audio gemini-3-5-flash | $1.50 / M | $9 / M | — | 1049K | −50% |
| Gemini 3.1 Pro vision reasoning audio gemini-3-1-pro | $2 / M | $12 / M | — | 1049K | −50% |
| Gemini 3.6 Flash vision reasoning audio gemini-3-6-flash | $0.75 / M | $3.75 / M | — | 1049K | −50% |
| Gemini 3.7 Flash vision reasoning audio gemini-3-7-flash | $0.38 / M | $1.88 / M | — | 1049K | −50% |
| Model | Input / M | Output / M | Cache read | Context | Batch |
|---|---|---|---|---|---|
| Mistral Large vision mistral-large | $2 / M | $6 / M | — | 128K | — |
| Mistral Small vision mistral-small | $0.15 / M | $0.60 / M | — | 262K | — |
| Mixtral 8x7B mixtral-8x7b | $0.70 / M | $0.70 / M | — | 32K | — |
| Mistral Large 3 vision reasoning mistral-large-3 | $0.50 / M | $1.50 / M | — | 262K | −50% |
| Mistral Small 4 vision mistral-small-4 | $0.15 / M | $0.60 / M | — | 262K | −50% |
| Model | Input / M | Output / M | Cache read | Context | Batch |
|---|---|---|---|---|---|
| Llama 3.1 405B llama-3-1-405b | $3.50 / M | $3.50 / M | — | 128K | — |
| Llama 3.1 70B llama-3-1-70b | $0.40 / M | $0.40 / M | — | 131K | — |
| Llama 3.1 8B llama-3-1-8b | $0.05 / M | $0.08 / M | — | 131K | — |
| Model | Input / M | Output / M | Cache read | Context | Batch |
|---|---|---|---|---|---|
| Grok 4.6 vision reasoning grok-4-6 | $2 / M | $6 / M | $0.50 / M | 500K | — |
| Grok 4.3 vision reasoning grok-4-3 | $1.25 / M | $2.50 / M | $0.20 / M | 1000K | — |
| Grok Build vision grok-build | $1 / M | $2 / M | $0.20 / M | 256K | — |
| Model | Input / M | Output / M | Cache read | Context | Batch |
|---|---|---|---|---|---|
| MiniMax M3 vision minimax-m3 | $0.30 / M | $1.20 / M | $0.06 / M | 1049K | — |
| MiniMax M2.7 vision minimax-m2-7 | $0.30 / M | $1.20 / M | — | 205K | — |
| MiniMax M2 vision minimax-m2 | $0.26 / M | $1.02 / M | — | 205K | — |
| Model | Input / M | Output / M | Cache read | Context | Batch |
|---|---|---|---|---|---|
| DeepSeek V4 Flash vision reasoning deepseek-v4-flash | $0.14 / M | $0.28 / M | $0.00 / M | 1311K | −50% |
| DeepSeek V4 Pro vision reasoning deepseek-v4-pro | $1.19 / M | $3.56 / M | $0.00 / M | 1049K | −50% |
| Matchup | Focus |
|---|---|
| GPT-4o vs Claude Sonnet 4.6 | The default flagship showdown. |
| GPT-4o vs Gemini 2.0 Flash | Speed-vs-depth trade. |
| GPT-4o vs o3 | Generalist vs reasoner. |
| GPT-4o mini vs Gemini 2.5 Flash | The budget workhorses. |
| Claude Sonnet 4.6 vs Gemini 3.1 Pro | Long-context specialists. |
| DeepSeek V4 Pro vs GPT-4o | The value play. |
| Claude Opus 5 vs GPT-4o | Premium intelligence vs balanced cost. |
| Grok 4.6 vs GPT-4o | xAI’s challenger against the incumbent. |
| Llama 3.1 405B vs GPT-4o | Open weights vs closed API. |
| Kimi K3 vs Claude Sonnet 4.6 | Moonshot’s flagship against Claude’s workhorse. |
As of August 2026, GPT-4o mini ($0.15 input / $0.60 output per million tokens), Gemini Flash-class models and DeepSeek are the cheapest frontier-quality options. Open-weight models served through providers like Together AI often undercut even those prices.
GPT-4o costs $2.50 per million input tokens and $10.00 per million output tokens. Cached input costs $1.25/M and batch mode halves both rates to $1.25 / $5.00.
Anthropic Claude pricing spans from Haiku-class models (from $0.80/M input) to Opus-class flagship models ($15/M input, $75/M output). Prompt caching writes cost more than base input; cache reads cost roughly 10% of base input.
LLM APIs bill per token — a token is roughly 4 characters or ¾ of an English word. Prices in this table are USD per 1,000,000 tokens, so $2.50/M means about $0.0000025 per token.
Yes. Every major provider discounts cached prompt tokens heavily — typically 50–90% off base input price. If your app sends the same long system prompt repeatedly, caching is the single biggest cost saver.
Prices were last verified against each provider's official pricing page on August 20, 2026. Providers change prices frequently — always confirm on the official page before committing to a budget. AITokenCalculator is an estimator, not a billing source.