Every provider prices its models per token, but each in its own table. This page puts the public list prices side by side: input and output price per 1 million tokens, cached input, context window and what each model can do. Enter your own usage in the calculator to see the monthly cost per model.

Cheapest model

Nemotron 3 Nano 30B

NVIDIA
$0.0875 per 1M tokens, blended 3:1

Cheapest with vision

Gemma 4 26B A4B IT

Google Gemini
$0.1425 per 1M tokens, blended 3:1

Cheapest reasoning model

Nemotron 3 Nano 30B

NVIDIA
$0.0875 per 1M tokens, blended 3:1

Largest context window

GPT 5.4

OpenAI
1.1M tokens of context

Cost calculator

Enter your average request and monthly volume. The last column shows what each model would cost per month.

Presets:
Model Provider Input
per 1M tokens
Output
per 1M tokens
Cached input
per 1M tokens
Context Capabilities Per month
your scenario
Nemotron 3 Nano 30B
nvidia/nvidia/nemotron-3-nano-30b-a3b
NVIDIA $0.05 $0.20 $0.025 262K Tools Reasoning $1.30
DeepSeek V4 Flash
deepseek/deepseek-v4-flash
DeepSeek $0.088606 $0.177212 $0.017721 1M Tools Reasoning $1.59
GPT OSS 20B
groq/openai/gpt-oss-20b
Groq $0.075 $0.30 $0.0375 131K Tools Reasoning $1.95
Gemma 4 26B A4B IT
google/gemma-4-26b-a4b-it
Google Gemini $0.09 $0.30 $0.05 262K Vision Tools Reasoning $2.10
Gemma 4 31B IT
google/gemma-4-31b-it
Google Gemini $0.09 $0.34 $0.05 262K Vision Tools Reasoning $2.26
Nemotron 3 Super 120B
nvidia/nvidia/nemotron-3-super-120b-a12b
NVIDIA $0.08 $0.45 $0.06 262K Tools Reasoning $2.60
GPT OSS 120B
groq/openai/gpt-oss-120b
Groq $0.15 $0.60 $0.075 131K Tools Reasoning $3.90
Magistral Small
mistral/magistral-small-latest
Mistral $0.15 $0.60 $0.015 262K Vision Tools Reasoning $3.90
Mistral Small
mistral/mistral-small-latest
Mistral $0.15 $0.60 $0.015 262K Vision Tools Reasoning $3.90
GLM 4.5 Air
zai/glm-4.5-air
Z.AI $0.13 $0.85 $0.025 131K Tools Reasoning $4.70
deepseek-v4-flash-vision-exp
deepseek/deepseek-v4-flash-vision-exp
DeepSeek $0.22 $0.66 $0.007 1M Vision Tools Reasoning $4.84
Llama 3.3 Nemotron Super 49B
nvidia/nvidia/llama-3.3-nemotron-super-49b-v1.5
NVIDIA $0.40 $0.40 131K Tools Reasoning $5.60
MiniMax-M2
minimax/MiniMax-M2
MiniMax $0.255 $1.02 $0.03 205K Tools Reasoning $6.63
GPT 5.4 Nano
openai/gpt-5.4-nano
OpenAI $0.20 $1.25 $0.02 400K Vision Tools Reasoning $7.00
MiniMax-M2.5
minimax/MiniMax-M2.5
MiniMax $0.27 $1.08 $0.027 205K Tools Reasoning $7.02
MiniMax-M2.1
minimax/MiniMax-M2.1
MiniMax $0.30 $1.20 $0.03 205K Tools Reasoning $7.80
MiniMax-M2.7
minimax/MiniMax-M2.7
MiniMax $0.30 $1.20 $0.06 205K Tools Reasoning $7.80
MiniMax-M3
minimax/MiniMax-M3
MiniMax $0.30 $1.20 $0.06 1M Vision Tools Reasoning $7.80
GPT 5.5
yunwu/gpt-5.5
Yunwu $0.30 $1.50 Tools $9.00
Moonshot V1 8K
moonshot/moonshot-v1-8k
Moonshot $0.20 $2.00 8K Tools $10.00
GPT 4.1 Mini
openai/gpt-4.1-mini
OpenAI $0.40 $1.60 $0.10 1M Vision Tools $10.40
GLM 4.7
zai/glm-4.7
Z.AI $0.40 $1.75 $0.08 205K Tools Reasoning $11.00
GPT 3.5 Turbo
openai/gpt-3.5-turbo
OpenAI $0.50 $1.50 16K Tools $11.00
Mistral Large
mistral/mistral-large-latest
Mistral $0.50 $1.50 $0.05 262K Vision Tools $11.00
mistral-large-2512
mistral/mistral-large-2512
Mistral $0.50 $1.50 $0.05 262K Vision Tools $11.00
GLM 4.6
zai/glm-4.6
Z.AI $0.43 $1.75 $0.08 205K Tools Reasoning $11.30
Devstral
mistral/devstral-latest
Mistral $0.40 $2.00 256K Tools $12.00
Devstral Medium
mistral/devstral-medium-latest
Mistral $0.40 $2.00 256K Tools $12.00
Gemini Flash-Lite Latest
google/gemini-flash-lite-latest
Google Gemini $0.30 $2.50 $0.03 1M Vision Tools Reasoning $13.00
GLM 5
zai/glm-5
Z.AI $0.60 $1.92 $0.12 205K Tools Reasoning $13.68
Perplexity Sonar
sonar
Perplexity $1.00 $1.00 127K $14.00
GLM 4.5
zai/glm-4.5
Z.AI $0.60 $2.20 $0.11 131K Tools Reasoning $14.80
Grok Build 0.1
xai/grok-build-0.1
xAI $1.00 $2.00 $0.20 256K Vision Tools Reasoning $18.00
Kimi 2.5
moonshot/kimi-k2.5
Moonshot $0.60 $3.00 $0.10 262K Vision Tools Reasoning $18.00
Qwen/Qwen3.6-27B
groq/qwen/qwen3.6-27b
Groq $0.60 $3.00 131K Vision Tools Reasoning $18.00
Nemotron 3 Ultra 550B
nvidia/nvidia/nemotron-3-ultra-550b-a55b
NVIDIA $0.625 $3.13 $0.1875 262K Tools Reasoning $18.75
GLM 5.1
zai/glm-5.1
Z.AI $0.966 $3.04 $0.1794 205K Tools Reasoning $21.80
Moonshot V1 32K
moonshot/moonshot-v1-32k
Moonshot $1.00 $3.00 33K Tools $22.00
Gemini Flash Latest
google/gemini-flash-latest
Google Gemini $0.75 $3.75 $0.075 1M Vision Tools Reasoning $22.50
Grok 4.3
xai/grok-4.3
xAI $1.25 $2.50 $0.20 1M Vision Tools Reasoning $22.50
GPT 5.4 Mini
openai/gpt-5.4-mini
OpenAI $0.75 $4.50 $0.075 400K Vision Tools Reasoning $25.50
Kimi 2.6
moonshot/kimi-k2.6
Moonshot $0.95 $4.00 $0.16 262K Vision Tools Reasoning $25.50
GLM 5 Turbo
zai/glm-5-turbo
Z.AI $1.20 $4.00 $0.24 203K Tools Reasoning $28.00
DeepSeek V4 Pro
deepseek/deepseek-v4-pro
DeepSeek $1.60 $3.20 $0.135 1M Tools Reasoning $28.80
Claude Haiku 4.5
anthropic/claude-haiku-4-5-20251001
Anthropic $1.00 $5.00 $0.10 200K Vision Tools Reasoning $30.00
GLM 5.2
zai/glm-5.2
Z.AI $1.40 $4.40 $0.14 1M Tools Reasoning $31.60
Moonshot V1 128k
moonshot/moonshot-v1-128k
Moonshot $2.00 $5.00 131K Tools $40.00
Moonshot V1 Auto
moonshot/moonshot-v1-auto
Moonshot $2.00 $5.00 131K Tools $40.00
Mistral Medium
mistral/mistral-medium-latest
Mistral $1.50 $7.50 $0.15 262K Vision Tools Reasoning $45.00
Perplexity Sonar Deep Research
sonar-deep-research
Perplexity $2.00 $8.00 127K $52.00
Perplexity Sonar Reasoning Pro
sonar-reasoning-pro
Perplexity $2.00 $8.00 127K $52.00
GPT 5.1
openai/gpt-5.1
OpenAI $1.25 $10.00 $0.125 400K Vision Tools Reasoning $52.50
Claude Sonnet 5
anthropic/claude-sonnet-5
Anthropic $2.00 $10.00 $0.20 1M Vision Tools Reasoning $60.00
Nano Banana Pro
google/gemini-3-pro-image
Google Gemini $2.00 $12.00 $0.20 131K Vision Tools Reasoning $68.00
Claude Opus 4.8
yunwu/claude-opus-4-8
Yunwu $2.50 $12.00 $0.30 Vision Tools $73.00
GPT 5.2
openai/gpt-5.2
OpenAI $1.75 $14.00 $0.175 400K Vision Tools Reasoning $73.50
GPT 5.4
openai/gpt-5.4
OpenAI $2.50 $15.00 $0.25 1.1M Vision Tools Reasoning $85.00
Claude Sonnet 4.6
anthropic/claude-sonnet-4-6
Anthropic $3.00 $15.00 $0.30 1M Vision Tools Reasoning $90.00
claude-sonnet-5
yunwu/claude-sonnet-5
Yunwu $3.00 $15.00 $0.30 Vision $90.00
Kimi 3
moonshot/kimi-k3
Moonshot $3.00 $15.00 $0.30 1M Vision Tools Reasoning $90.00
Perplexity Sonar Pro
sonar-pro
Perplexity $3.00 $15.00 127K $90.00
Claude Fable 5
yunwu/claude-fable-5
Yunwu $5.00 $25.00 $1.00 Tools Reasoning $150.00
Claude Opus 4.6
anthropic/claude-opus-4-6
Anthropic $5.00 $25.00 $0.50 1M Vision Tools Reasoning $150.00
Claude Opus 4.7
anthropic/claude-opus-4-7
Anthropic $5.00 $25.00 $0.50 1M Vision Tools Reasoning $150.00
Claude Opus 4.8
anthropic/claude-opus-4-8
Anthropic $5.00 $25.00 $0.50 1M Vision Tools Reasoning $150.00
Claude Opus 5
anthropic/claude-opus-5
Anthropic $5.00 $25.00 $0.50 1M Vision Tools Reasoning $150.00
GPT 5.5
openai/gpt-5.5
OpenAI $5.00 $30.00 $0.50 1.1M Vision Tools Reasoning $170.00
GPT Image 2
openai/gpt-image-2
OpenAI $5.00 $30.00 $1.25 Vision $170.00
GPT Image 1.5
openai/gpt-image-1.5
OpenAI $5.00 $32.00 $1.25 Vision $178.00
GPT Image 1
openai/gpt-image-1
OpenAI $5.00 $40.00 $1.25 $210.00
Claude Fable 5
anthropic/claude-fable-5
Anthropic $10.00 $50.00 $1.00 1M Vision Tools Reasoning $300.00
GPT 5.5 Pro
openai/gpt-5.5-pro
OpenAI $30.00 $180.00 $3.00 1.1M Vision Tools Reasoning $1,020.00

No models match your search.

List prices in USD per 1 million tokens as published by the providers, without volume or batch discounts. Prices as of Thursday, September 17, 2026.

Spotted an outdated price? Let us know.

How LLM API pricing works

Four things to know before comparing numbers.

  1. 01

    Tokens, not words

    Models read and write tokens, pieces of about four characters. 1 million tokens is roughly 750,000 English words; German and Czech need a few more tokens per word.

  2. 02

    Input and output are priced separately

    Reading your prompt is cheap, generating the answer is expensive. Output tokens usually cost three to five times more than input tokens.

  3. 03

    Cached input is cheaper

    If the same long prefix (system prompt, document, chat history) is sent again and again, providers reuse it and bill those tokens at a fraction of the input price.

  4. 04

    Reasoning tokens count as output

    Reasoning models think before they answer. The hidden thinking tokens are billed as output, so a single request can cost several times more than the list price suggests.

Frequently asked questions

Short answers on tokens, prices and how this table is built.

A short question with a short answer uses about 500 to 1,500 tokens in total. On a mid-priced model that is a fraction of a cent. Long documents, large code files and reasoning models push a single request into the cent range.

The highlights at the top show the cheapest model overall, with vision and with reasoning, based on a 3:1 mix of input and output tokens. Cheapest is not always most economical: a small model may need more retries and longer prompts on hard tasks. Use the calculator with your real token counts.

From the providers' public price lists, synchronized regularly. They are on-demand list prices in USD per 1 million tokens, without volume, batch or regional discounts. The date below the table shows the last update. Check the provider's page before a purchase decision.

The prices on this page are the providers' list prices. Our own prices per model are on the pricing page. They include one account for every provider, invoicing with VAT, prepaid credit with no minimum spend and cost tracking to the cent.

Keep system prompts short, cap the maximum output length, route simple tasks to a small model and reserve large or reasoning models for the hard ones. Prompt caching pays off as soon as the same prefix is sent repeatedly, and batch APIs are often half price for jobs that can wait.

Not in the list prices. Some providers offer batch processing at around half price for asynchronous jobs, and enterprise contracts with negotiated rates. This page shows the standard on-demand prices everybody gets.

Every model on this page, one account

One credit for all providers, billed to the cent. No credit card to start.