LLM API Pricing
Cost per 1M tokens · 420 models · Verified Sep 4, 2026 via OpenRouter API · Prices move — check the date.
LLM Token Expenditure Index (SDLLMTK)usage-weighted effective price, USD per 1M tokens — as reported by Silicon Data
What the market actually pays per 1M inference tokens — an expenditure-weighted average across frontier & open-weight providers, normalized for input/output mix and batching. Daily readings are captured every night from Silicon Data's public page, with attribution; earlier points as reported by CIO.com & Bloomberg coverage (May 2026 peak ≈ $2.03, inception Dec 2025 ≈ $1.00). Source: Silicon Data
| Model | $ / 1M in | $ / 1M out | Cache read | Context | Provider |
|---|---|---|---|---|---|
| cohere/north-mini-code:free | free | free | — | 256K | |
| dots-studio/dots-3-note-preview:free | free | free | — | 512K | |
| google/gemma-4-26b-a4b-it:free | free | free | — | 262K | |
| google/gemma-4-31b-it:free | free | free | — | 262K | |
| google/lyria-3-clip-preview | free | free | — | 1049K | |
| google/lyria-3-pro-preview | free | free | — | 1049K | |
| inclusionai/ling-3.0-flash-fin:free | free | free | — | 262K | |
| liquid/lfm-2.5-2.6b:free | free | free | — | 66K | |
| minimax/minimax-m2.7:free | free | free | — | 197K | |
| minimax/minimax-m3:free | free | free | — | 1049K | |
| nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free | free | free | — | 256K | |
| nvidia/nemotron-3-super-120b-a12b:free | free | free | — | 262K | |
| nvidia/nemotron-3-ultra-550b-a55b:free | free | free | — | 1000K | |
| nvidia/nemotron-3.5-content-safety:free | free | free | — | 128K | |
| nvidia/nemotron-3.5-lightning:free | free | free | — | 1000K | |
| openrouter/free | free | free | — | 200K | |
| poolside/laguna-s-2.1:free | free | free | — | 262K | |
| poolside/laguna-xs-2.1:free | free | free | — | 262K | |
| thinkingmachines/inkling-small:free | free | free | — | 1049K | |
| thinkingmachines/inkling:free | free | free | — | 1049K | |
| z-ai/glm-5.2:free | free | free | — | 256K | |
| ibm-granite/granite-4.0-h-micro | $0.0170 | $0.1120 | — | 131K | |
| mistralai/mistral-nemo | $0.0190 | $0.0300 | — | 131K | |
| inclusionai/ling-3.0-flash | $0.0210 | $0.0630 | $0.0042 | 262K | |
| nex-agi/nex-n2-mini | $0.0250 | $0.1000 | $0.0025 | 262K | |
| meta-llama/llama-3.2-1b-instruct | $0.0270 | $0.2010 | — | 60K | |
| openai/gpt-oss-20b | $0.0300 | $0.1300 | $0.0300 | 131K | |
| qwen/qwen3.7-flash | $0.0300 | $0.1300 | $0.0060 | 1000K | |
| upstage/solar-pro4 | $0.0300 | $0.1200 | $0.0060 | 524K | |
| amazon/nova-micro-v1 | $0.0350 | $0.1400 | — | 128K | |
| openai/gpt-oss-120b | $0.0370 | $0.1700 | — | 131K | |
| cohere/command-r7b-12-2024 | $0.0375 | $0.1500 | — | 128K | |
| inception/mercury-2.5-preview | $0.0400 | $0.1500 | $0.0040 | 260K | |
| sao10k/l3-lunaris-8b | $0.0400 | $0.0500 | — | 8K | |
| tencent/hy-mt2-1.8b | $0.0440 | $0.1770 | — | 8K | |
| qwen/qwen3-30b-a3b-instruct-2507 | $0.0481 | $0.1930 | — | 262K | |
| google/gemma-3-12b-it | $0.0500 | $0.1500 | — | 131K | |
| google/gemma-3-4b-it | $0.0500 | $0.1000 | — | 131K | |
| ibm-granite/granite-4.1-8b | $0.0500 | $0.1000 | $0.0500 | 131K | |
| meta-llama/llama-3.1-8b-instruct | $0.0500 | $0.0800 | $0.0250 | 131K | |
| meta-llama/llama-3.2-3b-instruct | $0.0500 | $0.3300 | — | 131K | |
| mistralai/mistral-small-24b-instruct-2501 | $0.0500 | $0.0800 | — | 33K | |
| nvidia/nemotron-3-nano-30b-a3b | $0.0500 | $0.2000 | $0.0300 | 262K | |
| openai/gpt-5-nano | $0.0500 | $0.4000 | $0.0050 | 400K | |
| deepseek/deepseek-v4-flash-latest | $0.0500 | $0.1600 | $0.0130 | 1311K | |
| amazon/nova-lite-v1 | $0.0600 | $0.2400 | — | 300K | |
| gryphe/mythomax-l2-13b | $0.0600 | $0.0600 | — | 8K | |
| inclusionai/ling-3.0-flash-fin | $0.0600 | $0.1800 | $0.0120 | 262K | |
| poolside/laguna-xs-2.1 | $0.0600 | $0.1200 | $0.0300 | 262K | |
| z-ai/glm-4.7-flash | $0.0600 | $0.4000 | $0.0100 | 203K | |
| deepseek/deepseek-v4-flash-0731 | $0.0650 | $0.1800 | $0.0160 | 1311K | |
| qwen/qwen3.5-flash-02-23 | $0.0650 | $0.2600 | — | 1000K | |
| google/gemma-4-26b-a4b-it | $0.0700 | $0.3400 | — | 262K | |
| microsoft/phi-4 | $0.0700 | $0.1400 | — | 16K | |
| qwen/qwen3-coder-30b-a3b-instruct | $0.0700 | $0.2800 | — | 262K | |
| z-ai/glm-flash-latest | $0.0713 | $0.2375 | $0.0143 | 1311K | |
| tencent/hy-mt2-30b-a3b | $0.0740 | $0.2950 | — | 8K | |
| tencent/hy-mt2-7b | $0.0740 | $0.2950 | — | 8K | |
| bytedance-seed/seed-1.6-flash | $0.0750 | $0.3000 | — | 262K | |
| mistralai/mistral-small-3.2-24b-instruct | $0.0750 | $0.2000 | — | 131K |
Sorted by input price, cheapest first. 60 models shown — use the search bar to browse all 400+.
Data verified on Sep 4, 2026 via the OpenRouter API (all prices include OpenRouter's provider rates). "free" = genuinely $0 models. See also the GPU pricing tracker, the glossary, and our methodology.