Live Data · API

LLM API Pricing

Cost per 1M tokens · 420 models · Verified Sep 4, 2026 via OpenRouter API · Prices move — check the date.

LLM Token Expenditure Index (SDLLMTK)usage-weighted effective price, USD per 1M tokens — as reported by Silicon Data

Bloomberg: SDLLMTK
$0.98/ 1M tokens · as of Sep 2, 2026

What the market actually pays per 1M inference tokens — an expenditure-weighted average across frontier & open-weight providers, normalized for input/output mix and batching. Daily readings are captured every night from Silicon Data's public page, with attribution; earlier points as reported by CIO.com & Bloomberg coverage (May 2026 peak ≈ $2.03, inception Dec 2025 ≈ $1.00). Source: Silicon Data

Cheapest paid model right now: ibm-granite/granite-4.0-h-micro $0.0170/M input · $0.1120/M output
Model$ / 1M in$ / 1M outCache readContextProvider
cohere/north-mini-code:freefreefree256K
dots-studio/dots-3-note-preview:freefreefree512K
google/gemma-4-26b-a4b-it:freefreefree262K
google/gemma-4-31b-it:freefreefree262K
google/lyria-3-clip-previewfreefree1049K
google/lyria-3-pro-previewfreefree1049K
inclusionai/ling-3.0-flash-fin:freefreefree262K
liquid/lfm-2.5-2.6b:freefreefree66K
minimax/minimax-m2.7:freefreefree197K
minimax/minimax-m3:freefreefree1049K
nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:freefreefree256K
nvidia/nemotron-3-super-120b-a12b:freefreefree262K
nvidia/nemotron-3-ultra-550b-a55b:freefreefree1000K
nvidia/nemotron-3.5-content-safety:freefreefree128K
nvidia/nemotron-3.5-lightning:freefreefree1000K
openrouter/freefreefree200K
poolside/laguna-s-2.1:freefreefree262K
poolside/laguna-xs-2.1:freefreefree262K
thinkingmachines/inkling-small:freefreefree1049K
thinkingmachines/inkling:freefreefree1049K
z-ai/glm-5.2:freefreefree256K
ibm-granite/granite-4.0-h-micro$0.0170$0.1120131K
mistralai/mistral-nemo$0.0190$0.0300131K
inclusionai/ling-3.0-flash$0.0210$0.0630$0.0042262K
nex-agi/nex-n2-mini$0.0250$0.1000$0.0025262K
meta-llama/llama-3.2-1b-instruct$0.0270$0.201060K
openai/gpt-oss-20b$0.0300$0.1300$0.0300131K
qwen/qwen3.7-flash$0.0300$0.1300$0.00601000K
upstage/solar-pro4$0.0300$0.1200$0.0060524K
amazon/nova-micro-v1$0.0350$0.1400128K
openai/gpt-oss-120b$0.0370$0.1700131K
cohere/command-r7b-12-2024$0.0375$0.1500128K
inception/mercury-2.5-preview$0.0400$0.1500$0.0040260K
sao10k/l3-lunaris-8b$0.0400$0.05008K
tencent/hy-mt2-1.8b$0.0440$0.17708K
qwen/qwen3-30b-a3b-instruct-2507$0.0481$0.1930262K
google/gemma-3-12b-it$0.0500$0.1500131K
google/gemma-3-4b-it$0.0500$0.1000131K
ibm-granite/granite-4.1-8b$0.0500$0.1000$0.0500131K
meta-llama/llama-3.1-8b-instruct$0.0500$0.0800$0.0250131K
meta-llama/llama-3.2-3b-instruct$0.0500$0.3300131K
mistralai/mistral-small-24b-instruct-2501$0.0500$0.080033K
nvidia/nemotron-3-nano-30b-a3b$0.0500$0.2000$0.0300262K
openai/gpt-5-nano$0.0500$0.4000$0.0050400K
deepseek/deepseek-v4-flash-latest$0.0500$0.1600$0.01301311K
amazon/nova-lite-v1$0.0600$0.2400300K
gryphe/mythomax-l2-13b$0.0600$0.06008K
inclusionai/ling-3.0-flash-fin$0.0600$0.1800$0.0120262K
poolside/laguna-xs-2.1$0.0600$0.1200$0.0300262K
z-ai/glm-4.7-flash$0.0600$0.4000$0.0100203K
deepseek/deepseek-v4-flash-0731$0.0650$0.1800$0.01601311K
qwen/qwen3.5-flash-02-23$0.0650$0.26001000K
google/gemma-4-26b-a4b-it$0.0700$0.3400262K
microsoft/phi-4$0.0700$0.140016K
qwen/qwen3-coder-30b-a3b-instruct$0.0700$0.2800262K
z-ai/glm-flash-latest$0.0713$0.2375$0.01431311K
tencent/hy-mt2-30b-a3b$0.0740$0.29508K
tencent/hy-mt2-7b$0.0740$0.29508K
bytedance-seed/seed-1.6-flash$0.0750$0.3000262K
mistralai/mistral-small-3.2-24b-instruct$0.0750$0.2000131K

Sorted by input price, cheapest first. 60 models shown — use the search bar to browse all 400+.

Data verified on Sep 4, 2026 via the OpenRouter API (all prices include OpenRouter's provider rates). "free" = genuinely $0 models. See also the GPU pricing tracker, the glossary, and our methodology.