The cheapest paid LLM API is DeepSeek V4 Flash served by Morph (morph-dsv4flash) at $0.1234375/$0.3475 per 1M tokens, then GLM-5.3-Flash at $0.15/$0.50 and GPT-5.6 Luna at $0.20/$1.20. DeepSeek's own V4 Flash endpoint now bills $0.44/$1.32 at peak hours and $0.22/$0.66 off-peak. The best value is MiniMax M3 ($0.30/$1.20), the cheapest model scoring above 80% on SWE-bench Verified, and Claude Sonnet 5 ($2/$10, 85.2%) at the frontier. The top coding models are Claude Fable 5.1 and Fable 5 ($10/$50; Fable 5 scores 95.0%), Claude Opus 5 ($5/$25), and GPT-5.6 Sol ($4/$20). Fable 5 has been available again since July 1, 2026, and GPT-5.6 reached the public price list at $4/$20 (Sol), $2/$12 (Terra), and $0.20/$1.20 (Luna). Most production systems do not pick one. They route each request to the cheapest model that can handle it. See the Morph model router and the model list at Morph models.
What Is an LLM API
An LLM API is an HTTP interface to a hosted large language model. You send a prompt (plus optional tools, images, or files) to an endpoint such as /v1/chat/completions and receive generated tokens back, billed per million tokens (MTok) of input and output. It replaces running model weights on your own GPUs with a metered service: you run no infrastructure, manage no model updates, and pay only for tokens used.
Every provider on this page works that way. The differences are price, rate limits, context window, and model quality, and they are large. As of September 2, 2026, output tokens cost between $0.50/M (GLM-5.3-Flash) and $50/M among general flagships (Claude Fable 5.1), and Morph serves DeepSeek V4 Flash at $0.3475/M output. Five models sit within 0.4 points of each other on SWE-bench Verified (80.2 to 80.6 percent) while spanning a 10x price range. The cheapest rows are all open-weight models; the open source LLM guide covers their licenses and the point at which self-hosting beats any API rate.
Every price below comes from the provider's official pricing or docs page as of September 2, 2026. Changed since the June 28 revision: Anthropic shipped Opus 5 (July 24) and Fable 5.1 (September 1) and made Sonnet 5's $2/$10 permanent; GPT-5.6 Sol, Terra, and Luna are on OpenAI's public price list; DeepSeek raised V4 Flash from $0.14/$0.28 to $0.44/$1.32 peak with off-peak at half; Z.AI added GLM-5.3 and GLM-5.3-Flash; Moonshot shipped Kimi K3 at $3/$15; Alibaba shipped Qwen3.8 Max at $2/$6; MiniMax made M3's $0.30/$1.20 permanent. LLM prices change quarterly; check the linked provider before committing to a volume contract.
Pricing Table: Flagship and Cheapest Model per Provider
Prices are $ per 1M tokens, input / output, standard (non-batch) tier. Context is the maximum input window.
| Provider | Model | Input/MTok | Output/MTok | Context |
|---|---|---|---|---|
| Anthropic | Claude Fable 5.1 / Fable 5 | $10.00 | $50.00 | 1M |
| OpenAI | GPT-5.5 | $5.00 | $30.00 | 1M |
| Anthropic | Claude Opus 5 / Opus 4.8 | $5.00 | $25.00 | 1M |
| OpenAI | GPT-5.6 Sol | $4.00 | $20.00 | 1M |
| Moonshot | Kimi K3 | $3.00 | $15.00 | 1M |
| Anthropic | Claude Sonnet 4.6 | $3.00 | $15.00 | 1M |
| Gemini 3.1 Pro (preview) | $2.00 | $12.00 | 1M | |
| OpenAI | GPT-5.6 Terra | $2.00 | $12.00 | 1M |
| Anthropic | Claude Sonnet 5 | $2.00 | $10.00 | 1M |
| Gemini 3.5 Flash | $1.50 | $9.00 | n/a | |
| Alibaba | Qwen3.7 Max | $2.50 | $7.50 | 1M |
| Gemini 3.6 Flash | $1.50 | $7.50 | n/a | |
| Alibaba | Qwen3.8 Max | $2.00 | $6.00 | 1M |
| Anthropic | Claude Haiku 4.5 | $1.00 | $5.00 | 200K |
| Z.AI | GLM-5.3 / GLM-5.2 | $1.40 | $4.40 | 1M |
| Morph | morph-glm53-744b (GLM-5.3, 16-bit) | $1.25 | $4.40 | 1M |
| Moonshot | Kimi K2.6 | $0.95 | $4.00 | 256K |
| DeepSeek | V4 Pro (peak / off-peak) | $1.32 / $0.66 | $3.96 / $1.98 | 1M |
| DeepSeek | V4 Flash (peak / off-peak) | $0.44 / $0.22 | $1.32 / $0.66 | 1M |
| MiniMax | M3 (up to 512K input) | $0.30 | $1.20 | 1M |
| OpenAI | GPT-5.6 Luna | $0.20 | $1.20 | 1M |
| Z.AI | GLM-5.3-Flash | $0.15 | $0.50 | n/a |
| Morph | morph-glm53flash (GLM-5.3-Flash, 16-bit) | $0.097 | $0.3395 | n/a |
| Morph | morph-dsv4flash (DeepSeek V4 Flash, 16-bit) | $0.1234375 | $0.3475 | 1M |
| Z.AI | GLM-4.7-Flash | Free | Free | n/a |
Notes that change effective cost: Gemini 3.1 Pro rises to $4/$18 for prompts over 200K tokens. GPT-5.6 doubles input to $8 (Sol), $4 (Terra), and $0.40 (Luna) on long-context requests, with output at $30, $18, and $1.80. MiniMax M3 lists $0.60/$2.40 with a permanent 50% discount to $0.30/$1.20 up to 512K input tokens; above 512K it is $0.60/$2.40. DeepSeek bills peak rates 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays and half price at all other hours. Anthropic charges the same per-token rate at any context length (a 900K-token request bills like a 9K one), but models from Opus 4.7 onward use a new tokenizer that produces roughly 30% more tokens for the same text, which inflates effective per-request cost in cross-provider comparisons.
Batch and cache discounts: OpenAI cached input is 10x cheaper ($0.40/M on GPT-5.6 Sol, $0.50/M on GPT-5.5). Anthropic batch is 50% off and cache reads are 0.1x base input (0.025x on Fable 5.1 and Mythos 5.1, so $0.25/M). Gemini batch is half price. GLM-5.3 cached input is $0.26/M. Kimi K3 cache hits are $0.30/M. DeepSeek cache hits drop V4 Flash input to $0.014/M at peak, roughly 30x cheaper. If your workload re-sends the same system prompt or file context, the cache column matters more than the headline price.
Provider-by-Provider Breakdown
OpenAI
GPT-5.6 is now the flagship line on OpenAI's public price list in three sizes: Sol ($4/$20), Terra ($2/$12), and Luna ($0.20/$1.20), each with a long-context tier at $8/$30, $4/$18, and $0.40/$1.80. GPT-5.5 ($5/$30, 1,050,000-token context, 128K max output) remains available at its launch price. Cached input is billed at 10% of base across the line. OpenAI renamed priority processing to Fast mode on July 30, 2026. The Daybreak cyber models (gpt-5.6-cyber, $12.50/$75) are gated. On Scale's standardized SWE-bench Pro public leaderboard, the gpt-5.4 (xHigh) variant leads at 59.10%. See Codex pricing for the subscription side.
| Model | Input/MTok | Cached In | Output/MTok | Long-context In/Out |
|---|---|---|---|---|
| GPT-5.5 | $5.00 | $0.50 | $30.00 | n/a |
| GPT-5.6 Sol | $4.00 | $0.40 | $20.00 | $8.00 / $30.00 |
| GPT-5.6 Terra | $2.00 | $0.20 | $12.00 | $4.00 / $18.00 |
| GPT-5.6 Luna | $0.20 | $0.02 | $1.20 | $0.40 / $1.80 |
Anthropic (Claude)
Anthropic's current lineup is Claude Fable 5.1 ($10/$50, released September 1, 2026), Claude Opus 5 ($5/$25, released July 24, 2026, the model Anthropic recommends as the default), Claude Sonnet 5 ($2/$10), and Claude Haiku 4.5 ($1/$5). Sonnet 5's $2/$10 was announced as introductory pricing through August 31; Anthropic has made it the standard price and cancelled the planned increase to $3/$15. Fable 5 (95.0% on SWE-bench Verified) has been available again since July 1, 2026, after the June 12 export-control suspension was lifted, and stays on the price list as a legacy model alongside Opus 4.8 (88.6%), Opus 4.7, Opus 4.6, and Sonnet 4.6. Mythos 5.1 ($10/$50) is limited availability. Opus 5 has no published SWE-bench score; Anthropic reports it within 0.5% of Fable 5 on CursorBench at half the cost per task. All current Claude models carry a 1M context with no long-context surcharge (Haiku 4.5: 200K), but models from Opus 4.7 onward use a tokenizer that produces roughly 30% more tokens for the same text. Full breakdown: Anthropic API pricing.
| Model | Input/MTok | Cache Read | Output/MTok | Context |
|---|---|---|---|---|
| Claude Fable 5.1 | $10.00 | $0.25 | $50.00 | 1M |
| Claude Fable 5 (legacy) | $10.00 | $1.00 | $50.00 | 1M |
| Claude Opus 5 | $5.00 | $0.50 | $25.00 | 1M |
| Claude Opus 4.8 / 4.7 / 4.6 (legacy) | $5.00 | $0.50 | $25.00 | 1M |
| Claude Sonnet 5 | $2.00 | $0.20 | $10.00 | 1M |
| Claude Sonnet 4.6 (legacy) | $3.00 | $0.30 | $15.00 | 1M |
| Claude Haiku 4.5 | $1.00 | $0.10 | $5.00 | 200K |
Google (Gemini)
Gemini 3.1 Pro is still labelled preview at $2/$12 for prompts up to 200K tokens ($4/$18 above) and scores 80.6% on SWE-bench Verified, with a 1,048,576-token input context and 65,536 max output. Gemini 3.6 Flash ($1.50/$7.50) is the new stable model Google positions as its most intelligent speed-tier model; Gemini 3.5 Flash ($1.50/$9) remains listed. Gemini 3.5 Flash-Lite ($0.30/$2.50) is the budget tier. Batch and Flex run at half price across the line.
DeepSeek
DeepSeek is no longer the price floor on its own API. The current models are DeepSeek-V4-Flash-0731 and DeepSeek-V4-Pro-0813, both with a 1M-token context and 384K max output. V4 Flash bills $0.44/$1.32 at peak (01:00 to 04:00 and 06:00 to 10:00 UTC, weekdays) and $0.22/$0.66 off-peak; V4 Pro bills $1.32/$3.96 peak and $0.66/$1.98 off-peak; cache hits are $0.014/M and $0.044/M at peak. V4 (released April 24, 2026) is open weights: V4 Pro is 1.6T total / 49B active parameters, V4 Flash 284B / 13B. DeepSeek-V4-Pro-Max scores 80.6% on SWE-bench Verified, the highest open-weights entry, tied with Gemini 3.1 Pro. The API accepts both OpenAI and Anthropic request formats and the Responses API. Deep dive: DeepSeek V4.
Where you run DeepSeek changes its output and its price. Most serverless providers quantize activations to fp8 to cut cost, which degrades quality away from the reference weights. Morph serves open-source models (DeepSeek V4 Flash, GLM-5.3, GLM-5.3-Flash, Kimi K3) with 16-bit (bf16) activations and does not quantize them, so output matches the released weights. morph-dsv4flash is $0.1234375/M input and $0.3475/M output with no peak surcharge, 4.5x below DeepSeek's own peak rate, and morph-glm53-744b is $1.25/$4.40 with a 1M context. For coding agents specifically, Morph adds codegen-tuned speculative decoding and custom low-level inference kernels, so it is the fastest and highest-fidelity way to run open-source models for codegen. See Morph models and pricing.
Moonshot (Kimi)
Kimi K3 ($3.00/$15.00, cache hits $0.30) is Moonshot's flagship: 2.8T parameters, a 1,048,576-token context, always-on reasoning with low, high, and max effort, and flat pricing with no context-length tiers. Kimi K2.6 ($0.95/$4.00, 256K context, 80.2% on SWE-bench Verified) and the coding-tuned Kimi K2.7 Code (256K) remain available. Morph serves Kimi K3 as morph-kimik3 at $2.60/$14.00 with 16-bit activations. Deep dive: Kimi K3.
Z.AI (GLM)
GLM-5.3 ($1.40/$4.40, cached input $0.26/M, 1M context, 128K max output) is Z.AI's open-weights flagship, priced identically to GLM-5.2. It always runs with reasoning on at low, high, or max effort; Z.AI reports 31.4% on its agentic coding suite at high effort against 29.5% for Claude Opus 4.8, and 34.5% at max against 39.5% for Claude Fable 5. GLM-5.3-Flash lists $0.15/$0.50 with a 50% promotion to $0.075/$0.25 through September 9, 2026. GLM-4.7-Flash and GLM-4.5-Flash are free, the only free named text models from a major provider. Morph serves morph-glm53-744b at $1.25/$4.40 and morph-glm53flash at $0.13/$0.45 with 16-bit activations. Deep dive: GLM-5 family.
MiniMax
MiniMax M3 ($0.30/$1.20 up to 512K input tokens, a permanent 50% discount off the $0.60/$2.40 list price; $0.60/$2.40 above 512K; cache reads $0.06) scores 80.5% on SWE-bench Verified and is open weights (released June 1, 2026). It is the cheapest 80%+ SWE-bench model on a hosted API, about one-twentieth of Opus 4.8's output price for a score 8 points lower. A priority tier runs $0.45/$1.80. MiniMax M2.7 ($0.30/$1.20, 200K context) is the prior generation.
Alibaba (Qwen)
Qwen3.8 Max ($2.00/$6.00 international, 1M context, released September 2, 2026 as qwen3.8-max-0902) is Alibaba's new proprietary flagship and undercuts Qwen3.7 Max, which now lists at $2.50/$7.50 and scores 80.4% on SWE-bench Verified. Qwen3.6 Plus (78.8%) and the open-weights Qwen3.6-27B (Apache 2.0, 77.2%) remain available. New Model Studio users get 1M free tokens valid 90 days; batch is half price and cannot stack with the context-cache discount. The Qwen API guide has the current model IDs, base URLs by region, rate limits, and the open-weight self-hosting path.
Aggregators and Cloud Resellers
OpenRouter fronts hundreds of models behind one OpenAI-compatible key, useful for evaluation and failover at a small markup. AWS Bedrock and Azure / Microsoft Foundry resell first-party models (Claude Opus 4.8, GPT-5.5) with VPC integration, compliance certifications, and consolidated cloud billing. If you need a routing layer rather than a reseller, see LLM gateways and LLM routers.
Morph (specialized)
Morph serves task-specific models rather than general chat: Fast Apply merges code edits at 10,500 tok/s, and WarpGrep does agentic codebase search at $0.80 per 100K tokens. Covered in Specialized APIs below.
LLM API Rate Limits by Provider and Tier
Rate limits decide whether your launch survives traffic, and almost no comparison page publishes them. Here is what each provider enforces, from official docs as of June 28, 2026.
Anthropic: spend-based tiers, cache-aware token counting
Tiers advance automatically by cumulative credit purchase: Tier 1 at $5, Tier 2 at $40, Tier 3 at $200, Tier 4 at $400. Monthly spend caps are $500 / $500 / $1,000 / $200,000. Limits are per model class, measured in requests per minute (RPM), input tokens per minute (ITPM), and output tokens per minute (OTPM). Cache reads do not count toward ITPM: with a 2M ITPM limit and an 80% cache hit rate you can process 10M total input tokens per minute.
| Model class | Tier 1 (RPM / ITPM / OTPM) | Tier 4 (RPM / ITPM / OTPM) |
|---|---|---|
| Claude Opus 4.x | 50 / 500K / 80K | 4,000 / 10M / 800K |
| Claude Sonnet 4.x | 50 / 30K / 8K | 4,000 / 2M / 400K |
| Claude Haiku 4.5 | 50 / 50K / 10K | 4,000 / 4M / 800K |
Counterintuitive: Opus gets 16x the Tier 1 input throughput of Sonnet (500K vs 30K ITPM). If you are rate-limit-bound on a new Anthropic account, the expensive model is also the one you can call hardest.
OpenAI: spend-unlocked tiers, per-model limits
| Tier | Qualification | Monthly usage cap |
|---|---|---|
| Free | Allowed geography | $100 |
| Tier 1 | $5 paid | $100 |
| Tier 2 | $50 paid | $500 |
| Tier 3 | $100 paid | $1,000 |
| Tier 4 | $250 paid | $5,000 |
| Tier 5 | $1,000 paid | $200,000 |
Per-model RPM/TPM numbers are published on each model's page in the OpenAI console rather than a single table, and GPT-5.5 carries a separate limit for long-context requests.
DeepSeek: concurrency caps instead of token budgets
DeepSeek publishes no RPM/TPM limits. It caps concurrent requests: 2,500 in flight on V4 Flash, 500 on V4 Pro. For batch-style pipelines this is friendlier than token-per-minute budgets; for bursty single requests it makes no difference.
Others
Moonshot, Z.AI, MiniMax, and Alibaba scale limits with account spend and publish them in their consoles rather than public tables. MiniMax sells a priority tier for latency-sensitive traffic. Bedrock and Azure enforce cloud-account-level quotas you raise through support tickets.
OpenAI-Compatibility Matrix
Most providers accept OpenAI-format /v1/chat/completions requests, so switching is a base_url and api_key change. The exceptions matter when you build against provider-specific features.
| Provider | OpenAI chat format | Anthropic format | Streaming | Tool calls |
|---|---|---|---|---|
| OpenAI | Native | No | Yes | Yes |
| Anthropic | Via OpenAI SDK compat layer | Native | Yes | Yes |
| Google Gemini | Compat endpoint | No | Yes | Yes |
| DeepSeek | Yes | Yes | Yes | Yes |
| Moonshot (Kimi) | Yes | No | Yes | Yes |
| Z.AI (GLM) | Yes | No | Yes | Yes |
| MiniMax | Yes | No | Yes | Yes |
| Alibaba (Qwen) | Compat mode | No | Yes | Yes |
| OpenRouter | Native (aggregator) | No | Yes | Yes |
| Morph | Yes | No | Yes | n/a (task models) |
DeepSeek is the only first-party provider that natively speaks both OpenAI and Anthropic formats, which means it drops into Claude-Code-style agents without a proxy. If you point a coding agent at a custom provider, the agent must support it: Codex, for example, configures custom providers in config.toml via model_providers entries with base_url, env_key, and wire_api = "responses" (the only wire API it supports), covered in Codex provider configuration.
Benchmarks vs Price: What You Get per Dollar
SWE-bench Verified (real GitHub issues, verified fixes) is the most cited coding benchmark. Scores below are from the llm-stats tracker, September 2, 2026; prices are official API rates. Claude Fable 5.1, Claude Opus 5, and GPT-5.6 have no entry on the tracker yet.
| Model | SWE-bench Verified | Input/MTok | Output/MTok | Value ($/MTok out per point) |
|---|---|---|---|---|
| Claude Fable 5 | 95.0% | $10.00 | $50.00 | $0.53 |
| Claude Mythos Preview (limited) | 93.9% | restricted | restricted | n/a |
| Claude Opus 4.8 | 88.6% | $5.00 | $25.00 | $0.28 |
| Claude Opus 4.7 | 87.6% | $5.00 | $25.00 | $0.29 |
| Claude Sonnet 5 | 85.2% | $2.00 | $10.00 | $0.12 |
| DeepSeek V4 Pro Max | 80.6% | open weights | open weights | open weights |
| Gemini 3.1 Pro | 80.6% | $2.00 | $12.00 | $0.15 |
| MiniMax M3 | 80.5% | $0.30 | $1.20 | $0.015 |
| Qwen3.7 Max | 80.4% | $2.50 | $7.50 | $0.09 |
| Kimi K2.6 | 80.2% | $0.95 | $4.00 | $0.05 |
Five models score between 80.2% and 80.6% on SWE-bench Verified: DeepSeek V4 Pro Max, Gemini 3.1 Pro, MiniMax M3, Qwen3.7 Max, and Kimi K2.6. Within that 0.4-point band, output prices run from $1.20/M (MiniMax M3) to $12/M (Gemini 3.1 Pro), and DeepSeek V4 Pro Max ships it in open weights. Claude Sonnet 5 adds 4.6 points for $10/M output. The next step to Opus 4.8 (88.6%) costs $25/M, and Fable 5 (95.0%) costs $50/M, a 42x step from MiniMax M3. Whether those points are worth it depends on whether your tasks live in the gap.
On the harder SWE-bench Pro (1,865 tasks, 41 professional repos), Scale's standardized public leaderboard tops out at gpt-5.4 (xHigh) 59.10%, Claude Opus 4.6 (thinking) 51.90%, and Gemini 3.1 Pro (thinking) 46.10%. Vendor self-reported aggregates run much higher (llm-stats lists Claude Opus 4.8 at 69.2% and GLM-5.2 at 62.1%), so compare scores only within the same harness.
The SWE-bench Verified figures above are vendor self-reported. The llm-stats tracker lists 102 self-reported results and 0 independently verified, so treat the rankings as vendor claims. Scale's SWE-bench Pro public set runs on a single standardized harness (Pass@1), which is why its numbers are much lower and more directly comparable across models.
LLM API Free Tiers: Exact Amounts
| Provider | Free offer | Limit |
|---|---|---|
| Z.AI | GLM-4.7-Flash, GLM-4.5-Flash | Free models, no token charge |
| Alibaba Model Studio | 1M tokens for new users | 90-day validity |
| OpenAI API | Free tier in allowed regions | $100/month usage cap |
| OpenAI Codex | Included with ChatGPT Free | Lowest 5-hour-window message limits |
| Google Gemini API | Free development tier | Reduced rate limits |
| Anthropic | None | Tier 1 starts at $5 credit |
| Morph WarpGrep | None | $0.80 per 100K tokens |
For experimentation, the practical order is: GLM Flash models (unlimited free), Qwen's 1M tokens, then Gemini's free tier. For coding agents specifically, Codex CLI works with a free ChatGPT sign-in, with the lowest usage limits.
Cost Calculator: Real Workloads
Per-token prices mean nothing until mapped to usage. Two reference workloads, 30-day months, no cache discounts applied (caching reduces all of these).
Coding agent: 50M input / 5M output tokens per day
| Model | Daily cost | Monthly cost |
|---|---|---|
| Claude Fable 5.1 | $750 | $22,500 |
| GPT-5.5 | $400 | $12,000 |
| Claude Opus 5 | $375 | $11,250 |
| GPT-5.6 Sol | $300 | $9,000 |
| Kimi K3 | $225 | $6,750 |
| Gemini 3.1 Pro (≤200K prompts) | $160 | $4,800 |
| Claude Sonnet 5 | $150 | $4,500 |
| GLM-5.3 | $92 | $2,760 |
| DeepSeek V4 Flash (peak) | $28.60 | $858 |
| MiniMax M3 | $21 | $630 |
| GLM-5.3-Flash | $10 | $300 |
| Morph morph-dsv4flash | $6.33 | $190 |
Support chatbot: 20M input / 5M output tokens per day
| Model | Daily cost | Monthly cost |
|---|---|---|
| Claude Haiku 4.5 | $45.00 | $1,350 |
| Gemini 3.5 Flash-Lite | $18.50 | $555 |
| DeepSeek V4 Flash (peak) | $15.40 | $462 |
| MiniMax M3 | $12.00 | $360 |
| GPT-5.6 Luna | $10.00 | $300 |
| GLM-5.3-Flash | $5.50 | $165 |
| Morph morph-dsv4flash | $3.37 | $101 |
The same coding-agent workload costs $22,500/month on Claude Fable 5.1 and $190/month on Morph's morph-dsv4flash, a 118x gap; against DeepSeek's own peak rate it is 26x. The production answer is rarely either extreme: route routine edits to a cheap model, escalate multi-file refactors to a frontier one, and cache aggressively (DeepSeek cache hits bill input at $0.014/M at peak; Anthropic cache reads are 0.1x and do not count against rate limits). Model the routing split with the LLM cost calculator.
Latency and Throughput
Two metrics matter: time to first token (how fast streaming starts) and tokens per second (how fast it finishes). Provider speed claims vary with load and region, so measure on your own traffic.
Measured medians from public endpoint benchmarks (Artificial Analysis, June 2026). General-purpose chat models cluster in the tens-to-low-hundreds of tokens per second; task-specific models built for one operation run far higher.
| Model | Output tok/s (median) | TTFT | Provider |
|---|---|---|---|
| Morph Fast Apply | 10,500 | sub-second | Morph |
Why Morph Fast Apply runs two orders of magnitude above frontier chat models: applying a code edit (merging a lazy edit snippet into the full file) does not need frontier reasoning, so the model is tuned for throughput with codegen-specific speculative decoding. For published per-model output-speed and TTFT figures across the general-purpose providers, see the live Artificial Analysis benchmark.
What the official docs also establish:
- Reasoning adds latency by design. DeepSeek V4 exposes separate thinking and non-thinking modes so you can opt out per request; Claude Opus 4.8 runs adaptive thinking, and thinking tokens are generated and billed before visible output.
- OpenAI sells speed explicitly: Fast mode (renamed from priority processing on July 30, 2026) bills a premium over the standard per-token rate.
- MiniMax sells a priority tier for faster scheduling on M3.
- Specialized models break the general-purpose ceiling: Morph Fast Apply sustains 10,500 tok/s on code-edit application, two orders of magnitude above frontier chat models, because the task (merging an edit into a file) does not need frontier reasoning.
Specialized APIs: When General-Purpose Falls Short
Coding agents spend most of their compute on two operations: searching codebases for context and applying edits to files. Cognition (the team behind Devin) measured 60% of agent time on search alone. Both operations run through general-purpose LLMs by default, at general-purpose prices and speeds.
Morph Fast Apply
Code-edit application at 10,500 tok/s with 98% accuracy. The agent outputs a lazy edit snippet; Fast Apply merges it into the full file in 1-3 seconds. OpenAI-compatible /v1/chat/completions endpoint.
Morph WarpGrep
RL-trained agentic codebase search: 8 parallel tool calls per turn, 4 turns, sub-6s searches, 0.73 F1. $0.80 per 100K tokens. Ships as an MCP server for any agent.
The pattern generalizes: a frontier model reasons and decides what to change; narrow, fast models execute the mechanical steps. The frontier model's output shrinks (edit snippets instead of whole files), which is exactly the token class that costs $12-30/M. See Fast Apply and WarpGrep.
Best LLM API by Use Case
One ranking does not fit every job. Pick by the constraint that binds you. Each pick below comes from the data already on this page.
Best for reasoning and hard coding
Claude Fable 5.1 and Fable 5 ($10/$50; Fable 5 scores 95.0% on SWE-bench Verified) lead, with Claude Opus 5 ($5/$25) as Anthropic's recommended default and GPT-5.6 Sol ($4/$20) as OpenAI's flagship. On the harder SWE-bench Pro, gpt-5.4 (xHigh) leads Scale's standardized public set at 59.10%. Use these for multi-file refactors and tasks where a cheaper model measurably fails.
Best value
Claude Sonnet 5 ($2/$10) scores 85.2% on SWE-bench Verified at $0.12 of output per benchmark point versus $0.28 for Opus 4.8 and $0.53 for Fable 5. MiniMax M3 ($0.30/$1.20) scores 80.5%, the cheapest model above 80%, at $0.015 per point. DeepSeek V4 Pro Max (80.6%, open weights) matches it for self-hosting.
Best for speed
For code edits, Morph Fast Apply is the uncontested pick: 10,500 tok/s with 98% accuracy, two orders of magnitude above frontier chat models, because merging an edit into a file does not need frontier reasoning. For frontier chat speed, OpenAI and Anthropic both sell paid fast tiers (OpenAI Fast mode, formerly priority processing; Anthropic fast mode on Opus).
Morph Fast Apply: best for code edits
10,500 tok/s, 98% accuracy. The frontier model decides what to change; Fast Apply merges the edit snippet into the full file in 1-3 seconds, behind an OpenAI-compatible endpoint. See /products/compact.
Morph model router
Route each request to the cheapest model that can handle it, roughly 430ms of routing overhead at about $0.001 per request. One key, OpenAI-compatible. See /llm-router.
Best on a budget
Morph morph-dsv4flash ($0.1234375/$0.3475, 1M context, DeepSeek V4 Flash at 16-bit) is the price floor on a paid API, followed by GLM-5.3-Flash ($0.15/$0.50) and GPT-5.6 Luna ($0.20/$1.20). DeepSeek's own V4 Flash is $0.44/$1.32 at peak and $0.22/$0.66 off-peak. GLM-4.7-Flash and GLM-4.5-Flash are free on the Z.AI API. For high-volume simple traffic, start here and escalate only the requests that measurably need a frontier model.
How to Choose an LLM API
- 1. Establish your quality floor cheaply. Run your real prompts through morph-dsv4flash ($0.3475/M out), GLM-5.3-Flash ($0.50/M), MiniMax M3 ($1.20/M), and GPT-5.6 Luna ($1.20/M). If one passes, you are done at a fraction of frontier cost.
- 2. Escalate only measured gaps. Move to Claude Sonnet 5 ($10/M), Gemini 3.1 Pro ($12/M), GPT-5.6 Sol ($20/M), Opus 5 ($25/M), or Fable 5.1 ($50/M) for the tasks where the cheap tier measurably fails.
- 3. Check the constraint that binds you. Rate-limited on day one? Anthropic Tier 1 gives Opus 500K ITPM vs Sonnet's 30K. Need self-hosting or data control? DeepSeek V4, GLM-5.3, Kimi K3, and MiniMax M3 are open weights. Compliance? Bedrock or Foundry.
- 4. Use specialized APIs for mechanical steps. Edit application, search, embeddings, and reranking all have purpose-built models that beat $15-30/M generalists on both speed and cost.
The most common mistake is anchoring on one provider's flagship and never testing down. Five models clear 80% on SWE-bench Verified for $1.20 to $12 per million output; the frontier (GPT-5.6 Sol, Opus 5, Fable 5.1) costs $20 to $50. The second most common is ignoring caching: at an 80% cache hit rate, Anthropic bills cached reads at 0.1x (0.025x on Fable 5.1) and exempts them from rate limits, and DeepSeek drops cached input to $0.014/M. For long-context workloads, compare windows in detail at LLM context window comparison.
Frequently Asked Questions
What is an LLM API?
An HTTP interface to a hosted large language model. You POST a prompt to an endpoint such as /v1/chat/completions and receive generated tokens, billed per million tokens of input and output. It replaces self-hosted GPU inference with a metered service.
What are the best LLM API providers in 2026?
First-party: Anthropic (Claude Fable 5.1, Opus 5, Sonnet 5), OpenAI (GPT-5.6 Sol, Terra, Luna; GPT-5.5), Google (Gemini 3.1 Pro, 3.6 Flash), DeepSeek (V4 Flash, V4 Pro), Moonshot (Kimi K3, K2.6), Z.AI (GLM-5.3, GLM-5.3-Flash), MiniMax (M3), and Alibaba (Qwen3.8 Max, Qwen3.7 Max). Aggregator: OpenRouter. Cloud resellers: Bedrock and Azure / Microsoft Foundry. Specialized: Morph (open-weight models at 16-bit, Fast Apply, WarpGrep). Which is best depends on the binding constraint: benchmark score (Anthropic), price (Morph, Z.AI, MiniMax), free tier (Z.AI), or compliance (cloud resellers).
What is the cheapest LLM API in 2026?
Among paid APIs, Morph's morph-dsv4flash at $0.1234375/M input and $0.3475/M output (DeepSeek V4 Flash, 1M context, 16-bit), then GLM-5.3-Flash at $0.15/$0.50 and GPT-5.6 Luna at $0.20/$1.20. DeepSeek's own V4 Flash is $0.44/$1.32 at peak, $0.22/$0.66 off-peak, $0.014/M on cache hits. MiniMax M3 ($0.30/$1.20) is the cheapest model above 80% on SWE-bench Verified. GLM-4.7-Flash is free outright.
Which LLM API is best for coding?
Claude Fable 5 (95.0% on SWE-bench Verified, $10/$50) is available again, and Fable 5.1 (September 1, 2026, same price) succeeds it; Claude Opus 4.8 (88.6%, $5/$25) and Claude Sonnet 5 (85.2%, $2/$10) are the next tiers, and Opus 5 ($5/$25) has no published SWE-bench score. On Scale's standardized SWE-bench Pro, gpt-5.4 (xHigh) leads the public set at 59.10%. For budget coding, MiniMax M3 (80.5%, $0.30/$1.20) and open-weights DeepSeek V4 Pro Max (80.6%) are the value picks. For applying edits an agent has already decided on, Morph Fast Apply runs at 10,500 tok/s with 98% accuracy.
What rate limits do LLM APIs have?
Anthropic: tiered by cumulative deposit ($5 to $400); Tier 1 gives Opus 4.x 50 RPM / 500K input tokens per minute, Tier 4 gives 4,000 RPM / 10M ITPM, and cached tokens are exempt. OpenAI: tiers unlock at $5 to $1,000 paid with $100 to $200,000 monthly usage caps; per-model RPM/TPM live on each model page. DeepSeek: concurrency caps of 2,500 (V4 Flash) and 500 (V4 Pro) instead of token budgets.
Which LLM API has the largest context window?
1M tokens is the 2026 flagship standard: Claude Fable 5.1, Opus 5, Sonnet 5, and the legacy Opus 4.x and Sonnet 4.6 (all with no long-context surcharge), GPT-5.6 and GPT-5.5 (1,050,000), Gemini 3.1 Pro (1,048,576), DeepSeek V4 Pro and Flash, Kimi K3 (1,048,576), MiniMax M3 (1,048,576), Qwen3.8 Max, Qwen3.7 Max, and GLM-5.3. Below 1M: Kimi K2.6 and K2.7 Code at 256K, Claude Haiku 4.5 and MiniMax M2.7 at 200K.
Are LLM APIs interchangeable / OpenAI-compatible?
DeepSeek, Moonshot, Z.AI, MiniMax, Qwen, OpenRouter, and Morph accept OpenAI-format requests directly; Google exposes a compatibility endpoint; Anthropic offers an OpenAI SDK compatibility layer over its native Messages API. DeepSeek also accepts Anthropic-format requests. In practice, switching providers is a base URL and key change, with re-testing for tool-calling behavior.
Should I use one provider or multiple?
Multiple, behind a router. Send high-volume simple traffic to a sub-$1.50/M model and escalate hard tasks to a frontier model. The 118x spread between Fable 5.1 and morph-dsv4flash on identical traffic is the budget you are leaving on the table with a single-provider setup. See LLM gateways for the plumbing.
Which LLM API is best by use case?
Reasoning and hard coding: Claude Fable 5.1 ($10/$50), Claude Opus 5 ($5/$25), and GPT-5.6 Sol ($4/$20). Value: Claude Sonnet 5 (85.2%, $2/$10) and MiniMax M3 ($0.30/$1.20), the cheapest model above 80%, with open-weights DeepSeek V4 Pro Max (80.6%) for self-hosting. Speed: Morph Fast Apply for code edits at 10,500 tok/s. Budget: Morph morph-dsv4flash ($0.1234375/$0.3475), GLM-5.3-Flash ($0.15/$0.50), or the free GLM Flash models on Z.AI. Most teams route across these with the Morph model router.
Sources
Every price and benchmark on this page traces to a primary source, checked September 2, 2026 (rate-limit tiers were last checked June 28, 2026):
- OpenAI API pricing: developers.openai.com/api/docs/pricing and the GPT-5.5 model page
- Anthropic (Claude) pricing and models: platform.claude.com pricing, models overview, Introducing Claude Opus 5
- Google Gemini API pricing: ai.google.dev/gemini-api/docs/pricing
- DeepSeek API pricing: api-docs.deepseek.com pricing
- Z.AI pricing and GLM-5.3 docs: docs.z.ai pricing, docs.z.ai GLM-5.3
- Moonshot Kimi K3 pricing: platform.kimi.ai Kimi K3 pricing
- MiniMax pay-as-you-go pricing: platform.minimax.io pricing
- Alibaba Model Studio pricing: alibabacloud.com Model Studio pricing
- Morph model pricing: morphllm.com/api/models/json
- SWE-bench Verified leaderboard: llm-stats.com SWE-bench Verified
- SWE-bench Pro public leaderboard: labs.scale.com SWE-bench Pro
Related Resources
Code Editing at 10,500 tok/s
Frontier LLM APIs bill $12-30 per million output tokens to rewrite whole files. Morph Fast Apply merges edit snippets into files at 10,500 tok/s with 98% accuracy, behind an OpenAI-compatible API.
