LLM API Providers (2026): 12 APIs Compared by Price per 1M Tokens, Rate Limits, and Context

Pricing for 12 LLM API providers, verified September 2, 2026: Claude Fable 5.1 $10/$50, GPT-5.6 Sol $4/$20, Claude Opus 5 $5/$25, Sonnet 5 $2/$10, Gemini 3.1 Pro $2/$12, GLM-5.3 $1.40/$4.40, MiniMax M3 $0.30/$1.20, DeepSeek V4 Flash $0.44/$1.32 per 1M tokens. Plus rate-limit tiers, OpenAI-compatibility matrix, free tiers, and SWE-bench scores.

June 28, 2026 · 1 min read
LLM API Providers (2026): 12 APIs Compared by Price per 1M Tokens, Rate Limits, and Context
Quick answer

The cheapest paid LLM API is DeepSeek V4 Flash served by Morph (morph-dsv4flash) at $0.1234375/$0.3475 per 1M tokens, then GLM-5.3-Flash at $0.15/$0.50 and GPT-5.6 Luna at $0.20/$1.20. DeepSeek's own V4 Flash endpoint now bills $0.44/$1.32 at peak hours and $0.22/$0.66 off-peak. The best value is MiniMax M3 ($0.30/$1.20), the cheapest model scoring above 80% on SWE-bench Verified, and Claude Sonnet 5 ($2/$10, 85.2%) at the frontier. The top coding models are Claude Fable 5.1 and Fable 5 ($10/$50; Fable 5 scores 95.0%), Claude Opus 5 ($5/$25), and GPT-5.6 Sol ($4/$20). Fable 5 has been available again since July 1, 2026, and GPT-5.6 reached the public price list at $4/$20 (Sol), $2/$12 (Terra), and $0.20/$1.20 (Luna). Most production systems do not pick one. They route each request to the cheapest model that can handle it. See the Morph model router and the model list at Morph models.

What Is an LLM API

An LLM API is an HTTP interface to a hosted large language model. You send a prompt (plus optional tools, images, or files) to an endpoint such as /v1/chat/completions and receive generated tokens back, billed per million tokens (MTok) of input and output. It replaces running model weights on your own GPUs with a metered service: you run no infrastructure, manage no model updates, and pay only for tokens used.

Every provider on this page works that way. The differences are price, rate limits, context window, and model quality, and they are large. As of September 2, 2026, output tokens cost between $0.50/M (GLM-5.3-Flash) and $50/M among general flagships (Claude Fable 5.1), and Morph serves DeepSeek V4 Flash at $0.3475/M output. Five models sit within 0.4 points of each other on SWE-bench Verified (80.2 to 80.6 percent) while spanning a 10x price range. The cheapest rows are all open-weight models; the open source LLM guide covers their licenses and the point at which self-hosting beats any API rate.

100x
Output price spread ($0.50 to $50 per MTok, general flagships)
1M
Context window now standard on flagships
95.0%
Top SWE-bench Verified (Claude Fable 5, llm-stats)
$0
GLM-4.7-Flash API price (free)
All prices verified September 2, 2026

Every price below comes from the provider's official pricing or docs page as of September 2, 2026. Changed since the June 28 revision: Anthropic shipped Opus 5 (July 24) and Fable 5.1 (September 1) and made Sonnet 5's $2/$10 permanent; GPT-5.6 Sol, Terra, and Luna are on OpenAI's public price list; DeepSeek raised V4 Flash from $0.14/$0.28 to $0.44/$1.32 peak with off-peak at half; Z.AI added GLM-5.3 and GLM-5.3-Flash; Moonshot shipped Kimi K3 at $3/$15; Alibaba shipped Qwen3.8 Max at $2/$6; MiniMax made M3's $0.30/$1.20 permanent. LLM prices change quarterly; check the linked provider before committing to a volume contract.

Pricing Table: Flagship and Cheapest Model per Provider

Prices are $ per 1M tokens, input / output, standard (non-batch) tier. Context is the maximum input window.

ProviderModelInput/MTokOutput/MTokContext
AnthropicClaude Fable 5.1 / Fable 5$10.00$50.001M
OpenAIGPT-5.5$5.00$30.001M
AnthropicClaude Opus 5 / Opus 4.8$5.00$25.001M
OpenAIGPT-5.6 Sol$4.00$20.001M
MoonshotKimi K3$3.00$15.001M
AnthropicClaude Sonnet 4.6$3.00$15.001M
GoogleGemini 3.1 Pro (preview)$2.00$12.001M
OpenAIGPT-5.6 Terra$2.00$12.001M
AnthropicClaude Sonnet 5$2.00$10.001M
GoogleGemini 3.5 Flash$1.50$9.00n/a
AlibabaQwen3.7 Max$2.50$7.501M
GoogleGemini 3.6 Flash$1.50$7.50n/a
AlibabaQwen3.8 Max$2.00$6.001M
AnthropicClaude Haiku 4.5$1.00$5.00200K
Z.AIGLM-5.3 / GLM-5.2$1.40$4.401M
Morphmorph-glm53-744b (GLM-5.3, 16-bit)$1.25$4.401M
MoonshotKimi K2.6$0.95$4.00256K
DeepSeekV4 Pro (peak / off-peak)$1.32 / $0.66$3.96 / $1.981M
DeepSeekV4 Flash (peak / off-peak)$0.44 / $0.22$1.32 / $0.661M
MiniMaxM3 (up to 512K input)$0.30$1.201M
OpenAIGPT-5.6 Luna$0.20$1.201M
Z.AIGLM-5.3-Flash$0.15$0.50n/a
Morphmorph-glm53flash (GLM-5.3-Flash, 16-bit)$0.097$0.3395n/a
Morphmorph-dsv4flash (DeepSeek V4 Flash, 16-bit)$0.1234375$0.34751M
Z.AIGLM-4.7-FlashFreeFreen/a

Notes that change effective cost: Gemini 3.1 Pro rises to $4/$18 for prompts over 200K tokens. GPT-5.6 doubles input to $8 (Sol), $4 (Terra), and $0.40 (Luna) on long-context requests, with output at $30, $18, and $1.80. MiniMax M3 lists $0.60/$2.40 with a permanent 50% discount to $0.30/$1.20 up to 512K input tokens; above 512K it is $0.60/$2.40. DeepSeek bills peak rates 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays and half price at all other hours. Anthropic charges the same per-token rate at any context length (a 900K-token request bills like a 9K one), but models from Opus 4.7 onward use a new tokenizer that produces roughly 30% more tokens for the same text, which inflates effective per-request cost in cross-provider comparisons.

Batch and cache discounts: OpenAI cached input is 10x cheaper ($0.40/M on GPT-5.6 Sol, $0.50/M on GPT-5.5). Anthropic batch is 50% off and cache reads are 0.1x base input (0.025x on Fable 5.1 and Mythos 5.1, so $0.25/M). Gemini batch is half price. GLM-5.3 cached input is $0.26/M. Kimi K3 cache hits are $0.30/M. DeepSeek cache hits drop V4 Flash input to $0.014/M at peak, roughly 30x cheaper. If your workload re-sends the same system prompt or file context, the cache column matters more than the headline price.

Provider-by-Provider Breakdown

OpenAI

GPT-5.6 is now the flagship line on OpenAI's public price list in three sizes: Sol ($4/$20), Terra ($2/$12), and Luna ($0.20/$1.20), each with a long-context tier at $8/$30, $4/$18, and $0.40/$1.80. GPT-5.5 ($5/$30, 1,050,000-token context, 128K max output) remains available at its launch price. Cached input is billed at 10% of base across the line. OpenAI renamed priority processing to Fast mode on July 30, 2026. The Daybreak cyber models (gpt-5.6-cyber, $12.50/$75) are gated. On Scale's standardized SWE-bench Pro public leaderboard, the gpt-5.4 (xHigh) variant leads at 59.10%. See Codex pricing for the subscription side.

ModelInput/MTokCached InOutput/MTokLong-context In/Out
GPT-5.5$5.00$0.50$30.00n/a
GPT-5.6 Sol$4.00$0.40$20.00$8.00 / $30.00
GPT-5.6 Terra$2.00$0.20$12.00$4.00 / $18.00
GPT-5.6 Luna$0.20$0.02$1.20$0.40 / $1.80

Anthropic (Claude)

Anthropic's current lineup is Claude Fable 5.1 ($10/$50, released September 1, 2026), Claude Opus 5 ($5/$25, released July 24, 2026, the model Anthropic recommends as the default), Claude Sonnet 5 ($2/$10), and Claude Haiku 4.5 ($1/$5). Sonnet 5's $2/$10 was announced as introductory pricing through August 31; Anthropic has made it the standard price and cancelled the planned increase to $3/$15. Fable 5 (95.0% on SWE-bench Verified) has been available again since July 1, 2026, after the June 12 export-control suspension was lifted, and stays on the price list as a legacy model alongside Opus 4.8 (88.6%), Opus 4.7, Opus 4.6, and Sonnet 4.6. Mythos 5.1 ($10/$50) is limited availability. Opus 5 has no published SWE-bench score; Anthropic reports it within 0.5% of Fable 5 on CursorBench at half the cost per task. All current Claude models carry a 1M context with no long-context surcharge (Haiku 4.5: 200K), but models from Opus 4.7 onward use a tokenizer that produces roughly 30% more tokens for the same text. Full breakdown: Anthropic API pricing.

ModelInput/MTokCache ReadOutput/MTokContext
Claude Fable 5.1$10.00$0.25$50.001M
Claude Fable 5 (legacy)$10.00$1.00$50.001M
Claude Opus 5$5.00$0.50$25.001M
Claude Opus 4.8 / 4.7 / 4.6 (legacy)$5.00$0.50$25.001M
Claude Sonnet 5$2.00$0.20$10.001M
Claude Sonnet 4.6 (legacy)$3.00$0.30$15.001M
Claude Haiku 4.5$1.00$0.10$5.00200K

Google (Gemini)

Gemini 3.1 Pro is still labelled preview at $2/$12 for prompts up to 200K tokens ($4/$18 above) and scores 80.6% on SWE-bench Verified, with a 1,048,576-token input context and 65,536 max output. Gemini 3.6 Flash ($1.50/$7.50) is the new stable model Google positions as its most intelligent speed-tier model; Gemini 3.5 Flash ($1.50/$9) remains listed. Gemini 3.5 Flash-Lite ($0.30/$2.50) is the budget tier. Batch and Flex run at half price across the line.

DeepSeek

DeepSeek is no longer the price floor on its own API. The current models are DeepSeek-V4-Flash-0731 and DeepSeek-V4-Pro-0813, both with a 1M-token context and 384K max output. V4 Flash bills $0.44/$1.32 at peak (01:00 to 04:00 and 06:00 to 10:00 UTC, weekdays) and $0.22/$0.66 off-peak; V4 Pro bills $1.32/$3.96 peak and $0.66/$1.98 off-peak; cache hits are $0.014/M and $0.044/M at peak. V4 (released April 24, 2026) is open weights: V4 Pro is 1.6T total / 49B active parameters, V4 Flash 284B / 13B. DeepSeek-V4-Pro-Max scores 80.6% on SWE-bench Verified, the highest open-weights entry, tied with Gemini 3.1 Pro. The API accepts both OpenAI and Anthropic request formats and the Responses API. Deep dive: DeepSeek V4.

Where you run DeepSeek changes its output and its price. Most serverless providers quantize activations to fp8 to cut cost, which degrades quality away from the reference weights. Morph serves open-source models (DeepSeek V4 Flash, GLM-5.3, GLM-5.3-Flash, Kimi K3) with 16-bit (bf16) activations and does not quantize them, so output matches the released weights. morph-dsv4flash is $0.1234375/M input and $0.3475/M output with no peak surcharge, 4.5x below DeepSeek's own peak rate, and morph-glm53-744b is $1.25/$4.40 with a 1M context. For coding agents specifically, Morph adds codegen-tuned speculative decoding and custom low-level inference kernels, so it is the fastest and highest-fidelity way to run open-source models for codegen. See Morph models and pricing.

Moonshot (Kimi)

Kimi K3 ($3.00/$15.00, cache hits $0.30) is Moonshot's flagship: 2.8T parameters, a 1,048,576-token context, always-on reasoning with low, high, and max effort, and flat pricing with no context-length tiers. Kimi K2.6 ($0.95/$4.00, 256K context, 80.2% on SWE-bench Verified) and the coding-tuned Kimi K2.7 Code (256K) remain available. Morph serves Kimi K3 as morph-kimik3 at $2.60/$14.00 with 16-bit activations. Deep dive: Kimi K3.

Z.AI (GLM)

GLM-5.3 ($1.40/$4.40, cached input $0.26/M, 1M context, 128K max output) is Z.AI's open-weights flagship, priced identically to GLM-5.2. It always runs with reasoning on at low, high, or max effort; Z.AI reports 31.4% on its agentic coding suite at high effort against 29.5% for Claude Opus 4.8, and 34.5% at max against 39.5% for Claude Fable 5. GLM-5.3-Flash lists $0.15/$0.50 with a 50% promotion to $0.075/$0.25 through September 9, 2026. GLM-4.7-Flash and GLM-4.5-Flash are free, the only free named text models from a major provider. Morph serves morph-glm53-744b at $1.25/$4.40 and morph-glm53flash at $0.13/$0.45 with 16-bit activations. Deep dive: GLM-5 family.

MiniMax

MiniMax M3 ($0.30/$1.20 up to 512K input tokens, a permanent 50% discount off the $0.60/$2.40 list price; $0.60/$2.40 above 512K; cache reads $0.06) scores 80.5% on SWE-bench Verified and is open weights (released June 1, 2026). It is the cheapest 80%+ SWE-bench model on a hosted API, about one-twentieth of Opus 4.8's output price for a score 8 points lower. A priority tier runs $0.45/$1.80. MiniMax M2.7 ($0.30/$1.20, 200K context) is the prior generation.

Alibaba (Qwen)

Qwen3.8 Max ($2.00/$6.00 international, 1M context, released September 2, 2026 as qwen3.8-max-0902) is Alibaba's new proprietary flagship and undercuts Qwen3.7 Max, which now lists at $2.50/$7.50 and scores 80.4% on SWE-bench Verified. Qwen3.6 Plus (78.8%) and the open-weights Qwen3.6-27B (Apache 2.0, 77.2%) remain available. New Model Studio users get 1M free tokens valid 90 days; batch is half price and cannot stack with the context-cache discount. The Qwen API guide has the current model IDs, base URLs by region, rate limits, and the open-weight self-hosting path.

Aggregators and Cloud Resellers

OpenRouter fronts hundreds of models behind one OpenAI-compatible key, useful for evaluation and failover at a small markup. AWS Bedrock and Azure / Microsoft Foundry resell first-party models (Claude Opus 4.8, GPT-5.5) with VPC integration, compliance certifications, and consolidated cloud billing. If you need a routing layer rather than a reseller, see LLM gateways and LLM routers.

Morph (specialized)

Morph serves task-specific models rather than general chat: Fast Apply merges code edits at 10,500 tok/s, and WarpGrep does agentic codebase search at $0.80 per 100K tokens. Covered in Specialized APIs below.

LLM API Rate Limits by Provider and Tier

Rate limits decide whether your launch survives traffic, and almost no comparison page publishes them. Here is what each provider enforces, from official docs as of June 28, 2026.

Anthropic: spend-based tiers, cache-aware token counting

Tiers advance automatically by cumulative credit purchase: Tier 1 at $5, Tier 2 at $40, Tier 3 at $200, Tier 4 at $400. Monthly spend caps are $500 / $500 / $1,000 / $200,000. Limits are per model class, measured in requests per minute (RPM), input tokens per minute (ITPM), and output tokens per minute (OTPM). Cache reads do not count toward ITPM: with a 2M ITPM limit and an 80% cache hit rate you can process 10M total input tokens per minute.

Model classTier 1 (RPM / ITPM / OTPM)Tier 4 (RPM / ITPM / OTPM)
Claude Opus 4.x50 / 500K / 80K4,000 / 10M / 800K
Claude Sonnet 4.x50 / 30K / 8K4,000 / 2M / 400K
Claude Haiku 4.550 / 50K / 10K4,000 / 4M / 800K

Counterintuitive: Opus gets 16x the Tier 1 input throughput of Sonnet (500K vs 30K ITPM). If you are rate-limit-bound on a new Anthropic account, the expensive model is also the one you can call hardest.

OpenAI: spend-unlocked tiers, per-model limits

TierQualificationMonthly usage cap
FreeAllowed geography$100
Tier 1$5 paid$100
Tier 2$50 paid$500
Tier 3$100 paid$1,000
Tier 4$250 paid$5,000
Tier 5$1,000 paid$200,000

Per-model RPM/TPM numbers are published on each model's page in the OpenAI console rather than a single table, and GPT-5.5 carries a separate limit for long-context requests.

DeepSeek: concurrency caps instead of token budgets

DeepSeek publishes no RPM/TPM limits. It caps concurrent requests: 2,500 in flight on V4 Flash, 500 on V4 Pro. For batch-style pipelines this is friendlier than token-per-minute budgets; for bursty single requests it makes no difference.

Others

Moonshot, Z.AI, MiniMax, and Alibaba scale limits with account spend and publish them in their consoles rather than public tables. MiniMax sells a priority tier for latency-sensitive traffic. Bedrock and Azure enforce cloud-account-level quotas you raise through support tickets.

OpenAI-Compatibility Matrix

Most providers accept OpenAI-format /v1/chat/completions requests, so switching is a base_url and api_key change. The exceptions matter when you build against provider-specific features.

ProviderOpenAI chat formatAnthropic formatStreamingTool calls
OpenAINativeNoYesYes
AnthropicVia OpenAI SDK compat layerNativeYesYes
Google GeminiCompat endpointNoYesYes
DeepSeekYesYesYesYes
Moonshot (Kimi)YesNoYesYes
Z.AI (GLM)YesNoYesYes
MiniMaxYesNoYesYes
Alibaba (Qwen)Compat modeNoYesYes
OpenRouterNative (aggregator)NoYesYes
MorphYesNoYesn/a (task models)

DeepSeek is the only first-party provider that natively speaks both OpenAI and Anthropic formats, which means it drops into Claude-Code-style agents without a proxy. If you point a coding agent at a custom provider, the agent must support it: Codex, for example, configures custom providers in config.toml via model_providers entries with base_url, env_key, and wire_api = "responses" (the only wire API it supports), covered in Codex provider configuration.

Benchmarks vs Price: What You Get per Dollar

SWE-bench Verified (real GitHub issues, verified fixes) is the most cited coding benchmark. Scores below are from the llm-stats tracker, September 2, 2026; prices are official API rates. Claude Fable 5.1, Claude Opus 5, and GPT-5.6 have no entry on the tracker yet.

ModelSWE-bench VerifiedInput/MTokOutput/MTokValue ($/MTok out per point)
Claude Fable 595.0%$10.00$50.00$0.53
Claude Mythos Preview (limited)93.9%restrictedrestrictedn/a
Claude Opus 4.888.6%$5.00$25.00$0.28
Claude Opus 4.787.6%$5.00$25.00$0.29
Claude Sonnet 585.2%$2.00$10.00$0.12
DeepSeek V4 Pro Max80.6%open weightsopen weightsopen weights
Gemini 3.1 Pro80.6%$2.00$12.00$0.15
MiniMax M380.5%$0.30$1.20$0.015
Qwen3.7 Max80.4%$2.50$7.50$0.09
Kimi K2.680.2%$0.95$4.00$0.05
The 80% club spans a 10x price range

Five models score between 80.2% and 80.6% on SWE-bench Verified: DeepSeek V4 Pro Max, Gemini 3.1 Pro, MiniMax M3, Qwen3.7 Max, and Kimi K2.6. Within that 0.4-point band, output prices run from $1.20/M (MiniMax M3) to $12/M (Gemini 3.1 Pro), and DeepSeek V4 Pro Max ships it in open weights. Claude Sonnet 5 adds 4.6 points for $10/M output. The next step to Opus 4.8 (88.6%) costs $25/M, and Fable 5 (95.0%) costs $50/M, a 42x step from MiniMax M3. Whether those points are worth it depends on whether your tasks live in the gap.

On the harder SWE-bench Pro (1,865 tasks, 41 professional repos), Scale's standardized public leaderboard tops out at gpt-5.4 (xHigh) 59.10%, Claude Opus 4.6 (thinking) 51.90%, and Gemini 3.1 Pro (thinking) 46.10%. Vendor self-reported aggregates run much higher (llm-stats lists Claude Opus 4.8 at 69.2% and GLM-5.2 at 62.1%), so compare scores only within the same harness.

Read SWE-bench numbers as vendor claims

The SWE-bench Verified figures above are vendor self-reported. The llm-stats tracker lists 102 self-reported results and 0 independently verified, so treat the rankings as vendor claims. Scale's SWE-bench Pro public set runs on a single standardized harness (Pass@1), which is why its numbers are much lower and more directly comparable across models.

LLM API Free Tiers: Exact Amounts

ProviderFree offerLimit
Z.AIGLM-4.7-Flash, GLM-4.5-FlashFree models, no token charge
Alibaba Model Studio1M tokens for new users90-day validity
OpenAI APIFree tier in allowed regions$100/month usage cap
OpenAI CodexIncluded with ChatGPT FreeLowest 5-hour-window message limits
Google Gemini APIFree development tierReduced rate limits
AnthropicNoneTier 1 starts at $5 credit
Morph WarpGrepNone$0.80 per 100K tokens

For experimentation, the practical order is: GLM Flash models (unlimited free), Qwen's 1M tokens, then Gemini's free tier. For coding agents specifically, Codex CLI works with a free ChatGPT sign-in, with the lowest usage limits.

Cost Calculator: Real Workloads

Per-token prices mean nothing until mapped to usage. Two reference workloads, 30-day months, no cache discounts applied (caching reduces all of these).

Coding agent: 50M input / 5M output tokens per day

ModelDaily costMonthly cost
Claude Fable 5.1$750$22,500
GPT-5.5$400$12,000
Claude Opus 5$375$11,250
GPT-5.6 Sol$300$9,000
Kimi K3$225$6,750
Gemini 3.1 Pro (≤200K prompts)$160$4,800
Claude Sonnet 5$150$4,500
GLM-5.3$92$2,760
DeepSeek V4 Flash (peak)$28.60$858
MiniMax M3$21$630
GLM-5.3-Flash$10$300
Morph morph-dsv4flash$6.33$190

Support chatbot: 20M input / 5M output tokens per day

ModelDaily costMonthly cost
Claude Haiku 4.5$45.00$1,350
Gemini 3.5 Flash-Lite$18.50$555
DeepSeek V4 Flash (peak)$15.40$462
MiniMax M3$12.00$360
GPT-5.6 Luna$10.00$300
GLM-5.3-Flash$5.50$165
Morph morph-dsv4flash$3.37$101
The 118x gap on identical traffic

The same coding-agent workload costs $22,500/month on Claude Fable 5.1 and $190/month on Morph's morph-dsv4flash, a 118x gap; against DeepSeek's own peak rate it is 26x. The production answer is rarely either extreme: route routine edits to a cheap model, escalate multi-file refactors to a frontier one, and cache aggressively (DeepSeek cache hits bill input at $0.014/M at peak; Anthropic cache reads are 0.1x and do not count against rate limits). Model the routing split with the LLM cost calculator.

Latency and Throughput

Two metrics matter: time to first token (how fast streaming starts) and tokens per second (how fast it finishes). Provider speed claims vary with load and region, so measure on your own traffic.

Measured medians from public endpoint benchmarks (Artificial Analysis, June 2026). General-purpose chat models cluster in the tens-to-low-hundreds of tokens per second; task-specific models built for one operation run far higher.

ModelOutput tok/s (median)TTFTProvider
Morph Fast Apply10,500sub-secondMorph

Why Morph Fast Apply runs two orders of magnitude above frontier chat models: applying a code edit (merging a lazy edit snippet into the full file) does not need frontier reasoning, so the model is tuned for throughput with codegen-specific speculative decoding. For published per-model output-speed and TTFT figures across the general-purpose providers, see the live Artificial Analysis benchmark.

What the official docs also establish:

  • Reasoning adds latency by design. DeepSeek V4 exposes separate thinking and non-thinking modes so you can opt out per request; Claude Opus 4.8 runs adaptive thinking, and thinking tokens are generated and billed before visible output.
  • OpenAI sells speed explicitly: Fast mode (renamed from priority processing on July 30, 2026) bills a premium over the standard per-token rate.
  • MiniMax sells a priority tier for faster scheduling on M3.
  • Specialized models break the general-purpose ceiling: Morph Fast Apply sustains 10,500 tok/s on code-edit application, two orders of magnitude above frontier chat models, because the task (merging an edit into a file) does not need frontier reasoning.

Specialized APIs: When General-Purpose Falls Short

Coding agents spend most of their compute on two operations: searching codebases for context and applying edits to files. Cognition (the team behind Devin) measured 60% of agent time on search alone. Both operations run through general-purpose LLMs by default, at general-purpose prices and speeds.

Morph Fast Apply

Code-edit application at 10,500 tok/s with 98% accuracy. The agent outputs a lazy edit snippet; Fast Apply merges it into the full file in 1-3 seconds. OpenAI-compatible /v1/chat/completions endpoint.

Morph WarpGrep

RL-trained agentic codebase search: 8 parallel tool calls per turn, 4 turns, sub-6s searches, 0.73 F1. $0.80 per 100K tokens. Ships as an MCP server for any agent.

The pattern generalizes: a frontier model reasons and decides what to change; narrow, fast models execute the mechanical steps. The frontier model's output shrinks (edit snippets instead of whole files), which is exactly the token class that costs $12-30/M. See Fast Apply and WarpGrep.

Best LLM API by Use Case

One ranking does not fit every job. Pick by the constraint that binds you. Each pick below comes from the data already on this page.

Best for reasoning and hard coding

Claude Fable 5.1 and Fable 5 ($10/$50; Fable 5 scores 95.0% on SWE-bench Verified) lead, with Claude Opus 5 ($5/$25) as Anthropic's recommended default and GPT-5.6 Sol ($4/$20) as OpenAI's flagship. On the harder SWE-bench Pro, gpt-5.4 (xHigh) leads Scale's standardized public set at 59.10%. Use these for multi-file refactors and tasks where a cheaper model measurably fails.

Best value

Claude Sonnet 5 ($2/$10) scores 85.2% on SWE-bench Verified at $0.12 of output per benchmark point versus $0.28 for Opus 4.8 and $0.53 for Fable 5. MiniMax M3 ($0.30/$1.20) scores 80.5%, the cheapest model above 80%, at $0.015 per point. DeepSeek V4 Pro Max (80.6%, open weights) matches it for self-hosting.

Best for speed

For code edits, Morph Fast Apply is the uncontested pick: 10,500 tok/s with 98% accuracy, two orders of magnitude above frontier chat models, because merging an edit into a file does not need frontier reasoning. For frontier chat speed, OpenAI and Anthropic both sell paid fast tiers (OpenAI Fast mode, formerly priority processing; Anthropic fast mode on Opus).

Morph Fast Apply: best for code edits

10,500 tok/s, 98% accuracy. The frontier model decides what to change; Fast Apply merges the edit snippet into the full file in 1-3 seconds, behind an OpenAI-compatible endpoint. See /products/compact.

Morph model router

Route each request to the cheapest model that can handle it, roughly 430ms of routing overhead at about $0.001 per request. One key, OpenAI-compatible. See /llm-router.

Best on a budget

Morph morph-dsv4flash ($0.1234375/$0.3475, 1M context, DeepSeek V4 Flash at 16-bit) is the price floor on a paid API, followed by GLM-5.3-Flash ($0.15/$0.50) and GPT-5.6 Luna ($0.20/$1.20). DeepSeek's own V4 Flash is $0.44/$1.32 at peak and $0.22/$0.66 off-peak. GLM-4.7-Flash and GLM-4.5-Flash are free on the Z.AI API. For high-volume simple traffic, start here and escalate only the requests that measurably need a frontier model.

How to Choose an LLM API

Decision framework
  • 1. Establish your quality floor cheaply. Run your real prompts through morph-dsv4flash ($0.3475/M out), GLM-5.3-Flash ($0.50/M), MiniMax M3 ($1.20/M), and GPT-5.6 Luna ($1.20/M). If one passes, you are done at a fraction of frontier cost.
  • 2. Escalate only measured gaps. Move to Claude Sonnet 5 ($10/M), Gemini 3.1 Pro ($12/M), GPT-5.6 Sol ($20/M), Opus 5 ($25/M), or Fable 5.1 ($50/M) for the tasks where the cheap tier measurably fails.
  • 3. Check the constraint that binds you. Rate-limited on day one? Anthropic Tier 1 gives Opus 500K ITPM vs Sonnet's 30K. Need self-hosting or data control? DeepSeek V4, GLM-5.3, Kimi K3, and MiniMax M3 are open weights. Compliance? Bedrock or Foundry.
  • 4. Use specialized APIs for mechanical steps. Edit application, search, embeddings, and reranking all have purpose-built models that beat $15-30/M generalists on both speed and cost.

The most common mistake is anchoring on one provider's flagship and never testing down. Five models clear 80% on SWE-bench Verified for $1.20 to $12 per million output; the frontier (GPT-5.6 Sol, Opus 5, Fable 5.1) costs $20 to $50. The second most common is ignoring caching: at an 80% cache hit rate, Anthropic bills cached reads at 0.1x (0.025x on Fable 5.1) and exempts them from rate limits, and DeepSeek drops cached input to $0.014/M. For long-context workloads, compare windows in detail at LLM context window comparison.

Frequently Asked Questions

What is an LLM API?

An HTTP interface to a hosted large language model. You POST a prompt to an endpoint such as /v1/chat/completions and receive generated tokens, billed per million tokens of input and output. It replaces self-hosted GPU inference with a metered service.

What are the best LLM API providers in 2026?

First-party: Anthropic (Claude Fable 5.1, Opus 5, Sonnet 5), OpenAI (GPT-5.6 Sol, Terra, Luna; GPT-5.5), Google (Gemini 3.1 Pro, 3.6 Flash), DeepSeek (V4 Flash, V4 Pro), Moonshot (Kimi K3, K2.6), Z.AI (GLM-5.3, GLM-5.3-Flash), MiniMax (M3), and Alibaba (Qwen3.8 Max, Qwen3.7 Max). Aggregator: OpenRouter. Cloud resellers: Bedrock and Azure / Microsoft Foundry. Specialized: Morph (open-weight models at 16-bit, Fast Apply, WarpGrep). Which is best depends on the binding constraint: benchmark score (Anthropic), price (Morph, Z.AI, MiniMax), free tier (Z.AI), or compliance (cloud resellers).

What is the cheapest LLM API in 2026?

Among paid APIs, Morph's morph-dsv4flash at $0.1234375/M input and $0.3475/M output (DeepSeek V4 Flash, 1M context, 16-bit), then GLM-5.3-Flash at $0.15/$0.50 and GPT-5.6 Luna at $0.20/$1.20. DeepSeek's own V4 Flash is $0.44/$1.32 at peak, $0.22/$0.66 off-peak, $0.014/M on cache hits. MiniMax M3 ($0.30/$1.20) is the cheapest model above 80% on SWE-bench Verified. GLM-4.7-Flash is free outright.

Which LLM API is best for coding?

Claude Fable 5 (95.0% on SWE-bench Verified, $10/$50) is available again, and Fable 5.1 (September 1, 2026, same price) succeeds it; Claude Opus 4.8 (88.6%, $5/$25) and Claude Sonnet 5 (85.2%, $2/$10) are the next tiers, and Opus 5 ($5/$25) has no published SWE-bench score. On Scale's standardized SWE-bench Pro, gpt-5.4 (xHigh) leads the public set at 59.10%. For budget coding, MiniMax M3 (80.5%, $0.30/$1.20) and open-weights DeepSeek V4 Pro Max (80.6%) are the value picks. For applying edits an agent has already decided on, Morph Fast Apply runs at 10,500 tok/s with 98% accuracy.

What rate limits do LLM APIs have?

Anthropic: tiered by cumulative deposit ($5 to $400); Tier 1 gives Opus 4.x 50 RPM / 500K input tokens per minute, Tier 4 gives 4,000 RPM / 10M ITPM, and cached tokens are exempt. OpenAI: tiers unlock at $5 to $1,000 paid with $100 to $200,000 monthly usage caps; per-model RPM/TPM live on each model page. DeepSeek: concurrency caps of 2,500 (V4 Flash) and 500 (V4 Pro) instead of token budgets.

Which LLM API has the largest context window?

1M tokens is the 2026 flagship standard: Claude Fable 5.1, Opus 5, Sonnet 5, and the legacy Opus 4.x and Sonnet 4.6 (all with no long-context surcharge), GPT-5.6 and GPT-5.5 (1,050,000), Gemini 3.1 Pro (1,048,576), DeepSeek V4 Pro and Flash, Kimi K3 (1,048,576), MiniMax M3 (1,048,576), Qwen3.8 Max, Qwen3.7 Max, and GLM-5.3. Below 1M: Kimi K2.6 and K2.7 Code at 256K, Claude Haiku 4.5 and MiniMax M2.7 at 200K.

Are LLM APIs interchangeable / OpenAI-compatible?

DeepSeek, Moonshot, Z.AI, MiniMax, Qwen, OpenRouter, and Morph accept OpenAI-format requests directly; Google exposes a compatibility endpoint; Anthropic offers an OpenAI SDK compatibility layer over its native Messages API. DeepSeek also accepts Anthropic-format requests. In practice, switching providers is a base URL and key change, with re-testing for tool-calling behavior.

Should I use one provider or multiple?

Multiple, behind a router. Send high-volume simple traffic to a sub-$1.50/M model and escalate hard tasks to a frontier model. The 118x spread between Fable 5.1 and morph-dsv4flash on identical traffic is the budget you are leaving on the table with a single-provider setup. See LLM gateways for the plumbing.

Which LLM API is best by use case?

Reasoning and hard coding: Claude Fable 5.1 ($10/$50), Claude Opus 5 ($5/$25), and GPT-5.6 Sol ($4/$20). Value: Claude Sonnet 5 (85.2%, $2/$10) and MiniMax M3 ($0.30/$1.20), the cheapest model above 80%, with open-weights DeepSeek V4 Pro Max (80.6%) for self-hosting. Speed: Morph Fast Apply for code edits at 10,500 tok/s. Budget: Morph morph-dsv4flash ($0.1234375/$0.3475), GLM-5.3-Flash ($0.15/$0.50), or the free GLM Flash models on Z.AI. Most teams route across these with the Morph model router.

Sources

Every price and benchmark on this page traces to a primary source, checked September 2, 2026 (rate-limit tiers were last checked June 28, 2026):

Related Resources

Code Editing at 10,500 tok/s

Frontier LLM APIs bill $12-30 per million output tokens to rewrite whole files. Morph Fast Apply merges edit snippets into files at 10,500 tok/s with 98% accuracy, behind an OpenAI-compatible API.