Free cloud & LLM pricing tools — AWS EC2 + LLM API Pricing Explorer

Long-Context Pricing

A surcharge that raises the per-token rate once a prompt crosses a token threshold — common on large-window models.

Long-context pricing means the per-token rate goes up once a single prompt exceeds a set size. Instead of one flat rate across the whole context window, the provider charges a base rate up to a threshold and a higher rate above it. Google’s Gemini Pro models, for instance, price prompts up to 200K tokens at a base rate and multiply input and output above that; xAI’s Grok 4.5 doubles both rates past 200K tokens.

The rationale is that very long prompts are more expensive to serve, so the surcharge kicks in exactly when you lean on the model’s biggest capability. It means a 1M-token window is not uniformly priced — the last 800K tokens can cost markedly more than the first 200K. OpenAI applies a surcharge above 272K tokens on its largest-window flagships as well, though the exact multiplier is not published; this explorer notes that rather than guessing it.

This explorer models the threshold and multipliers where a provider publishes them, showing both the base and above-threshold rates on the model page, and adds a note where the surcharge exists but the numbers are unconfirmed. If your workload routinely sends very large prompts, the effective rate is the long-context rate, not the headline — factor it in before assuming a big window is cheap.

FAQ

Which models have long-context surcharges?

Among those here, Google’s Gemini Pro models (above 200K tokens) and xAI’s Grok 4.5 (rates double above 200K) publish explicit tiers; OpenAI notes a surcharge above 272K tokens on its largest flagships.

Does the surcharge apply to the whole prompt or just the excess?

It depends on the provider — read the model page and the provider’s docs. Either way, a large prompt makes the higher tier the rate that actually governs your bill.

More LLM pricing terms

← Back to the LLM pricing hub