Qwen3.8-Max pricing
Flagship LLM from Alibaba (Qwen) —$2.00 input and$6.00 output per 1M tokens.
GAFlagshipHost: Alibaba Model Studio (Singapore)
- Input /1M
- $2.00
- Output /1M
- $6.00
- Cached /1M
- $0.25
- Blended
- $4.00
- Context
- 1M
- Max output
- 131.1K
API model id: qwen3.8-max
Full price matrix
| Rate | Input /1M | Output /1M | Notes |
|---|---|---|---|
| Standard | $2.00 | $6.00 | |
| Cached input (read) | $0.25 | — | Prompt-cache hit |
List prices, USD, directional. Rates are provider list prices per 1M tokens and are meant for comparison, not billing. Batch rates shown as 50% off are derived where a provider offers batch but does not publish a separate figure. Preview, promo, intro, peak/off-peak, long-context, and third-party-host prices are labeled where they apply. Token counts vary by tokenizer, so per-token price is not always a like-for-like cost. Always confirm with the provider before relying on a number. Prices as of 2026-09-22.
Source: https://www.alibabacloud.com/help/en/model-studio/qwen3-8-max
Pricing notes
- Regional pricing: Priced on Alibaba Model Studio (Singapore); prices shown are the Singapore region — other regions differ.
- Alibaba's recommended flagship. One flat rate up to 1M tokens, with thinking and non-thinking output billed the same. Cached input shown is the implicit (automatic) cache rate, $0.25; explicit cache costs $2.50 to create and $0.17 to read. Alibaba lists the Qwen3.8 models as explicit exceptions to its usual cache rates (implicit 20% / explicit read 10% of input), which is why the implicit-cache rate here sits above the explicit-cache read. Batch is not supported in the Singapore region.