Qwen3.7-Plus pricing
Mid-range LLM from Alibaba (Qwen) —$0.40 input and$1.60 output per 1M tokens.
GAMid-rangeHost: Alibaba Model Studio (Singapore)Long-context tierPromo available
- Input /1M
- $0.40
- Output /1M
- $1.60
- Cached /1M
- $0.08
- Blended
- $1.00
- Context
- 1M
- Max output
- 131.1K
API model id: qwen3.7-plus
Full price matrix
| Rate | Input /1M | Output /1M | Notes |
|---|---|---|---|
| Standard | $0.40 | $1.60 | |
| Cached input (read) | $0.08 | — | Prompt-cache hit |
| Long context (> 256,000 tokens) | $1.20 | $4.80 | Input ×3, output ×3 |
List prices, USD, directional. Rates are provider list prices per 1M tokens and are meant for comparison, not billing. Batch rates shown as 50% off are derived where a provider offers batch but does not publish a separate figure. Preview, promo, intro, peak/off-peak, long-context, and third-party-host prices are labeled where they apply. Token counts vary by tokenizer, so per-token price is not always a like-for-like cost. Always confirm with the provider before relying on a number. Prices as of 2026-09-29.
Source: https://www.alibabacloud.com/help/en/model-studio/qwen3-7-plus
Pricing notes
- Long-context pricing: above 256,000 tokens, input is billed at ×3 ($1.20) and output at ×3 ($4.80). See long-context pricing.
- Promotion: Alibaba's pricing page shows a 20% limited-time discount on this model; the list price is shown here — verify the current rate..
- Regional pricing: Priced on Alibaba Model Studio (Singapore); prices shown are the Singapore region — other regions differ.
- Alibaba's recommended Plus model. Thinking and non-thinking output are billed the same. Above 256K input tokens, input and output triple ($1.20 / $4.80). Cached input shown is the implicit-cache rate (20% of input); explicit cache costs $0.50 to create and $0.04 to read. Batch is not supported in the Singapore region.