Qwen3.8-Omni-Flash
An omnimodal model for audio and video understanding, content analysis, and tool-based workflows. Accepts text, images, audio, and video and generates text with adjustable thinking.
Source: Official model limits, modalities, tools, and reasoning controls, First documented Model Studio API availability: September 17, 2026, Implicit caching minimum and variable retention, JSON Object output support in thinking and non-thinking modes; strict JSON Schema not established, Thinking token billing
- Text
- Image
- Audio
- Video
- Text
Pricing
Every billing leg from the live catalog, in USD per 1M tokens.
| Billing leg | Rate |
|---|---|
| Input | $0.15 / M |
| Output | $0.47 / M |
| Cache read | $0.016 / M |
Cost calculator
Billing-grade math over the rate card above, including the thinking tokens most estimates miss.
Hidden reasoning tokens bill as output. 1.0x = thinking off. Medium-effort reasoning typically lands near 3x; heavy reasoning runs 6x to 8x. Set it to match your workload.
Provider
Where this model runs, and what it charges.
| Platform | Input / M | Output / M | Cache read / M | Batch | Access |
|---|---|---|---|---|---|
Alibaba | $0.15 | $0.47 | $0.016 | - | via OpenRouter |
Price history
Has this model gotten cheaper or more expensive?
More from Qwen
Other models from the same provider.
Stop estimating. Track it.
SuperPenguin meters every Qwen3.8-Omni-Flash call your team makes, prices it with this exact rate card, and shows spend by feature, customer, and team.
Rates are generated from SuperPenguin's live pricing catalog.