Qwen

Qwen3.8-Omni-Flash

An omnimodal model for audio and video understanding, content analysis, and tool-based workflows. Accepts text, images, audio, and video and generates text with adjustable thinking.

ActiveReleased September 17, 2026
Tool callingCan invoke external tools and APIs (function calling) mid-response.Structured outputCan return JSON or schema-constrained responses instead of free text.ReasoningUses extended thinking / chain-of-thought for harder multi-step problems.Effort controlLets you dial thinking depth (low / medium / high) to trade cost for quality.

Source: Official model limits, modalities, tools, and reasoning controls, First documented Model Studio API availability: September 17, 2026, Implicit caching minimum and variable retention, JSON Object output support in thinking and non-thinking modes; strict JSON Schema not established, Thinking token billing

Price per 1M tokens via Alibaba (OpenRouter)
$0.15 in / $0.47 out
Context window1,000,000
Max output131,072
Cache read$0.016 / M
Min cache prefix1,024 tokens
Input
  • Text
  • Image
  • Audio
  • Video
Output
  • Text

Pricing

Every billing leg from the live catalog, in USD per 1M tokens.

Billing legRate
Input$0.15 / M
Output$0.47 / M
Cache read$0.016 / M

Cost calculator

Billing-grade math over the rate card above, including the thinking tokens most estimates miss.

Hidden reasoning tokens bill as output. 1.0x = thinking off. Medium-effort reasoning typically lands near 3x; heavy reasoning runs 6x to 8x. Set it to match your workload.

Estimated monthly cost
$58.20
at current list prices
Fresh input$30.00
Cached input$0.00
Output incl. thinking$28.20
Track this automatically →

Provider

Where this model runs, and what it charges.

PlatformInput / MOutput / MCache read / MBatchAccess
Alibaba
$0.15$0.47$0.016-via OpenRouter

Price history

Has this model gotten cheaper or more expensive?

No changePriced at $0.15 / $0.47 per 1M since September 23, 2026.

More from Qwen

Other models from the same provider.

Stop estimating. Track it.

SuperPenguin meters every Qwen3.8-Omni-Flash call your team makes, prices it with this exact rate card, and shows spend by feature, customer, and team.

Get started free

Rates are generated from SuperPenguin's live pricing catalog.