Qwen3.8-Omni-Flash Cuts Audio API Cost by 98 Percent and Adds 1-Million-Token Multimodal Context
Alibaba's Qwen team released Qwen3.8-Omni-Flash on September 18, 2026. Per Alibaba's own estimates, the model prices audio input at more than 98 percent below what Qwen3.5-Omni-Plus cost per hour, a cost reduction that changes the economics of audio and video analysis pipelines at the API layer.
What
Qwen3.8-Omni-Flash accepts text, images, audio, and video through the Chat Completions and Responses APIs with a 1-million-token context window, per TechNode. Its native output is text only. Built-in function calling and web search let the model plan multistep tasks and call external tools for work that goes beyond analysis.
International pricing is $0.15 per million input tokens, $0.47 per million output tokens, and $0.016 per million cache-hit input tokens, per Alibaba Cloud's pricing page. Alibaba says the estimated API cost per hour of audio input runs more than 98 percent below Qwen3.5-Omni-Plus, with audiovisual input down more than 93 percent. Both figures use Alibaba's own methodology: two minutes of source material priced and scaled to 60 minutes, with video sampled at 720p and one frame per second, per RuntimeWire. Independent testing has not confirmed either figure.
The model is available in six regions: Beijing, Singapore, Hong Kong, Tokyo, Frankfurt, and Virginia, per Alibaba Cloud's model documentation.
Repeated passes over long recordings get much cheaper
The cache-hit rate of $0.016 per million tokens is roughly one-tenth the standard input rate. For audio and video agents that query the same long recording multiple times, such as a meeting-analysis pipeline running several analytical passes over the same session, the cost of repeated inference drops substantially relative to re-sending the full token representation each time.
The 98 percent cost reduction is Alibaba's own figure, not an independent benchmark, so the real-world savings depend on the specific workload. The published token rates are verifiable and competitive. Builders currently using OpenAI or Google endpoints for audio understanding now have a direct cost comparison point. One constraint to account for: text-only output means the model functions as an understanding and routing layer, so teams still need separate tools to generate any media the analysis triggers.
Context and reactions
Alibaba released the open-source Qwen-MM-Plugins repository for multimodal agent workflows alongside the model, plus Qwen-Live Harness for continuous audiovisual interaction. Per Alibaba, the model improved average scores by more than 25 percent across 29 evaluations compared to Qwen3.5-Omni-Plus, per TechNode. The Qwen blog post announcing the release was marked "draft" as of September 18; Alibaba Cloud's production documentation was updated the same day and lists the model as available.
What to watch next
Independent benchmarks comparing Qwen3.8-Omni-Flash against GPT-6 Astra and Gemini 3.8 Flash on audio-visual agent tasks would clarify whether the pricing efficiency carries through to task-completion cost. A finalized Qwen blog post with full technical documentation is the next expected update from Alibaba.
Sources
- Alibaba's Qwen releases Qwen3.8-Omni-Flash with 1M-token context: TechNode, September 18, 2026
- Alibaba ships Qwen3.8-Omni-Flash to watch, listen and call tools: RuntimeWire, September 17, 2026
- Alibaba Qwen releases Qwen3.8-Omni-Flash: a 1M-context omni-modal model: MarkTechPost, September 18, 2026
