Qwen3.8-Omni-Flash: Alibaba's New Model Cuts AI Audio and Video Costs by Up to 98%
Alibaba's Qwen team launched a native omni-modal model with 1M-token context and far cheaper audio/video processing for AI agents.
Most of the cost in building a voice agent, a video-summarizing workflow, or any AI system that has to watch, listen, and respond in real time isn't the reasoning — it's the audio and video processing underneath it. On September 18, Alibaba's Qwen team released Qwen3.8-Omni-Flash, a native omni-modal model that handles text, images, audio, and video in a single workflow, and cut the price of that processing by as much as 98%. For agencies and teams building automation on top of multimodal AI, that's a more consequential number than any benchmark score.
Qwen3.8-Omni-Flash is a native omni-modal model, meaning it processes text, images, audio, and video together rather than bolting separate encoders onto a text model. It supports a 1-million-token context window and is available now through the Qwen AI platform's API. Alibaba reports the model improved its average score by more than 26% across 30 evaluations compared with its predecessor, Qwen3.5-Omni-Plus, with specific gains in audio-video agent tasks, coding, long-context handling, and real-time multimodal interaction. Alibaba positions its audio-video capability as approaching Google's Gemini 3.8 Flash, with audio performance reported to exceed it — though Gemini reportedly still leads on pure video-reasoning benchmarks.
Benchmark gains are easy to find in every model release this year. What's less common is a price cut this steep: Alibaba cut the API price for audio input processing by more than 98%, and audio-visual input by more than 93%, with input pricing available for as low as RMB 0.8 (roughly $0.11) per million tokens. For any team running production workloads that involve transcription, meeting analysis, call review, or video content processing, that's the difference between a multimodal AI feature being a line item and being a rounding error.
Alongside the model, Alibaba released Qwen-MM-Plugins and a Qwen-Live Harness, aimed specifically at long-running and real-time agentic workflows rather than single-turn chat responses. The model supports long-video analysis, automated meeting summaries, video research tasks, and multimodal tool use — the building blocks for agents that need to sit inside a longer workflow rather than answer one question and stop.
Alibaba's own demonstration is a useful signal of what "agentic" means here in practice: Qwen3.8-Omni-Flash was used to autonomously improve a smaller model's (Qwen2.5-Omni-3B) recognition of Sichuan-dialect speech, cutting the character error rate from 25.79% to 15.30% within twelve hours by building 3,413 training samples across four self-directed rounds. That's a model being used to run its own improvement loop on a narrow, real-world problem — not just answering prompts.
For automation and no-code builders, the practical use case is clear: voice-driven customer support agents, automated meeting-notes pipelines, video content moderation or tagging, and any workflow that currently avoids audio/video AI because of cost now has a meaningfully cheaper option to prototype against. The 1M-token context window also matters for anything that needs to hold a long call transcript, a full video, or an extended multi-turn agent session in memory at once.
The catch, for teams that prefer to self-host or fine-tune: weights for Qwen3.8-Omni-Flash have not been open-sourced, unlike some earlier Qwen releases. Access is API-only for now, so this is a build-on-top-of-the-vendor decision rather than a run-it-yourself one. Teams that need on-premises or fully open-weight multimodal models will want to keep watching for a possible later open release, as Alibaba has done with other models in the Qwen family.
Our cutting-edge features simplify collaboration and creativity, making your workflow intuitive and efficient. Transform your vision into reality effortlessly with Hadidiz Flow.



