Qwen3.8-Omni-Flash Undercuts Gemini Flash on Price
Alibaba's new omnimodal model prices at $0.15/$0.47 per million tokens, far below Gemini 3.8 Flash, with benchmarks that are close but mixed.
Alibaba's new omnimodal model prices at $0.15/$0.47 per million tokens, far below Gemini 3.8 Flash, with benchmarks that are close but mixed.
Alibaba's Qwen Team Launches an Agent-Focused Omnimodal Model
Alibaba's Qwen team announced Qwen3.8-Omni-Flash in an official blog post, "Qwen3.8-Omni-Flash: Omni Senses. Agentic Delivery." The model is now live and accessible through Qwen Studio, Qwen Cloud, and the API (Alibaba Cloud Model Studio / DashScope), with regional endpoints in China (Beijing), Singapore, Hong Kong, Tokyo, Frankfurt, and US (Virginia). It is Qwen's next-generation native omnimodal model, built specifically to strengthen agentic capabilities — planning tasks, calling tools, and completing creative work with audio and video, not just describing what's in them.
What the model does
Qwen3.8-Omni-Flash accepts text, image, audio, and video input with a 1-million-token context window, while Qwen says it maintains text performance comparable to a text-only model of the same size. Notably, the base Qwen3.8-Omni-Flash API is text-out only: Alibaba Cloud's own model documentation lists its output modality as text, even though inputs span text, image, audio, and video. Alongside the base model, Qwen introduced a separate real-time endpoint, Qwen3.8-Omni-Flash-Realtime, purpose-built for continuous, low-latency interaction: Alibaba's documentation describes it as processing streaming audio and image input (including video frames) and generating both text and audio responses in real time over WebSocket or WebRTC, extending the model from "understanding content" to "participating in an interaction." According to Qwen's own realtime performance table, time to first audio packet ranges from about 978 milliseconds for short 6-second audio inputs up to roughly 1.35 seconds for 20-second audio-visual inputs, with an audio generation speed of roughly 6-7x real time.
Qwen highlights several concrete workflows built on the model: editing and summarizing long-form video, producing music videos by understanding a song's rhythm and lyrics well enough to time visuals to it, generating film commentary, and running audio-visual "deep research" reports. Two open-source tools ship alongside the model to support these workflows: Qwen-MM-Plugins, a multimodal plugin suite that adds perception, tool use, and workflow execution to existing agent harnesses including Claude Code, Gemini CLI, Qwen Code, Codex, Qoder, CodeBuddy, and OpenClaw; and Qwen-Live Harness, a native runtime for real-time omnimodal interaction that supports task delegation and long-term memory. Notably, while those two tools are open-sourced on GitHub, Qwen3.8-Omni-Flash's own model weights are not — unlike several prior Qwen releases, there is no HuggingFace model card or downloadable checkpoint for this model; it's accessible only through Qwen's hosted platform and API.
Pricing: a real, verifiable gap
The clearest advantage in this release is price. Alibaba Cloud's own Model Studio pricing page lists Qwen3.8-Omni-Flash's international rate at $0.15 per million input tokens, $0.016 per million cache-hit input tokens, and $0.47 per million output tokens. Google's published pricing for Gemini 3.8 Flash is $0.75 per million input tokens and $3.75 per million output tokens at its current introductory rate — roughly 5x and 8x higher than Qwen's input and output prices, respectively — and that Gemini rate is set to double, to $1.50/$7.50 per million tokens, once the introductory window ends on January 1, 2027. Qwen also reports that, compared with its own predecessor Qwen3.5-Omni-Plus, hourly API cost for audio input is down more than 98% and for audio-visual input down more than 93%, using a standardized methodology (720p video at 1 frame per second, matched API settings across vendors) disclosed in the blog post's footnotes.
Benchmarks: genuinely close, not a clean sweep
Qwen's own framing is careful and worth quoting precisely: it describes "audio-visual performance close to Gemini 3.8 Flash and overall audio performance that exceeds Gemini 3.8 Flash" — not a claim of beating Gemini across the board. The published benchmark tables bear that out as a mixed picture rather than a rout in either direction. On agentic multimodal tool use, Qwen3.8-Omni-Flash scores 71.0 on WildClawBench-MM against Gemini 3.8 Flash's 58.9, a clear lead. On general audio tasks, the gap is much larger: on the AliMeeting multi-speaker transcription benchmark, Qwen posts a diarization error rate of 3.4 versus Gemini's 72.6 (lower is better), a dramatic improvement that anchors Qwen's "exceeds on audio" claim. But on several audio-visual reasoning benchmarks, Gemini leads: Video-MME-v2 (71.0 vs. Qwen's 65.0), LVOmniBench (70.7 vs. 63.3), and AgenticVBench (45.0 vs. Qwen's 36.8) — though those are the static, one-shot numbers. Qwen also publishes a second table running both models through its Qwen Code agent harness, where the gap narrows or reverses: Qwen3.8-Omni-Flash reaches 73.6 on LVOmniBench against Gemini's 70.7, and 71.3 on Video-MME-v2 against Gemini's 72.7. Given that agentic use is the whole point of this release, that second table is arguably the more relevant comparison. It's also worth separating Qwen's two kinds of comparisons in the blog post: figures like "+25% average score across 29 evaluations" and "+36.5 points on WildClawBench-MM" describe improvement over Qwen's own prior model, Qwen3.5-Omni-Plus, not a lead over Gemini — a distinction the source material is explicit about but that's easy to blur in secondary coverage.
Usability
For developers, the realtime API is the more novel piece: a documented Python example using DashScope's WebSocket endpoint streams microphone audio and webcam video (at 1 fps by default) to the model and gets back streaming audio and transcript deltas, aimed at building live conversational agents rather than one-shot API calls. The model also supports an adjustable reasoning_effort parameter to trade off latency and cost, and broad language coverage — 74 languages plus 39 Chinese dialects for speech recognition, and 29 languages plus 7 dialects for speech generation — which matters for the transcription and translation workflows Qwen is targeting.
Pros and cons
The price advantage here is not a marketing estimate; it's confirmed directly from Alibaba Cloud's own published rate card against Google's own published Gemini pricing, and it holds up as a genuine multiple rather than a marginal discount. The model also backs up its agentic and general-audio claims with concrete wins, particularly the large margin on multi-speaker transcription. Shipping open-source integration tooling (Qwen-MM-Plugins, Qwen-Live Harness) that plugs into existing agent harnesses lowers the integration barrier meaningfully compared with a bare API.
The tradeoffs are real, too. This is not an open-weight release — despite Qwen's track record of open-sourcing flagship models, Qwen3.8-Omni-Flash itself is API-only, which limits self-hosting and fine-tuning options that some previous Qwen releases offered. The "close to Gemini" framing is honest but does mean Gemini 3.8 Flash retains a lead on several audio-visual reasoning benchmarks, including long-video understanding. And every comparison so far comes from Qwen's own published evaluation; independent third-party benchmark results aren't yet available to cross-check these numbers.
Outlook
Qwen3.8-Omni-Flash is a clear signal that Alibaba is competing on cost and agent integration rather than trying to claim outright superiority over Gemini's multimodal reasoning. If the realtime API's sub-1.5-second latency holds up under real production load, and the open-source plugin ecosystem gets adopted inside popular agent harnesses, the price gap alone could make this an attractive default for cost-sensitive, high-volume audio and video agent workloads — even without open weights.
Conclusion
Qwen3.8-Omni-Flash earns its headline: a verifiable, multiple-fold price advantage over Gemini 3.8 Flash, backed by real (if uneven) benchmark performance and genuinely useful open-source tooling for agent developers. It's best suited for teams building cost-sensitive audio/video agents who don't need open weights or best-in-class audio-visual reasoning specifically — for that narrower need, Gemini 3.8 Flash still leads on several of its own benchmarks.
Editor's Verdict
Qwen3.8-Omni-Flash Undercuts Gemini Flash on Price earns a solid recommendation within the Other LLM space.
The strongest case for paying attention: independently confirmed price advantage of roughly 5-8x over Gemini 3.8 Flash's introductory per-token rate, verified directly against both companies' own pricing pages. That alone raises the bar for what readers should expect in this space. Reinforcing that, decisive lead on general audio benchmarks, including multi-speaker transcription (AliMeeting DER 3.4 vs. Gemini's 72.6) and on WildClawBench-MM multimodal tool use — practical value rather than just headline appeal. The broader signal worth registering is straightforward: the vendor's own framing is deliberately narrow — "close to Gemini 3.8 Flash" on audio-visual performance and "exceeds" only on general audio tasks, not a claim of beating Gemini across the board. On the other side of the ledger, one constraint is real rather than a marketing footnote: not open-weight — unlike several previous flagship Qwen releases, no HuggingFace model card or downloadable checkpoint exists; access is API-only. It should factor into any serious decision. Layered on top of that, parity is not universal — Gemini 3.8 Flash still leads on several audio-visual reasoning benchmarks in the static setting (Video-MME-v2, LVOmniBench, AgenticVBench), though Qwen closes much of that gap when run in agent mode — which narrows the set of teams for whom this is an obvious yes.
For multi-model deployment teams, cost-conscious operators, and developers willing to evaluate beyond the major labs, this is a serious evaluation candidate, not just a curiosity to bookmark. For everyone else, the safer posture is to monitor coverage and revisit once the use cases that matter to your team are demonstrated in the wild.
Pros
- Independently confirmed price advantage of roughly 5-8x over Gemini 3.8 Flash's introductory per-token rate, verified directly against both companies' own pricing pages
- Decisive lead on general audio benchmarks, including multi-speaker transcription (AliMeeting DER 3.4 vs. Gemini's 72.6) and on WildClawBench-MM multimodal tool use
- 1M-token context window supports genuinely long-form audio and video workflows rather than short clips
- Ships with open-source integration tooling (Qwen-MM-Plugins, Qwen-Live Harness) that plugs into existing agent harnesses like Claude Code and Gemini CLI
- Realtime API delivers sub-1.5-second time-to-first-audio-packet, making live conversational use cases practical
Cons
- Not open-weight — unlike several previous flagship Qwen releases, no HuggingFace model card or downloadable checkpoint exists; access is API-only
- Parity is not universal — Gemini 3.8 Flash still leads on several audio-visual reasoning benchmarks in the static setting (Video-MME-v2, LVOmniBench, AgenticVBench), though Qwen closes much of that gap when run in agent mode
- All benchmark comparisons come from Qwen's own published evaluation, with no independent third-party results yet available to cross-check
- Text output only on the base API — Alibaba Cloud's own documentation lists Qwen3.8-Omni-Flash's output modality as text; only the separate Qwen3.8-Omni-Flash-Realtime endpoint adds generated audio for live voice conversations
References
Comments0
Key Features
Qwen3.8-Omni-Flash is a native omnimodal model accepting text, image, audio, and video input with a 1M-token context window, paired with a low-latency Qwen3.8-Omni-Flash-Realtime endpoint for live camera-and-microphone agent interaction over WebSocket/WebRTC. It supports 74 languages for speech recognition and 29 for speech generation, plus dozens of Chinese dialects. Two open-source companion tools ship alongside it: Qwen-MM-Plugins, which adds multimodal perception and workflow execution to agent harnesses like Claude Code and Gemini CLI, and Qwen-Live Harness, a runtime for continuous real-time interaction. The model itself, however, is API-only — its weights are not published on HuggingFace.
Key Insights
- The vendor's own framing is deliberately narrow — "close to Gemini 3.8 Flash" on audio-visual performance and "exceeds" only on general audio tasks, not a claim of beating Gemini across the board.
- The price gap is independently verifiable: Alibaba Cloud's pricing page lists $0.15/$0.47 per million tokens versus Google's published $0.75/$3.75 introductory rate for Gemini 3.8 Flash, a roughly 5-8x difference that widens further once Gemini's intro pricing ends January 1, 2027.
- Many of the biggest reported point gains (+25% average score, +36.5 points on WildClawBench-MM) describe improvement over Qwen's own predecessor, Qwen3.5-Omni-Plus, not a lead over Gemini.
- On individual benchmarks the picture is mixed: Qwen leads decisively on multi-speaker transcription (AliMeeting DER 3.4 vs. Gemini's 72.6) and on WildClawBench-MM tool use, while Gemini leads in the static setting on Video-MME-v2, LVOmniBench, and AgenticVBench.
- Despite Qwen's history of open-sourcing flagship models, Qwen3.8-Omni-Flash's weights are not published on HuggingFace — only the surrounding agent tooling (Qwen-MM-Plugins, Qwen-Live Harness) is open-source.
- The realtime API reports time-to-first-audio-packet under 1.4 seconds even for 20-second audio-visual inputs, aimed at live conversational agent use cases rather than batch processing.
- Qwen discloses its pricing comparison methodology in footnotes (720p at 1fps, matched API parameters, fixed USD-to-CNY conversion rate), a level of transparency that makes the price claim easier to verify than most vendor comparisons.
Was this review helpful?
Share
Related AI Reviews
Meta Expands Muse Agent With Video Chat, Glasses, Charm
Meta expanded its Muse AI agent at Connect 2026 with video chat, smart glasses, Mac control, and a Muse Charm wearable shipping in December.
Grok 4.7 Review: Same Price, Trails Fable 5.1 and GPT-6
xAI's Grok 4.7 keeps Grok 4.6's $2/$6 pricing on a larger base model, but Artificial Analysis scores it 46 versus 53 for Fable 5.1 and GPT-6.
TypeSafe AI Launches Jev, a New System One Model
Ex-OpenAI researcher Diogo Almeida's TypeSafe AI released Jev in early access, a structured-decision model priced at $0.042 per million input tokens.
PrismML's Bonsai 2 27B Shrinks Qwen3.8 27B by More Than 9x
PrismML's Bonsai 2 27B compresses Qwen3.8 27B to 5.9GB, over 9x smaller, while retaining 98.2% of its full-precision benchmark performance.
