Back to list
Jul 20, 2026
25
0
0
Other LLMNEW

Qwen3.8-Max-Preview: Alibaba's 2.4T Multimodal Model Launches at WAIC

Alibaba previewed Qwen3.8-Max, a 2.4-trillion-parameter multimodal MoE model, at WAIC 2026. It handles text, images, video, and documents with a 1M-token context, though key claims remain unverified.

#Qwen#Alibaba#LLM#Multimodal#Mixture of Experts
Qwen3.8-Max-Preview: Alibaba's 2.4T Multimodal Model Launches at WAIC
AI Summary

Alibaba previewed Qwen3.8-Max, a 2.4-trillion-parameter multimodal MoE model, at WAIC 2026. It handles text, images, video, and documents with a 1M-token context, though key claims remain unverified.

Introduction

At the World AI Conference (WAIC) in Shanghai, Alibaba's Qwen team unveiled Qwen3.8-Max-Preview on July 19, 2026. The model is a sparse Mixture-of-Experts (MoE) system that Alibaba says contains 2.4 trillion parameters. This figure is Alibaba's own disclosure; it has not been independently audited or verified. Qwen3.8-Max-Preview is the first Qwen model above one trillion parameters to accept multimodal input, processing text, images, video, and documents in a single request. The preview arrived days after a rival open-weight launch from another Chinese lab, placing Qwen3.8-Max-Preview in a competitive window for large-scale LLM releases.

Feature Overview

Qwen3.8-Max-Preview retains the 1-million-token context window first introduced with the prior Qwen3.7-Max, allowing it to process long documents, codebases, or video transcripts without chunking. Its sparse MoE architecture routes each token through a subset of expert subnetworks rather than the full parameter set, a design intended to control inference cost at this scale. Alibaba has not disclosed how many parameters are actually active per token, an important figure for estimating real-world serving costs. The model exposes both OpenAI-compatible and Anthropic-compatible API protocols, which should ease migration for developers already building against those ecosystems. Multimodal support now spans four input types, text, images, video, and documents, making it Alibaba's most broadly capable model released to date. Alongside the preview, Alibaba's own marketing describes the model as "second only to Fable 5," a claim not accompanied by any published benchmark table, model card, or third-party evaluation as of this writing.

Usability Analysis

Qwen3.8-Max-Preview is live now through Alibaba Cloud's Token Plan, as well as through Qoder and QoderWork. During the preview period, pricing sits at 10% of standard rates, letting developers test the 1M-token context and multimodal pipeline at a fraction of the eventual cost. For reference, the prior Qwen3.7-Max baseline charged $1.25 per million input tokens and $3.75 per million output tokens; Qwen3.8-Max-Preview's discounted pricing should undercut that substantially, though Alibaba has not published an exact preview rate card beyond the stated 10% figure. The OpenAI- and Anthropic-compatible endpoints mean teams already using either ecosystem's SDKs can point existing code at the new model with minimal rewiring, lowering the barrier to evaluation.

Pros and Cons

Pros:

  • Native multimodal handling of text, images, video, and documents in one model
  • 1M-token context window supports large-document and long-video workflows
  • Dual API compatibility (OpenAI and Anthropic) simplifies integration for existing developers
  • Steep preview discount (10% of standard rates) lowers the cost of evaluation

Cons:

  • The "second only to Fable 5" ranking is Alibaba's own unverified marketing claim, with no published benchmarks or independent evaluation supporting it
  • Active parameters per token are undisclosed, leaving real inference cost opaque
  • The promised open-weight release has no confirmed date, license terms, or Hugging Face repository
  • The headline 2.4-trillion-parameter figure is self-reported and not independently audited

Outlook

Alibaba has said a full open-weight release is coming "soon," but as of the WAIC announcement there is no confirmed timeline, license, or repository. If that release arrives with a transparent model card and reproducible benchmarks, outside researchers could finally verify Alibaba's performance claims. Until then, Qwen3.8-Max-Preview's standing relative to other frontier models rests on Alibaba's own word rather than public evidence. The broader industry trend, larger MoE models with multimodal input and million-token context, continues to accelerate, and Qwen3.8-Max-Preview extends Alibaba's participation in that race.

Conclusion

Qwen3.8-Max-Preview is a notable technical step for the Qwen family, adding multimodal input to a trillion-plus-parameter model while carrying over the 1M-token context window from Qwen3.7-Max. Developers interested in large-scale multimodal context handling, and comfortable with preview-stage pricing and unverified claims, have reason to test it now at a steep discount. Technical decision-makers should treat the "second only to Fable 5" claim and the 2.4-trillion-parameter figure as marketing until Alibaba publishes benchmarks or open weights. Rating: 3/5, reflecting genuine technical capability alongside a meaningful transparency gap.

Editor's Verdict

Qwen3.8-Max-Preview: Alibaba's 2.4T Multimodal Model Launches at WAIC is a workable proposition that fills a clear gap, even if it doesn't fundamentally change the landscape.

The strongest case for paying attention is native multimodal support for text, images, video, and documents in a single model, which raises the bar for what readers should now expect from peers in this space. Reinforcing that, 1M-token context window enables large-document and long-video processing adds practical value rather than just headline appeal. The broader signal worth registering is straightforward: qwen3.8-Max-Preview is the first Qwen model over 1 trillion parameters to support multimodal input across text, images, video, and documents. On the other side of the ledger, the 'second only to Fable 5' performance claim is self-reported and unverified, with no published benchmarks is a real constraint, not a marketing footnote, and it should factor into any serious decision. Layered on top of that, active parameters per token are undisclosed, making real-world inference cost difficult to estimate narrows the set of teams for whom this is an obvious yes.

For multi-model deployment teams, cost-conscious operators, and developers willing to evaluate beyond the major labs, the smart move is to track its trajectory and revisit once the rough edges are filed down. For everyone else, the safer posture is to monitor coverage and revisit once the use cases that matter to your team are demonstrated in the wild.

Pros

  • Native multimodal support for text, images, video, and documents in a single model
  • 1M-token context window enables large-document and long-video processing
  • OpenAI- and Anthropic-compatible APIs simplify integration for existing developer workflows
  • Steep 10% preview discount lowers the barrier to evaluation

Cons

  • The 'second only to Fable 5' performance claim is self-reported and unverified, with no published benchmarks
  • Active parameters per token are undisclosed, making real-world inference cost difficult to estimate
  • The promised open-weight release has no confirmed date, license terms, or repository
  • The 2.4-trillion-parameter figure itself is Alibaba's own disclosure and not independently audited

Comments0

Key Features

1. 2.4-trillion-parameter sparse Mixture-of-Experts architecture (Alibaba's own disclosed figure, unaudited) 2. First Qwen model above 1 trillion parameters with multimodal input: text, images, video, and documents 3. 1-million-token context window, carried over from the prior Qwen3.7-Max 4. OpenAI-compatible and Anthropic-compatible API protocols for easier migration 5. Preview-period pricing at 10% of standard rates via Alibaba Cloud's Token Plan, Qoder, and QoderWork

Key Insights

  • Qwen3.8-Max-Preview is the first Qwen model over 1 trillion parameters to support multimodal input across text, images, video, and documents
  • The 2.4-trillion-parameter figure is Alibaba's own claim and has not been independently verified or audited
  • Alibaba's claim that the model is 'second only to Fable 5' is unverified marketing with no published benchmark table or model card
  • Active parameters per token in the MoE architecture remain undisclosed, obscuring true inference cost
  • The model carries over the 1M-token context window from the prior Qwen3.7-Max rather than expanding it
  • Dual OpenAI- and Anthropic-compatible API support could accelerate developer adoption by easing migration
  • A promised open-weight release has no confirmed date, license, or Hugging Face repository as of the announcement
  • The 10% preview pricing signals aggressive customer acquisition ahead of standard-rate billing

Was this review helpful?

Share

Twitter/X