Back to list
Aug 04, 2026
9
0
0
Other LLMNEW

Qwen3.8-Max Ships With Benchmarks, Rivals Claude Opus 4.8

Alibaba moved Qwen3.8-Max out of preview on August 3, publishing benchmark scores against Claude and GPT-5.6, with open weights due next week.

#Qwen#Alibaba#LLM#Mixture of Experts#Benchmark
Qwen3.8-Max Ships With Benchmarks, Rivals Claude Opus 4.8
AI Summary

Alibaba moved Qwen3.8-Max out of preview on August 3, publishing benchmark scores against Claude and GPT-5.6, with open weights due next week.

Introduction

On August 3, 2026, Alibaba's Qwen team moved Qwen3.8-Max out of preview and into general availability, publishing benchmark scores that were missing when the model first appeared at WAIC in July. Qwen3.8-Max is a 2.4-trillion-parameter sparse Mixture-of-Experts model that processes text, images, and video through a 1-million-token context window. The July preview carried an unverified marketing claim, that the model ranked "second only to Fable 5" (Anthropic's flagship), with no published benchmark table to back it up. Monday's release changes that: Alibaba shared results directly comparing Qwen3.8-Max against Claude Opus 4.8, OpenAI's GPT-5.6 Sol, and its own predecessor, and confirmed open weights are coming the following week. Alibaba Group Holding shares rose about 6% in Hong Kong trading on the news.

Feature Overview

On Terminal-Bench 2.1, an agentic terminal-use benchmark, Qwen3.8-Max scored 86.6, edging out Claude Opus 4.8's 84.6 but trailing GPT-5.6 Sol's 88.8. It posted 92.6 on GPQA Diamond, a graduate-level science reasoning test, and 93.0 on PaperBench, which evaluates a model's ability to reproduce results from published research papers. The more striking gains are in coding: FrontierSWE rose from 40.7 in the prior Qwen generation to 73.5, and DeepSWE 1.1 climbed from 21.6 to 56.6, both roughly doubling. Alibaba also had the model complete a multi-day software engineering project autonomously, without human intervention, placing it fourth on the Frontend Code Arena leaderboard with a score of 1,668. Not every metric favors Qwen3.8-Max: on SWE-bench Pro, a harder real-world coding benchmark, it scored 67.7 against Claude Fable 5's 80.0, a gap Alibaba's own marketing has not addressed. Pricing is set at $2.00 per million input tokens and $6.00 per million output tokens, with cached input priced at $0.25 per million tokens, an eightfold discount that rewards workloads reusing the same context repeatedly. The context window tops out at 991,000 input tokens (983,000 with reasoning enabled) and 131,000 output tokens. Alibaba also closed one of the preview's biggest information gaps: of the 2.4 trillion total parameters, 95 billion activate per token. That figure matters more than the headline count, because per-token compute, serving cost, and self-hosting feasibility all scale with active parameters rather than total ones.

Usability Analysis

Qwen3.8-Max is live now through Alibaba Cloud's Model Studio APIs and the QwenWork platform, reachable through an OpenAI-compatible interface that should let teams already building against OpenAI's SDKs point existing code at the new model with minimal rewiring. The bigger usability shift from the preview is that open weights for both the full Qwen3.8-Max and a smaller Qwen3.8-27B checkpoint are now scheduled for release the following week, according to Alibaba, with the 27B version aimed at teams that want to run the model on their own GPUs rather than through a hosted API. That is a meaningfully more concrete commitment than the open-ended "coming soon" language Alibaba used in July. For teams evaluating coding assistants or research tooling, the published Terminal-Bench and coding benchmark gains give a genuine basis for comparison against Claude and GPT-5.6 tiers, something the preview release did not offer.

Pros and Cons

Pros:

  • Published, reproducible benchmark scores replace Alibaba's earlier unverified "second only to Fable 5" marketing claim
  • Terminal-Bench 2.1 and GPQA Diamond scores land within striking distance of Claude Opus 4.8 and GPT-5.6 Sol
  • Aggressive cached-input pricing ($0.25 per million tokens, an 8x discount) rewards repeated-context workloads
  • Open weights for both the flagship and a smaller 27B checkpoint now have a concrete release window
  • Native text, image, and video input broadens use cases beyond a text-only chat model

Cons:

  • SWE-bench Pro (67.7) still trails Claude Fable 5's 80.0 by a wide margin, undercutting Alibaba's top-tier positioning
  • Active parameters per inference pass remain undisclosed, so real-world serving cost is still hard to model
  • Open weights are promised but not yet published, so independent reproduction of Monday's benchmark numbers isn't possible yet
  • The benchmark selection is Alibaba's own choice, and independent third-party leaderboards have not yet replicated the figures

Outlook

If Alibaba follows through on next week's open-weight release with a transparent model card, independent researchers will finally be able to reproduce Monday's benchmark numbers rather than take Alibaba's comparisons on faith, closing the credibility gap that weighed down the July preview. The remaining gap on SWE-bench Pro against Claude Fable 5 suggests Qwen3.8-Max is a strong second-tier contender rather than an outright leader on every coding benchmark. Still, doubling scores on FrontierSWE and DeepSWE 1.1 within a single generation, combined with aggressive cached-token pricing, positions Qwen3.8-Max as a serious option for cost-sensitive teams running high-volume agentic workloads, and keeps pricing pressure on Western labs.

Conclusion

Qwen3.8-Max's move from an unverified preview to a benchmarked, generally available model with a concrete open-weights date addresses the core criticism of July's WAIC announcement. It is not a clean win over Claude or GPT-5.6 on every metric, but the published Terminal-Bench, GPQA Diamond, and coding scores give developers real numbers to evaluate rather than marketing language. Teams building agentic coding tools or high-volume applications with repeated context are the clearest fit. Rating: 4/5.

Editor's Verdict

Qwen3.8-Max Ships With Benchmarks, Rivals Claude Opus 4.8 earns a solid recommendation within the other llm space.

The strongest case for paying attention is published, reproducible benchmark scores replace Alibaba's earlier unverified "second only to Fable 5" marketing claim, which raises the bar for what readers should now expect from peers in this space. Reinforcing that, terminal-Bench 2.1 and GPQA Diamond scores land within striking distance of Claude Opus 4.8 and GPT-5.6 Sol adds practical value rather than just headline appeal. The broader signal worth registering is straightforward: qwen3.8-Max's Terminal-Bench 2.1 score of 86.6 sits between Claude Opus 4.8 (84.6) and GPT-5.6 Sol (88.8), a direct benchmark comparison rather than marketing language. On the other side of the ledger, SWE-bench Pro (67.7) still trails Claude Fable 5's 80.0 by a wide margin, undercutting Alibaba's top-tier positioning is a real constraint, not a marketing footnote, and it should factor into any serious decision. Layered on top of that, independent leaderboard coverage was still thin at launch, with Artificial Analysis and community boards yet to publish their own scores narrows the set of teams for whom this is an obvious yes.

For multi-model deployment teams, cost-conscious operators, and developers willing to evaluate beyond the major labs, this is a serious evaluation candidate, not just a curiosity to bookmark. For everyone else, the safer posture is to monitor coverage and revisit once the use cases that matter to your team are demonstrated in the wild.

Pros

  • Published, reproducible benchmark scores replace Alibaba's earlier unverified "second only to Fable 5" marketing claim
  • Terminal-Bench 2.1 and GPQA Diamond scores land within striking distance of Claude Opus 4.8 and GPT-5.6 Sol
  • Aggressive cached-input pricing ($0.25 per million tokens, an 8x discount) rewards repeated-context workloads
  • Open weights for both the flagship and a smaller 27B checkpoint now have a concrete release window
  • Native text, image, and video input broadens use cases beyond a text-only chat model

Cons

  • SWE-bench Pro (67.7) still trails Claude Fable 5's 80.0 by a wide margin, undercutting Alibaba's top-tier positioning
  • Independent leaderboard coverage was still thin at launch, with Artificial Analysis and community boards yet to publish their own scores
  • Open weights are promised but not yet published, so independent reproduction of the benchmark numbers isn't possible yet
  • The benchmark selection is Alibaba's own choice, so the comparison set is curated rather than neutral

Comments0

Key Features

1. 2.4-trillion-parameter sparse Mixture-of-Experts architecture activating 95 billion parameters per token 2. Terminal-Bench 2.1 score of 86.6, ahead of Claude Opus 4.8 (84.6) but behind GPT-5.6 Sol (88.8) 3. FrontierSWE and DeepSWE 1.1 coding scores roughly doubled from the prior Qwen generation 4. 1-million-token context window (991K input tokens, 131K output tokens) 5. Open weights for Qwen3.8-Max and a smaller Qwen3.8-27B checkpoint scheduled for the following week

Key Insights

  • Qwen3.8-Max's Terminal-Bench 2.1 score of 86.6 sits between Claude Opus 4.8 (84.6) and GPT-5.6 Sol (88.8), a direct benchmark comparison rather than marketing language
  • Of the 2.4 trillion total parameters, 95 billion activate per token, letting developers model serving cost and self-hosting feasibility
  • FrontierSWE score nearly doubled from the prior Qwen generation, rising from 40.7 to 73.5
  • DeepSWE 1.1 score more than doubled, climbing from 21.6 to 56.6
  • SWE-bench Pro at 67.7 still trails Claude Fable 5's 80.0, leaving a real gap despite gains elsewhere
  • Alibaba shares rose about 6% in Hong Kong trading on the day of the announcement
  • Open weights for Qwen3.8-Max and a smaller Qwen3.8-27B checkpoint are scheduled for release the following week
  • Pricing lands at $2.00 per million input tokens and $6.00 per million output tokens, with an 8x discount for cached input

Was this review helpful?

Share

Twitter/X