AIThis post was created with the assistance of artificial intelligence (AI).

📊 Full opportunity report: Qwen3.8-Max's AI Performance: A Deeper Look At The Latest Figures on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Alibaba officially released detailed benchmarks for its Qwen3.8-Max model, confirming a 2.4 trillion-parameter size and top scores in several AI benchmarks. The open weights will be available next week, marking a significant step in large-scale open AI models.

Alibaba has officially published comprehensive benchmark results for its Qwen3.8-Max model, confirming it as a 2.4 trillion-parameter, multimodal AI system with strong performance across several benchmarks. The release includes the full benchmark table and an announcement that open weights will be available next week, marking a significant milestone in large-language-model deployment.

On August 3, Alibaba disclosed the full benchmark results for its Qwen3.8-Max model, which had previously been previewed stealthily. The model features approximately 95 billion active parameters per query within a 2.4 trillion-parameter sparse mixture-of-experts architecture, built on the Qwen3.5 foundation. It demonstrates top-tier scores on benchmarks such as Terminal-Bench 2.1 (86.6), PaperBench (93.0), and others, often outperforming competitors like Claude Fable 5 and approaching GPT-5.6 Sol performance at its peak.

The model excels particularly in multimodal and agentic tasks, with notable scores on OSWorld-Verified (86.1), Parametric CAD Bench (91.5), and OmniDocBench (92.1). It also demonstrated the ability to reproduce research results and outperform its own previous iterations in long-horizon agent tasks, signifying significant progress in agentic AI capabilities. However, it trails in deep software-engineering benchmarks like SWE-bench Pro and FrontierSWE, with gaps of twelve to fifteen points compared to Fable 5.

The announcement also clarified that the 2.4 trillion parameters are primarily a model size metric, with active parameters used per query around 95 billion. The benchmark results were achieved using Alibaba’s own testing harness, and the full benchmark table was released alongside the announcement. Open weights for the model are scheduled to ship next week, though the licensing details remain unpublished, and the weights are intended for multi-node datacenter deployment rather than individual use.

At a glance
reportWhen: announced August 3, 2023; benchmark res…
The developmentAlibaba announced the full benchmark results for Qwen3.8-Max, confirming its size and performance, and revealed that open weights will be released next week.
AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
“Next week”
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
“kaleb” appears on Code Arena
Anonymous model introduces itself as “Claude” — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: “second only to Fable 5”
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

“Qwen3.8 is going open-weight” describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
“Next week” · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says “second only” depends entirely on which row you read.

Implications of Alibaba's Benchmark Reveal for AI Leadership

The publication of detailed benchmarks and the confirmation of the 2.4 trillion parameters establish Alibaba as a serious contender in the large-language-model space, especially with its focus on multimodal and agentic capabilities. The upcoming release of open weights will enable wider adoption and experimentation, potentially accelerating innovation in open AI deployment. However, the model's performance gaps in certain engineering benchmarks highlight ongoing challenges in scaling deep software tasks.

This development signals a shift toward more transparent disclosure of large-model capabilities and sets a new benchmark for open-weight models, which could influence industry standards and competitive positioning among AI labs and companies.

AI Systems Performance Engineering: Optimizing Model Training and Inference Workloads with GPUs, CUDA, and PyTorch

AI Systems Performance Engineering: Optimizing Model Training and Inference Workloads with GPUs, CUDA, and PyTorch

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of Alibaba’s AI Model Releases and Benchmark Strategy

Alibaba's AI journey has been marked by stealth previews and selective disclosures, culminating in the July preview of Qwen3.8-Max, which was initially identified through community detection methods. The model's size and capabilities have been a subject of speculation, with previous models like Kimi K3 and the anonymous 'kaleb' contributing to the narrative. The company's strategy involved a staged rollout, culminating in this comprehensive benchmark publication, which provides transparency and sets the stage for open-weight deployment.

Earlier models, including Qwen3.5 and Qwen3.7-Max, laid the groundwork for the current release, with incremental improvements in multimodal understanding, agentic reasoning, and benchmark performance. The focus on RL-environment scaling and long-horizon tasks reflects Alibaba's emphasis on advancing AI capabilities beyond traditional benchmarks.

"We are committed to advancing AI capabilities and providing open access to our latest models, starting with next week's open weights release."

— Alibaba spokesperson

stedi Model Tools Kit for Beginners 14 PCS, Modeler Premium Basic Tools

stedi Model Tools Kit for Beginners 14 PCS, Modeler Premium Basic Tools

  • Professional Basic Model Kit: Includes essential tools for model building
  • Complete Tool Set: Nippers, tweezers, sanding sticks, utility knife, and accessories
  • High-Quality & Durable: All-metal craft knife and washable sanding sticks

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unconfirmed Details About Model Licensing and Deployment

It remains unclear what the exact licensing terms will be for the open weights, as Alibaba has not yet published the license details. Additionally, the deployment scope—whether the open weights will be suitable for individual or small-scale use—is still uncertain, given the model's size and infrastructure requirements. The impact of potential licensing restrictions on broader adoption also remains to be seen.

LAFVIN ESP32-S3 1.69" LCD Development Board with Camera, AI Vision Voice Development Kit, Programmable IoT Board with Mic Speaker for STEM Education

LAFVIN ESP32-S3 1.69" LCD Development Board with Camera, AI Vision Voice Development Kit, Programmable IoT Board with Mic Speaker for STEM Education

  • Powerful Microcontroller: ESP32-S3 with 16MB Flash and 8MB PSRAM
  • AI Vision & Voice Capabilities: Camera and audio for AI interactions
  • Multiple Development Platforms: Supports Arduino IDE and ESP-IDF

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Alibaba’s Open-Source AI Model Strategy

Alibaba is scheduled to release the open weights of Qwen3.8-Max next week, which will likely trigger broader testing and deployment by researchers and developers. The company may also publish licensing details and usage guidelines at that time. Observers will watch for how the community adopts and adapts the model, particularly in multimodal and agentic applications, and whether the performance in deep engineering tasks improves with future updates.

Doom's Benchmark: The Game That Measures Machines (Prompt Engineering with AI)

Doom's Benchmark: The Game That Measures Machines (Prompt Engineering with AI)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

When will Alibaba release the open weights for Qwen3.8-Max?

The open weights are scheduled to be released next week, following the benchmark publication on August 3, 2023.

What are the key performance strengths of Qwen3.8-Max?

The model excels in multimodal understanding, agentic reasoning, and certain benchmarks like PaperBench and OSWorld-Verified, approaching top-tier scores among large models.

How does Qwen3.8-Max compare to other models like GPT-5.6 or Fable 5?

It outperforms models like Claude Fable 5 and approaches GPT-5.6 at its peak on some benchmarks, but trails significantly in deep software-engineering tasks.

Will the open weights be suitable for individual or small-scale deployment?

Given the model's size and infrastructure needs, the open weights are primarily intended for multi-node datacenter use, not individual deployment.

What are the limitations of Qwen3.8-Max based on current benchmarks?

The model shows significant gaps in deep engineering benchmarks like SWE-bench Pro and FrontierSWE, indicating ongoing challenges in software-specific tasks.

Source: ThorstenMeyerAI.com

You May Also Like

Boost AI Outcomes By Focusing On Talent Density

Exploring how AI amplifies talent density, enabling small teams of high performers to outperform larger organizations and reshape productivity metrics.

The unstoppable rise of China-made cars in Europe: 5 things to know

China-made electric and hybrid vehicles are rapidly increasing market share in Europe, driven by affordability and expanding model offerings, despite tariffs.

Fubo rolls out price increase on NBC-inclusive plans after new carriage deal

Fubo raises subscription prices for plans including NBC following a new carriage agreement, impacting customers and streaming options.

Forge Oder Eigenes Hosting? Die Kosten Für Souveräne KI Im Check

Analyse der Kosten für selbstgehostete KI im Vergleich zu europäischen Cloud-Anbietern, inklusive aktueller Entwicklungen und Unsicherheiten.