📊 Full opportunity report: The Confusing World Of AI Rankings: Qwen3.8-Max’s Place In The Spotlight on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Alibaba officially released details on Qwen3.8-Max, a 2.4 trillion-parameter model with 95 billion active parameters. The company published benchmark results, confirming its strong performance in some areas but also revealing limitations. The release highlights ongoing confusion in AI model rankings and open-weight model deployment.

Alibaba has officially made Qwen3.8-Max broadly available, confirming it as the largest open-weight AI model with 2.4 trillion parameters. The company published its benchmark table and announced that open weights will ship next week, marking a significant milestone in AI model transparency and deployment. This development follows weeks of speculation and partial previews, elevating Alibaba’s position in the AI landscape.

On 3 August, Alibaba revealed the full specifications of Qwen3.8-Max, a model built on the Qwen3.5 architecture with 2.4 trillion total parameters and approximately 95 billion active parameters during inference. The model employs sparse mixture-of-experts technology and supports multimodal input (text, image, video) with text output. Benchmark results, obtained using Alibaba’s own testing framework, show competitive performance, notably achieving top scores in several tasks such as Terminal-Bench 2.1 (86.6) and PaperBench (93.0). However, it trails in deep software engineering benchmarks like SWE-bench Pro, where it scored 67.7 against Fable 5’s 80.0, indicating limitations in certain specialized areas.

The company also confirmed the upcoming release of open weights for a 27B checkpoint, designed for local deployment on high-memory machines. The flagship model’s active parameters and benchmark performance have generated mixed reactions, with some praising the agentic improvements and others questioning the model’s overall capabilities relative to claims. Alibaba’s shares rose by up to 5.4% following the announcement, reflecting market optimism but also skepticism about the model’s real-world impact.

At a glance
reportWhen: announced August 3, 2023; details relea…
The developmentAlibaba announced the full specifications and benchmark results for Qwen3.8-Max, confirming its status as the largest open-weight AI model to date.
AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
“Next week”
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
“kaleb” appears on Code Arena
Anonymous model introduces itself as “Claude” — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: “second only to Fable 5”
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

“Qwen3.8 is going open-weight” describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
“Next week” · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says “second only” depends entirely on which row you read.

Implications of Alibaba’s Largest Open-Weight Model

This release underscores the ongoing race among AI labs to develop and deploy ever-larger models with open weights, which could democratize access but also complicate rankings and comparisons. The detailed benchmark data provides transparency but also reveals that claims of being "second only to Fable 5" are based on selective metrics. The demonstrated agentic improvements and long-horizon capabilities suggest potential for practical applications, yet the model’s limitations in certain benchmarks highlight the challenge of balancing size, performance, and deployability. For readers, this development emphasizes the shifting landscape of AI competitiveness and the importance of understanding what model size and benchmarks truly signify.

Kinupute AI Server Mini PC AMD Ryzen AI 9 HX 370(12 Cores | 80 Tops | Radeon 890M), Win-11 Pro, 32G DDR5, 2T PCIE4.0 SSD, Desktop Computers Dual 2.5G LAN, Oculink/USB4.0 /Triple Display 8K/WiFi7, VESA
  • High-Speed AI Performance: Powered by AMD Ryzen AI 9 HX370 processor
  • AI Acceleration: 50 TOPS AI acceleration with NPU
  • Expandable Memory: Supports up to 128GB DDR5 RAM

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Alibaba’s AI Model Developments

Over the past month, Alibaba’s AI efforts have been characterized by strategic secrecy and selective disclosures. In July, the company previewed Qwen3.8-Max as an anonymous entity called "kaleb" at the World AI Conference in Shanghai, without releasing detailed specifications or benchmarks. The model was briefly available via a paid preview endpoint, with a 983,616-token context window and multiple output settings. The hype was fueled by the model’s claimed size and performance, but concrete details remained scarce until the recent full release.

Prior to this, Alibaba’s AI models, such as Kimi K3 and earlier Qwen versions, had gained attention for their rapid development and competitive benchmarks. The new release marks a significant escalation, positioning Alibaba as a leader in the size and openness of AI models, even as the community debates the real-world significance of such large parameters and the transparency of benchmark claims.

"Qwen3.8-Max demonstrates our commitment to open AI development and transparency, with benchmark results that highlight its strengths and limitations."

— Alibaba spokesperson

Local LLM Inference Optimization: A Comprehensive Guide to Quantization, Hardware Acceleration, and Efficient Private AI Deployment

Local LLM Inference Optimization: A Comprehensive Guide to Quantization, Hardware Acceleration, and Efficient Private AI Deployment

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About Model Capabilities

While Alibaba has published benchmark scores and specifications, some questions remain about the model’s performance in real-world applications, especially in specialized tasks like software engineering. The impact of the upcoming open weights for the 27B checkpoint on local deployment and whether the agentic gains will persist after compression are still unclear. Additionally, the licensing terms for the open weights are yet to be announced, which could influence adoption and trust.

LAFVIN ESP32-S3 1.69" LCD Development Board with Camera, AI Vision Voice Development Kit, Programmable IoT Board with Mic Speaker for STEM Education

LAFVIN ESP32-S3 1.69" LCD Development Board with Camera, AI Vision Voice Development Kit, Programmable IoT Board with Mic Speaker for STEM Education

  • Powerful Microcontroller: ESP32-S3 with 16MB Flash and 8MB PSRAM
  • AI Vision & Voice Capabilities: Camera and audio for AI interactions
  • Supports OpenCV & YOLO: Face tracking and human pose estimation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Alibaba’s AI Model Rollout

Alibaba plans to release the open weights for Qwen3.8-27B next week, enabling local deployment on high-memory hardware. The company also intends to publish detailed licensing information and further benchmark results, particularly in software engineering and multimodal tasks. Industry watchers will monitor how the model performs in practical deployments and whether the claimed agentic improvements translate into real-world benefits. Market reactions and community evaluations will shape the model’s reputation in the coming weeks.

Etekcity Digital Body Weight Bathroom Scale, 440 lb Extra Wide Platform

Etekcity Digital Body Weight Bathroom Scale, 440 lb Extra Wide Platform

  • Large Platform Size: 13.8 x 11.8 inches with LCD display
  • High Weight Capacity: Supports up to 440 pounds
  • Accurate Measurements: High-precision sensors for reliable readings

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What makes Qwen3.8-Max different from other large language models?

Qwen3.8-Max is notable for its size—2.4 trillion parameters—and its open-weight availability, making it the largest openly accessible model of its kind. It also demonstrates strong multimodal and agentic capabilities, though it trails in some deep software engineering benchmarks.

When will the open weights for Qwen3.8-27B be available?

The open weights for the 27B checkpoint are scheduled for release next week, allowing local deployment on high-memory machines.

Does the model's size guarantee better performance?

Not necessarily. While larger size can improve certain tasks, benchmark results show that performance varies across different domains, and size alone does not ensure superiority in all areas.

What are the licensing implications for the open weights?

The licensing terms for the open weights have not yet been published, which could impact how the model is used and integrated into other systems.

Source: ThorstenMeyerAI.com

You May Also Like

Alibaba Launches Qwen3.8-Max, Its Largest AI Model Yet

Alibaba unveils Qwen3.8-Max, its largest AI model to date, aiming to enhance AI capabilities across various sectors. Details on size and capabilities announced.

Should You Use Mistral Forge? A Buyer’s Decision Guide

Evaluate if Mistral Forge suits your needs with this comprehensive decision guide, focusing on data sovereignty, technical capacity, and use case fit.

VigilSAR Benchmark: There Is No Best Model

VigilSAR’s new benchmark shows there is no universally best AI model; suitability depends on specific deployment needs and compliance requirements.

Twenty-five Years Ago It Was Cryptography, Today It’s Model Weights

A look at how AI’s focus shifted from cryptography in the 1990s to model weights today, highlighting key developments and ongoing uncertainties.