📊 Full opportunity report: The Confusing World Of AI Rankings: Qwen3.8-Max’s Place In The Spotlight on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
Alibaba officially released details on Qwen3.8-Max, a 2.4 trillion-parameter model with 95 billion active parameters. The company published benchmark results, confirming its strong performance in some areas but also revealing limitations. The release highlights ongoing confusion in AI model rankings and open-weight model deployment.
Alibaba has officially made Qwen3.8-Max broadly available, confirming it as the largest open-weight AI model with 2.4 trillion parameters. The company published its benchmark table and announced that open weights will ship next week, marking a significant milestone in AI model transparency and deployment. This development follows weeks of speculation and partial previews, elevating Alibaba’s position in the AI landscape.
On 3 August, Alibaba revealed the full specifications of Qwen3.8-Max, a model built on the Qwen3.5 architecture with 2.4 trillion total parameters and approximately 95 billion active parameters during inference. The model employs sparse mixture-of-experts technology and supports multimodal input (text, image, video) with text output. Benchmark results, obtained using Alibaba’s own testing framework, show competitive performance, notably achieving top scores in several tasks such as Terminal-Bench 2.1 (86.6) and PaperBench (93.0). However, it trails in deep software engineering benchmarks like SWE-bench Pro, where it scored 67.7 against Fable 5’s 80.0, indicating limitations in certain specialized areas.
The company also confirmed the upcoming release of open weights for a 27B checkpoint, designed for local deployment on high-memory machines. The flagship model’s active parameters and benchmark performance have generated mixed reactions, with some praising the agentic improvements and others questioning the model’s overall capabilities relative to claims. Alibaba’s shares rose by up to 5.4% following the announcement, reflecting market optimism but also skepticism about the model’s real-world impact.
For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.
▲ All performance figures: Alibaba’s own harnessThe claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.
“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.
“Qwen3.8 is going open-weight” describes three things with very different deployment realities.
OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.
A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.
The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.
Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.
- The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
- More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
- If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
- The 27B sibling could become the best local agent model on hardware people already own.
- Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
- The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
- “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
- Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
and it says “second only” depends entirely on which row you read.
Implications of Alibaba’s Largest Open-Weight Model
This release underscores the ongoing race among AI labs to develop and deploy ever-larger models with open weights, which could democratize access but also complicate rankings and comparisons. The detailed benchmark data provides transparency but also reveals that claims of being "second only to Fable 5" are based on selective metrics. The demonstrated agentic improvements and long-horizon capabilities suggest potential for practical applications, yet the model’s limitations in certain benchmarks highlight the challenge of balancing size, performance, and deployability. For readers, this development emphasizes the shifting landscape of AI competitiveness and the importance of understanding what model size and benchmarks truly signify.

Kinupute AI Server Mini PC AMD Ryzen AI 9 HX 370(12 Cores | 80 Tops | Radeon 890M), Win-11 Pro, 32G DDR5, 2T PCIE4.0 SSD, Desktop Computers Dual 2.5G LAN, Oculink/USB4.0 /Triple Display 8K/WiFi7, VESA
- High-Speed AI Performance: Powered by AMD Ryzen AI 9 HX370 processor
- AI Acceleration: 50 TOPS AI acceleration with NPU
- Expandable Memory: Supports up to 128GB DDR5 RAM
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background on Alibaba’s AI Model Developments
Over the past month, Alibaba’s AI efforts have been characterized by strategic secrecy and selective disclosures. In July, the company previewed Qwen3.8-Max as an anonymous entity called "kaleb" at the World AI Conference in Shanghai, without releasing detailed specifications or benchmarks. The model was briefly available via a paid preview endpoint, with a 983,616-token context window and multiple output settings. The hype was fueled by the model’s claimed size and performance, but concrete details remained scarce until the recent full release.
Prior to this, Alibaba’s AI models, such as Kimi K3 and earlier Qwen versions, had gained attention for their rapid development and competitive benchmarks. The new release marks a significant escalation, positioning Alibaba as a leader in the size and openness of AI models, even as the community debates the real-world significance of such large parameters and the transparency of benchmark claims.
"Qwen3.8-Max demonstrates our commitment to open AI development and transparency, with benchmark results that highlight its strengths and limitations."
— Alibaba spokesperson

Local LLM Inference Optimization: A Comprehensive Guide to Quantization, Hardware Acceleration, and Efficient Private AI Deployment
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unresolved Questions About Model Capabilities
While Alibaba has published benchmark scores and specifications, some questions remain about the model’s performance in real-world applications, especially in specialized tasks like software engineering. The impact of the upcoming open weights for the 27B checkpoint on local deployment and whether the agentic gains will persist after compression are still unclear. Additionally, the licensing terms for the open weights are yet to be announced, which could influence adoption and trust.

LAFVIN ESP32-S3 1.69" LCD Development Board with Camera, AI Vision Voice Development Kit, Programmable IoT Board with Mic Speaker for STEM Education
- Powerful Microcontroller: ESP32-S3 with 16MB Flash and 8MB PSRAM
- AI Vision & Voice Capabilities: Camera and audio for AI interactions
- Supports OpenCV & YOLO: Face tracking and human pose estimation
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps for Alibaba’s AI Model Rollout
Alibaba plans to release the open weights for Qwen3.8-27B next week, enabling local deployment on high-memory hardware. The company also intends to publish detailed licensing information and further benchmark results, particularly in software engineering and multimodal tasks. Industry watchers will monitor how the model performs in practical deployments and whether the claimed agentic improvements translate into real-world benefits. Market reactions and community evaluations will shape the model’s reputation in the coming weeks.

Etekcity Digital Body Weight Bathroom Scale, 440 lb Extra Wide Platform
- Large Platform Size: 13.8 x 11.8 inches with LCD display
- High Weight Capacity: Supports up to 440 pounds
- Accurate Measurements: High-precision sensors for reliable readings
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What makes Qwen3.8-Max different from other large language models?
Qwen3.8-Max is notable for its size—2.4 trillion parameters—and its open-weight availability, making it the largest openly accessible model of its kind. It also demonstrates strong multimodal and agentic capabilities, though it trails in some deep software engineering benchmarks.
When will the open weights for Qwen3.8-27B be available?
The open weights for the 27B checkpoint are scheduled for release next week, allowing local deployment on high-memory machines.
Does the model's size guarantee better performance?
Not necessarily. While larger size can improve certain tasks, benchmark results show that performance varies across different domains, and size alone does not ensure superiority in all areas.
What are the licensing implications for the open weights?
The licensing terms for the open weights have not yet been published, which could impact how the model is used and integrated into other systems.
Source: ThorstenMeyerAI.com