Qwen3.8-Max open weights (Qwen3.8-2.4T-A95B)
Alibaba's frontier Qwen-Max tier as downloadable open weights — a 2.4T-param sparse MoE (95B active across 512 experts, 262K native context). Released August, newly trending this week: Fixstars benchmarked day-1 serving at ~2x Kimi-K3 throughput on 8x B300, and AWS published a full SageMaker HyperPod deployment path.
Why it matters
What you could build with it
Does it hold up?
Grounded evidence is unusually strong for a datacenter-class release: Fixstars benchmarked it live on 8x B300 at ~2x Kimi-K3 throughput, AWS documented a full HyperPod path, and community quants (Unsloth GGUF, RadixArk NVFP4, Inferact NVFP4) exist. But the bar is a multi-GPU cluster — individuals cannot run it, so for most developers the 27B sibling is the usable artifact.
Built with Qwen3.8-Max open weights
- Fixstars: Day-1 deployment on 8x NVIDIA B300 with SGLang — ~2x Kimi-K3 throughputfixstars tech blog · Fixstars served the NVFP4 quant (~1.48TB) on a single 8x B300 node: ~11,833 tok/s total throughput at concurrency 50 with zero failures (Kimi-K3 failed 119/200 at the same level), 65-min startup, and ~1.54M-token effective context. They note PTQ may lose some accuracy vs official numbers and that the cheapest usable quant needs ~450GB+ of combined RAM/VRAM.
- AWS: Deploying Qwen3.8-2.4T-A95B on SageMaker HyperPod with vLLMaws.amazon.com · Full production deployment walkthrough (second in AWS's trillion-parameter series): HyperPod on ml.p6-b300, vLLM with NVFP4, built-in reasoning parser, tool calling, and native MTP speculative decoding. Confirms ~4.8TB BF16 memory footprint and the 72x Blackwell Ultra reference deployment.
- unsloth/Qwen3.8-2.4T-A95B-GGUFgithub · Unsloth's Dynamic 2.0 GGUF quants (including 1-bit) of the 2.4T model with a how-to-run guide, compatible with llama.cpp/Ollama/Unsloth Desktop.
- RadixArk/Qwen3.8-2.4T-A95B-NVFP4github · SGLang team's official NVFP4 quantized build (~1.48TB) used in Fixstars' day-one 8x B300 deployment; enabled single-node serving before the model itself even finished uploading to HF.
- lmsysorg/sglanggithub · SGLang shipped day-zero support for Qwen3.8 (qwen38 docker image, cookbook, NEXTN/MTP speculative decoding, reasoning+tool-call parsers) — the runtime behind the fastest reported community deployments.
Learn more
First spotted on modelscope: source.
More AI models
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.