The first half of 2026 has been a landmark period for open-source AI. Qwen3.5 from Alibaba launched in February with a natively multimodal MoE architecture supporting 201 languages. GLM-5 from Z.ai (formerly Zhipu AI) followed days later — 744B parameters, MIT-licensed, trained entirely on Huawei Ascend chips — proving that frontier-quality models no longer require Nvidia GPUs. Then in late April, DeepSeek dropped what many called the "second DeepSeek moment": DeepSeek-V4, a 1.6T-parameter model with a 1M-token context window as standard and pricing roughly 1/6th that of the leading proprietary rivals.
This review compares all three models across architecture innovation, benchmark performance, pricing strategy, and practical use cases — based on data available through May 2026.
1. Product Overview & Background
Qwen3.5 — Alibaba's Multimodal Flagship Family
Qwen3.5 is not a single model but an entire family released in three waves. The flagship Qwen3.5-397B-A17B (397B total, 17B activated parameters) launched February 16, 2026, with 256K context and coverage for 201 languages. Unlike its predecessor, Qwen3.5 integrates vision understanding natively — making it a true multimodal model without a separate visual adapter. All models under Apache 2.0 support both "thinking" (extended reasoning) and "non-thinking" (fast response) modes.
A mid-size series (27B dense, 35B-A3B, 122B-A10B) followed on February 24, with small models (0.8B–9B) completing the lineup on March 2. The standout efficiency story: Qwen3.5-35B-A3B with just 3B active parameters outperforms the previous-generation 235B-A22B — a testament to architecture quality over raw scale. Qwen3.5-Plus, the hosted API version, extends context to 1 million tokens.
GLM-5 — Z.ai's Hardware-Independence Statement
GLM-5 was released by Z.ai (formerly Zhipu AI, a Tsinghua University spinoff) on February 11, 2026. With ~745B total parameters and 40–44B activated per token, it operates with a 200K token context window under the permissive MIT license. Most notably, GLM-5 was trained entirely on Huawei Ascend chips using the MindSpore framework — making it the highest-capability open-source model with no Nvidia GPU dependency in its training pipeline.
Z.ai completed its Hong Kong IPO on January 8, 2026, raising HKD 4.35 billion (~USD 558M), becoming the world's first publicly traded foundation model company at a ~$31.3B valuation. The capital directly accelerated their release cadence: GLM-5 (Feb 11) → GLM-5 Turbo closed-source agent variant (Mar 15) → GLM-5.1 API (Mar 27) → GLM-5.1 open weights (Apr 7). Shares surged up to 34% on the GLM-5 announcement.
DeepSeek-V4 — The "Second DeepSeek Moment"
Released April 24, 2026 — exactly 484 days after V3 — DeepSeek-V4 comes in two flavors: Pro (1.6T total, 49B active) and Flash (284B total, 13B active). Both ship with a 1M-token context window as standard and MIT licensing. The core innovation is a hybrid Compressed Sparse Attention (CSA) + Heavily Compressed Attention (HCA) mechanism: at 1M-token context, V4-Pro requires only 27% of V3.2's per-token FLOPs and just 10% of its KV cache. Pre-trained on 32T+ tokens with the Muon optimizer, V4 also natively integrates with Claude Code, OpenClaw, and OpenCode agent frameworks.
2. Model Lineup at a Glance
| Model | Params (Total/Active) | Context | Architecture | License | Release |
|---|---|---|---|---|---|
| Qwen3.5-397B-A17B | 397B / 17B | 256K | Hybrid Linear MoE | Apache 2.0 | Feb 16, 2026 |
| Qwen3.5-Plus (API) | 397B / 17B (hosted) | 1M | Same (closed API) | API-only | Feb 16, 2026 |
| Qwen3.5 Mid Series | 27B / 35B-A3B / 122B | 256K | Dense + MoE | Apache 2.0 | Feb 24, 2026 |
| GLM-5 | ~745B / 40–44B | 200K | MoE (Ascend chips) | MIT | Feb 11, 2026 |
| GLM-5.1 (Upgraded) | 744B / 40B | 200K / 131K out | MoE, RL-aligned | MIT | Apr 7, 2026 |
| DeepSeek-V4-Pro | 1.6T / 49B | 1M | MoE (CSA+HCA hybrid) | MIT | Apr 24, 2026 |
| DeepSeek-V4-Flash | 284B / 13B | 1M | MoE (lightweight) | MIT | Apr 24, 2026 |
3. Benchmark Performance Deep Dive
All figures are drawn from official technical reports or independent third-party evaluators (SWE-bench, GPQA Diamond, AIME 2026, LiveCodeBench, Humanity's Last Exam). Flagship models are compared at maximum reasoning effort (Pro-Max / Max modes).
Key takeaways by domain:
- Coding: DeepSeek-V4-Pro-Max leads with a Codeforces rating of 3,206 and LiveCodeBench score of 93.5. GLM-5.1 tops SWE-Bench Pro at 58.4, outperforming Claude Opus 4.6 (57.3).
- Math/Reasoning: GLM-5 achieves 96.9 on HMMT Nov. 2025 (highest among all tested models); all three are clustered at 91–92.7 on AIME 2026 I.
- Graduate Science (GPQA): DeepSeek-V4-Pro-Max leads at 90.1%, followed by Qwen3.5 (88.4%) and GLM-5 (86.0%).
- Multimodal: Qwen3.5 dominates — MMMU 85.0 and MathVision 88.6 (beating GPT-5.2's 83.0 and Gemini 3 Pro's 86.6).
- Hard reasoning (HLE): DeepSeek-V4 scores 37.7% without tools; GLM-5 reaches 50.4% with tools, beating most proprietary models.
4. Pricing Strategy
API pricing below is sourced from OpenRouter and official vendor documentation as of May 2026.
| Model | Input ($/M tokens) | Output ($/M tokens) | Context | License |
|---|---|---|---|---|
| Qwen3.5-Flash | $0.13 | $0.52 | 256K | Apache 2.0 |
| Qwen3.5-Plus (API) | $0.40 | $2.40 | 1M | Closed API |
| GLM-5 | $0.60 | $1.92 | 200K | MIT |
| GLM-5.1 (latest) | $1.05–$1.40* | $3.50–$4.40* | 200K | MIT |
| DeepSeek-V4-Flash | $0.14 | $0.28 | 1M | MIT |
| DeepSeek-V4-Pro | $1.74 | $3.48 | 1M | MIT |
*GLM-5.1 charges 3× standard rate during peak hours (14:00–18:00 Beijing Time). Sources: OpenRouter and Z.ai official docs (May 2026).
💡 Value Highlights: DeepSeek-V4-Flash at $0.14/$0.28 (input/output) is the most cost-effective frontier-class model available today — roughly 17× cheaper on output than Qwen3.5-Plus. Qwen3.5-Flash offers the best value in multimodal scenarios. GLM-5 sits in the middle on price but delivers maximum freedom: MIT license, no hardware dependency, and self-hosting support via vLLM/SGLang on Ascend chips.
5. Features & Ecosystem Comparison
| Feature | Qwen3.5 | GLM-5/5.1 | DeepSeek-V4 |
|---|---|---|---|
| Thinking / Reasoning Mode | ✓ Dual-mode toggle | ✓ Reasoning mode | ✓ 3-tier effort control |
| Native Multimodal (Vision) | ✓ Image/Video/Docs | △ Separate GLM-5V-Turbo | ✗ Not supported |
| Max Context Window | 1M (Plus API) | 200K | 1M (standard) |
| Autonomous Agent Mode | ✓ Enhanced tool use | ✓ 8+ hr autonomous runs | ✓ Coding-agent optimized |
| Direct Document Generation | △ Limited | ✓ .docx/.xlsx/.pdf | ✗ |
| Non-Nvidia Chip Training | △ | ✓ Huawei Ascend (100%) | ✗ |
| OpenAI-Compatible API | ✓ | ✓ | ✓ + Anthropic API |
| Language Coverage | 201 languages | Primarily EN/ZH | Primarily EN/ZH |
6. Architecture Innovations
Qwen3.5 — Hybrid Linear MoE
Gated DeltaNet linear attention fused with sparse MoE routing. The 35B-A3B model activates just 3B parameters yet outperforms the prior-gen 235B-A22B — proving architecture quality beats parameter count. Early-fusion multimodal training, a 250K vocabulary, and multi-token prediction reduce token costs by 10–60% across 201 languages.
GLM-5 — Slime Async RL
Z.ai's custom Slime asynchronous RL infrastructure significantly improves post-training throughput. Trained entirely on Huawei Ascend chips via MindSpore — no Nvidia GPUs at all. GLM-5.1 achieves capability uplift through refined RL alignment alone (no additional pre-training), and can autonomously execute tasks for 8+ hours, iterating 655+ times in a single session.
DeepSeek-V4 — Dual Sparse Attention
Hybrid CSA + HCA attention: at 1M-token context, requires just 27% of V3.2's FLOPs and 10% of KV cache — making 1M context practical as a default. The Muon optimizer replaces AdamW for faster convergence at trillion-parameter scale, while FP4 quantization-aware training on expert weights enables efficient inference without post-training quality loss.
7. Real-World Experience
✅ Core Strengths
- Best-in-class visual understanding (MathVision 88.6, MMMU 85.0)
- 201-language support for global deployments
- Qwen3.5-9B punches above its weight: GPQA Diamond 81.7 vs 70B+ rivals
- Qwen Studio ecosystem: image, video, code, docs in one place
- #1 on SWE-Bench Pro (58.4), outperforms Claude Opus 4.6
- 8+ hour autonomous agent runs with self-correction loops
- Full hardware independence: Huawei Ascend end-to-end
- MIT license, zero vendor lock-in, commercially free
- Codeforces 3,206 — highest competitive programming score of any model
- LiveCodeBench 93.5 — best coding benchmark result
- V4-Flash at $0.14/M input — best price/performance ratio
- Both OpenAI and Anthropic API format compatibility
⚠️ Limitations to Consider
- 397B flagship has very high local deployment requirements
- AIME 2026 slightly below competitors (91.3 vs 92.7)
- HLE no-tools ~30%, below DeepSeek-V4's 37.7%
- 200K context window, well behind both rivals' 1M
- 3× peak-hour pricing creates unpredictable costs
- Vision requires separate GLM-5V-Turbo model
- International language coverage lags far behind Qwen3.5
- No native multimodal/vision capability
- Pro inference speed is slow (~32 tokens/s)
- NIST CAISI independent eval shows ~8 months behind US frontier
- GPQA Diamond 90.1% still trails GPT-5.5 (93.6%)
8. Scorecard & Recommendations
| Dimension | Qwen3.5 | GLM-5/5.1 | DeepSeek-V4 |
|---|---|---|---|
| Coding | ⭐⭐⭐⭐ 8.0 | ⭐⭐⭐⭐ 9.0 | ⭐⭐⭐⭐⭐ 10.0 |
| Multimodal | ⭐⭐⭐⭐⭐ 9.5 | ⭐⭐⭐ 6.0 | ⭐ 3.0 |
| Math & Reasoning | ⭐⭐⭐⭐ 8.0 | ⭐⭐⭐⭐ 9.0 | ⭐⭐⭐⭐ 8.5 |
| API Value | ⭐⭐⭐⭐ 9.0 | ⭐⭐⭐ 7.0 | ⭐⭐⭐⭐⭐ 9.5 |
| Agent Capability | ⭐⭐⭐⭐ 8.0 | ⭐⭐⭐⭐⭐ 9.5 | ⭐⭐⭐⭐ 8.5 |
| Openness | ⭐⭐⭐⭐ 8.5 | ⭐⭐⭐⭐ 8.5 | ⭐⭐⭐⭐ 8.5 |
| Weighted Average | 8.5 / 10 | 8.2 / 10 | 8.3 / 10 |
🎯 Quick Selection Guide
- Multimodal / global products: Choose Qwen3.5-Plus — 201 languages + native image/video understanding, most mature ecosystem
- Resource-constrained edge deployment: Choose Qwen3.5-9B — GPQA Diamond 81.7, outperforms models 13× its size
- Long-horizon coding agents / enterprise self-hosting: Choose GLM-5.1 — #1 SWE-Bench Pro, MIT license, Ascend chip self-host option
- Competitive programming / maximum coding power: Choose DeepSeek-V4-Pro-Max — Codeforces 3,206, unmatched
- High-volume cost-sensitive API workloads: Choose DeepSeek-V4-Flash — $0.14/M input, near-Pro capability
9. Conclusion
These three models collectively mark a turning point: open-source AI has caught up with — and in specific domains surpassed — the best proprietary models, at a fraction of the price.
Qwen3.5 is the versatile all-rounder with unmatched multimodal depth and international reach. Its small-model series makes on-device deployment a reality for the first time at this capability level. GLM-5/5.1 is the choice for teams who demand maximum sovereignty: MIT license, full Huawei Ascend training stack, and the world's best open-source coding agent. It also marks a significant milestone in China's AI hardware independence. DeepSeek-V4 once again proves that architectural innovation beats brute-force scaling — its coding and long-context economics are unmatched, making it the default choice for developer toolchains worldwide.
The next round of this competition will be fought simultaneously on three fronts: architectural efficiency, ecosystem integration, and hardware diversity. Whatever your use case, all three offer frontier-class capability without frontier-class price tags — the era of AI accessibility has arrived.
This review was generated with the assistance of AI technology.