Qwen 3.5 Plus is a commercial model with hybrid linear attention and sparse MoE architecture. Supports text, image, and video input with a 1M context window.
Added Feb 16, 2026
Context Window
983.6K
Max Output
65.5K
Input Price (Auto)
$0.40/1M
Output Price (Auto)
$2.40/1M
Cache Read (Auto)
$0.040/1M
Capabilities
Benchmarks
Benchmarks
Performance metrics and benchmarks
Sourced from Vectara.
Hallucination Rate
10.7%
Factual Consistency
89.3%
Answer Rate
99.8%
Avg Summary Length
Average generated summary length
92.1
Last updated 2026-05-11 · Matched as qwen/qwen3.5-plus-2026-02-15
Vectara LeaderboardProviders
Auto routing is available for this model. Explicit provider selection is not available.
Loading provider options…
Related text models
Compare Qwen3.5 Plus with similar models from the same provider or model family.
Qwen3.5 Plus Thinking
qwen/qwen3.5-plus-thinkingQwen 3.5 Plus with extended reasoning. A commercial model with hybrid linear attention and sparse MoE architecture. Supports text, image, and video input with a 1M context window.
Qwen3 Coder Plus
qwen/qwen3-coder-plusAlibaba’s proprietary upgrade to the open‑weights Qwen3 Coder 480B A35B. A coding‑first agent model with strong tool use and environment control for autonomous programming, while remaining capable at general tasks.
Qwen3.8 2.4T A95B (Max)
qwen/qwen3.8-2.4t-a95bThis is the same underlying model as Qwen3.8 Max, exposed under its architecture-based 2.4T A95B name for easier discovery. It uses the identical routing, pricing, capabilities, and non-thinking mode.
Qwen 3.8 27B Uncensored Thinking
qwen/qwen3.8-27b-uncensored:thinkingQwen 3.8 27B Uncensored with thinking enabled for more deliberate creative work, coding, multimodal analysis, tool use, and long-context problem solving.
Qwen 3.6 35B A3B Uncensored Thinking
qwen/qwen3.6-35b-a3b-uncensored:thinkingQwen 3.6 35B A3B Uncensored with thinking enabled for more deliberate coding, multimodal analysis, tool use, and complex chat tasks.
Qwen 3.8 27B Obliterated
qwen/qwen3.8-27b-obliteratedQwen 3.8 27B Obliterated is an open-weight multimodal model LoRA-tuned for fewer refusals across chat, coding, reasoning, tool use, and long-context work.