Qwen3.6-35B-A3B - GAIA
GAIA score 56.1 from Castle AI
View sourceQwen
Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated...
Running this yourself: likely needs a high-memory cloud gpu.
51.1
Quality Score
---
Arena ELO
36B
Parameters
262K
Context
This measures the amount of verifiable public evidence we have, not how capable the model is. A missing field means it has not been verified yet, not that its value is zero.
16 of 22 public signals
Sign in to join the discussion
5.0M
Downloads
2.8K
Likes
Apr 2026
Released
5/5 signals
3/4 signals
4/5 signals
2/4 signals
2/4 signals
Gaps we are still tracking
Benchmarks
4
Open Source
1
Research
2
Recent launch, pricing, benchmark, and API signals linked to this model or its provider.
GAIA score 56.1 from Castle AI
View sourceQwen3.6-35B-A3B: Agentic Coding Power, Now Open to All - Alibaba Cloud Community Community Blog Events Webinars Tutorials Forum Blog Events Webinars Tutorials Forum Create Account Log In × Community Blog Qwen3.6-35B-A3B: Agentic Coding Power, Now Open to All Qwen3.6-35B-A3B: Agentic Coding Power, Now Open to All Alibaba Cloud Community April 17, 2026 50,119 0 Alibaba open-sources Qwen3.6-35B-A3B, an efficient 35B/3B MoE model delivering top-tier agentic coding and multimodal
GAIA score 44.2 from WA0824
View sourceGAIA score 44.2 from WA0824
View sourceQwen3.6 35B A3B is now available through local Ollama runtime. 256K context window listed. Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models.
View sourceMulti-teacher on-policy distillation (MOPD) distills several domain-expert teachers into a single student by minimizing per-domain reverse-KL divergence on the student's own rollouts. Existing approaches typically fix the per-domain data mixture before training, overlooking the fact that different domains converge at substantially different rates: some plateau early while others continue to improve throughout the training budget. A fixed mixture therefore wastes compute on fast-converging domains and undertrains slower-converging ones. To address this, we propose D^3-MOPD (Dynamic Domain ScheDuling for MOPD), a zero-overhead scheduler that repurposes the per-domain reverse-KL signal already produced during training to adapt the domain mixture online. Running asynchronously outside the training process, an off-process watcher periodically tracks each domain's KL trajectory, estimates remaining headroom and current improvement rate, and accordingly adjusts the domain sampling ratios without altering the core training loop. Our D^3-MOPD scales naturally to arbitrary numbers of domains, and the expected benefit grows as more domains introduce more diverse convergence patterns for the scheduler to exploit. On a Qwen3.6-35B-A3B student distilled from four domain-expert teachers, D^3-MOPD closes 97% of the average student-to-teacher performance gap, compared with 63% for vanilla MOPD, reaches the same peak performance with an approximately 3times reduction in rollout steps, and surpasses the specialist teachers on three of seven benchmarks.
Existing benchmarks for MLLM-generated web artifacts assess interaction through local evidence and miss the requirement-induced states and transitions that determine whether a page works. We introduce WebRISE, which compiles task requirements into Interaction Contract Graphs (ICGs) of observable states, user-intent transitions, and DOM/visual assertions for implementation-agnostic browser execution. WebRISE spans 442 tasks across five input modalities (Text, Markdown, Sketch, Image, Video), with 5,495 transitions and 5,271 requirement checks that separate user-stated functions from implicit product-level constraints. Across 14 MLLMs, even the strongest model reaches only 65.6% transition validity and 66.3% requirement coverage, and visual quality is no proxy for behavior (Qwen3.6-35B-A3B on Markdown: V=80.8 yet T=15.5). Video gives the strongest interaction signal (+10.6 pp implicit coverage over Text), while implicit constraints persist; defect injection shows ICG-based scoring detects state errors at 2-16x the rate of checkpoint-style evaluation.
Qwen3.6 35B A3B is now available through local Ollama runtime. 256K context window listed. Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models.
GAIA score 56.1 from Castle AI
Qwen3.6-35B-A3B: Agentic Coding Power, Now Open to All - Alibaba Cloud Community Community Blog Events Webinars Tutorials Forum Blog Events Webinars Tutorials Forum Create Account Log In × Community Blog Qwen3.6-35B-A3B: Agentic Coding Power, Now Open to All Qwen3.6-35B-A3B: Agentic Coding Power, Now Open to All Alibaba Cloud Community April 17, 2026 50,119 0 Alibaba open-sources Qwen3.6-35B-A3B, an efficient 35B/3B MoE model delivering top-tier agentic coding and multimodal