Qwen3-235B-A22B - LiveCodeBench
LiveCodeBench pass@1 80.4 across 1055 tasks
View sourceQwen
Qwen3-235B-A22B-Thinking-2507 is a high-performance, open-weight Mixture-of-Experts (MoE) language model optimized for complex reasoning tasks. It activates 22B of its 235B parameters per forward pass and natively supports up to 262,144...
Running this yourself: likely needs a high-memory cloud gpu.
OpenRouter
Price record: 2026-10-03. Source: openrouter.
No current subscription pricing is tracked for this model.
Confirm this specific model, usage limits, and billing terms with the provider. A subscription does not automatically include API credits.
No login needed to compare. Prices are in USD; provider charges are separate from AI Market Cap plans. Context length, caching, tools, taxes, and regional terms can change the final cost. Open weights do not mean free hosting.
67.2
Quality Score
1367
Arena ELO
235B
Parameters
131K
Context
This measures the amount of verifiable public evidence we have, not how capable the model is. A missing field means it has not been verified yet, not that its value is zero.
17 of 22 public signals
Sign in to join the discussion
11.9K
Downloads
409
Likes
Jul 2025
Released
5/5 signals
4/4 signals
4/5 signals
2/4 signals
2/4 signals
Gaps we are still tracking
Benchmarks
7
Open Source
1
Research
1
Recent launch, pricing, benchmark, and API signals linked to this model or its provider.
LiveCodeBench pass@1 80.4 across 1055 tasks
View sourceLiveCodeBench pass@1 80.4 across 1055 tasks
View sourceArena-Hard-Auto official Gemini-2.5 judged score 58.4 with CI -1.9/2.1
SWE-Bench Verified resolved rate 69.6
View sourceSWE-Bench Verified resolved rate 69.6
View sourceWe present GoLongRL, a fully open-source, capability-oriented post-training recipe for long-context reinforcement learning with verifiable rewards (RLVR). Existing long-context RL methods often treat data construction as a matter of designing increasingly complex retrieval paths, leading to homogeneous task coverage and reward formulations that inadequately reflect practical long-context requirements. Our work offers two contributions. (1) Capability-oriented data construction with full open release. We openly release a dataset of 23K RLVR samples, the complete construction pipeline, and all training code. Guided by a taxonomy of long-context capabilities, the dataset spans 9 task types, each paired with its natural evaluation metric. It comprises curated open-source samples from established corpora and synthetic samples whose QA pairs are generated from real source documents such as books, academic papers, and multi-turn dialogues. Under the same vanilla GRPO setup, our dataset alone outperforms the closed-source QwenLong-L1.5 dataset. Moreover, our Qwen3-30B-A3B model trained on this data delivers long-context performance comparable to DeepSeek-R1-0528 and Qwen3-235B-A22B-Thinking-2507, suggesting that broader coverage and greater reward diversity substantially benefit long-context capability improvement. (2) TMN-Reweight for heterogeneous multitask optimization. To address optimization challenges from heterogeneous rewards, we propose TMN-Reweight, which combines task-level mean normalization for cross-task reward scale alignment with difficulty-adaptive weighting for more reliable advantage estimation. TMN-Reweight further improves average performance over vanilla GRPO, with general capabilities preserved or improved across reported evaluations.
Qwen3 235B A22B Thinking 2507 is now available through local Ollama runtime. 40K context window listed. Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models.
LiveCodeBench pass@1 80.4 across 1055 tasks
LiveCodeBench pass@1 80.4 across 1055 tasks
Arena-Hard-Auto official Gemini-2.5 judged score 58.4 with CI -1.9/2.1
SWE-Bench Verified resolved rate 69.6
SWE-Bench Verified resolved rate 69.6