Qwen/Qwen3-8B-Base · Hugging Face
Qwen published benchmark or leaderboard evidence for Qwen3-8B-Base.
View sourceQwen
Qwen3-8B-Base is a open-weight Qwen llm model with a 131,072 token context window.
Running this yourself: desktop gpu should be enough.
Live access is not confirmed for this model. No current purchase price is advertised.
No current subscription pricing is tracked for this model.
Confirm this specific model, usage limits, and billing terms with the provider. A subscription does not automatically include API credits.
No login needed to compare. Prices are in USD; provider charges are separate from AI Market Cap plans. Context length, caching, tools, taxes, and regional terms can change the final cost. Open weights do not mean free hosting.
40.8
Quality Score
---
Arena ELO
8B
Parameters
131K
Context
This measures the amount of verifiable public evidence we have, not how capable the model is. A missing field means it has not been verified yet, not that its value is zero.
16 of 22 public signals
Sign in to join the discussion
768.0K
Downloads
138
Likes
Apr 2025
Released
4/5 signals
3/4 signals
4/5 signals
3/4 signals
2/4 signals
Gaps we are still tracking
Benchmarks
2
Open Source
1
Research
1
Recent launch, pricing, benchmark, and API signals linked to this model or its provider.
Qwen published benchmark or leaderboard evidence for Qwen3-8B-Base.
View sourceGAIA score 44.2 from WA0824
View sourceQwen3-8B-Base is now available through local Ollama runtime. 256K context window listed. Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks.
View sourceAs LLMs are increasingly used for pre-submission self-review, there is growing demand for feedback that not only identifies weaknesses but also guides authors toward concrete revisions. We study this as Actionable Peer-review Generation and decompose it into two subtasks: diagnostic claim generation and revision suggestion generation. We introduce ActReview, a rebuttal-guided post-training framework that connects paper-specific diagnoses to concrete, grounded revision plans. Our central insight is that author rebuttals reveal plausible actions for addressing reviewer concerns and can therefore provide latent supervision for revision-oriented feedback. From real review-rebuttal threads on OpenReview, we construct ActReview-40K by aligning reviewer weaknesses with author responses and grounding the resulting feedback in localized paper evidence. We post-train Qwen3-8B-Base with multi-task supervised fine-tuning followed by GRPO using candidate-aware, weakness-specific rubric rewards. We also introduce ActReview-Bench, a human-curated benchmark of 1,000 instances for evaluating diagnostic quality and revision usefulness. Experiments show that ActReview outperforms prior specialized review-generation models on actionability and grounding while remaining competitive with strong prompt-based LLMs. Human evaluation confirms improved revision usefulness while revealing a remaining gap in technical accuracy, and additional analyses support generalization to held-out papers and robustness across independent judges.
View sourceAs LLMs are increasingly used for pre-submission self-review, there is growing demand for feedback that not only identifies weaknesses but also guides authors toward concrete revisions. We study this as Actionable Peer-review Generation and decompose it into two subtasks: diagnostic claim generation and revision suggestion generation. We introduce ActReview, a rebuttal-guided post-training framework that connects paper-specific diagnoses to concrete, grounded revision plans. Our central insight is that author rebuttals reveal plausible actions for addressing reviewer concerns and can therefore provide latent supervision for revision-oriented feedback. From real review-rebuttal threads on OpenReview, we construct ActReview-40K by aligning reviewer weaknesses with author responses and grounding the resulting feedback in localized paper evidence. We post-train Qwen3-8B-Base with multi-task supervised fine-tuning followed by GRPO using candidate-aware, weakness-specific rubric rewards. We also introduce ActReview-Bench, a human-curated benchmark of 1,000 instances for evaluating diagnostic quality and revision usefulness. Experiments show that ActReview outperforms prior specialized review-generation models on actionability and grounding while remaining competitive with strong prompt-based LLMs. Human evaluation confirms improved revision usefulness while revealing a remaining gap in technical accuracy, and additional analyses support generalization to held-out papers and robustness across independent judges.
Qwen3-8B-Base is now available through local Ollama runtime. 256K context window listed. Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks.
Qwen published benchmark or leaderboard evidence for Qwen3-8B-Base.