Qwen/Qwen2.5-Coder-7B · Hugging Face
Qwen published benchmark or leaderboard evidence for Qwen2.5-Coder-7B.
View sourceQwen
Significantly improvements in code generation, code reasoning and code fixing. Base on the strong Qwen2.5, we scale up the training tokens into 5.5 trillion including source code, text-code grounding, Synthetic data, etc.
Running this yourself: consumer gpu should be enough.
Live access is not confirmed for this model. No current purchase price is advertised.
No current subscription pricing is tracked for this model.
Confirm this specific model, usage limits, and billing terms with the provider. A subscription does not automatically include API credits.
No login needed to compare. Prices are in USD; provider charges are separate from AI Market Cap plans. Context length, caching, tools, taxes, and regional terms can change the final cost. Open weights do not mean free hosting.
26.5
Quality Score
---
Arena ELO
8B
Parameters
33K
Context
This measures the amount of verifiable public evidence we have, not how capable the model is. A missing field means it has not been verified yet, not that its value is zero.
18 of 22 public signals
Sign in to join the discussion
416.4K
Downloads
178
Likes
Sep 2024
Released
5/5 signals
3/4 signals
4/5 signals
3/4 signals
3/4 signals
Parameters
8B
Training compute
2.5e23 FLOP
Dataset scale
Not reported
Base model
Not reported
Source-reported access: Open weights (unrestricted) · Confident confidence
Gaps we are still tracking
Benchmarks
4
Open Source
1
Research
1
Recent launch, pricing, benchmark, and API signals linked to this model or its provider.
Qwen published benchmark or leaderboard evidence for Qwen2.5-Coder-7B.
View sourceGAIA score 4.7 from rft-2
View sourceGAIA score 4.7 from rft-2
View sourceGAIA score 4.7 from rft-2
View sourceQwen2.5-Coder-7B is now available through local Ollama runtime. 32K context window listed. The latest series of Code-Specific Qwen models, with significant improvements in code generation, code reasoning, and code fixing.
View sourceDevelopers can build LLM agents by adapting third-party models through benign post-training. We study a supply-chain threat in which an attacker supplies a model with a backdoor: hidden behavior that produces malicious outputs when a particular input pattern appears. Focusing on software-engineering agents, we ask whether such backdoors survive the developer's supervised fine-tuning (SFT) and subsequent task-level reinforcement learning (RL). We observe that benign SFT substantially reduces attack success, but subsequent RL often preserves the residual behavior and sometimes even increases attack success. Our analysis of backdoor erosion during SFT identifies two factors that may favor survival: initial backdoor strength and gradient compatibility with benign training. These factors motivate PersistBD, which refines an already-backdoored model before release to improve its persistency through the benign post-training process. On Qwen2.5-Coder-7B, PersistBD raises attack success from 20% to 74% after SFT and from 20% to 76% after SFT-RL, while maintaining comparable benign task performance. Together, our results show that backdoors can remain active through benign post-training and that adversaries can deliberately increase their persistence. This highlights a supply-chain risk for AI developers and motivates stronger techniques for detecting and mitigating inherited backdoors when adapting third-party models into agents. Our code is available at https://github.com/uiuc-kang-lab/PersistBD.
Qwen2.5-Coder-7B is now available through local Ollama runtime. 32K context window listed. The latest series of Code-Specific Qwen models, with significant improvements in code generation, code reasoning, and code fixing.
Qwen published benchmark or leaderboard evidence for Qwen2.5-Coder-7B.