Qwen/Qwen2.5-Coder-0.5B · Hugging Face
Qwen published benchmark or leaderboard evidence for Qwen2.5-Coder-0.5B.
View sourceQwen
Qwen2.5-Coder-0.5B is a open-weight Qwen llm model with a 32,768 token context window.
Running this yourself: can likely run on your own machine.
Live access is not confirmed for this model. No current purchase price is advertised.
No current subscription pricing is tracked for this model.
Confirm this specific model, usage limits, and billing terms with the provider. A subscription does not automatically include API credits.
No login needed to compare. Prices are in USD; provider charges are separate from AI Market Cap plans. Context length, caching, tools, taxes, and regional terms can change the final cost. Open weights do not mean free hosting.
23.9
Quality Score
---
Arena ELO
494M
Parameters
33K
Context
This measures the amount of verifiable public evidence we have, not how capable the model is. A missing field means it has not been verified yet, not that its value is zero.
16 of 22 public signals
Sign in to join the discussion
33.2K
Downloads
63
Likes
Nov 2024
Released
4/5 signals
3/4 signals
4/5 signals
3/4 signals
2/4 signals
Gaps we are still tracking
Benchmarks
4
Open Source
1
Research
1
Recent launch, pricing, benchmark, and API signals linked to this model or its provider.
Qwen published benchmark or leaderboard evidence for Qwen2.5-Coder-0.5B.
View sourceGAIA score 4.7 from rft-2
View sourceGAIA score 4.7 from rft-2
View sourceGAIA score 4.7 from rft-2
View sourceQwen2.5-Coder-0.5B is now available through local Ollama runtime. 32K context window listed. The latest series of Code-Specific Qwen models, with significant improvements in code generation, code reasoning, and code fixing.
View sourceLarge language models used for code editing can be trained and deployed in at least two output regimes: direct generation, where the model emits the entire modified file in one shot, and iterative diff-based generation ("steps"), where the model emits a sequence of localized search/replace edits applied one at a time until it signals completion or a step budget is exhausted. The diff-based regime is attractive because it mirrors how developers edit code and should require far fewer generated tokens per turn. We train two code models - a 100M-parameter model trained from scratch (Rainbow-Pony-100M) and a fine-tuned Qwen2.5-Coder-0.5B - in both regimes on a shared Flutter/Dart dataset, and evaluate all four resulting models on a held-out set of approx 1,790 tasks per model. Direct generation substantially outperforms diff-based generation on every metric we measure - compilation/static-analysis pass rate, bits-per-byte, character-level similarity to the reference, and blinded LLM-judge ratings of goal fulfillment, correctness, and code quality - and the gap persists after controlling for task difficulty via a matched-ID comparison and when restricting to code that compiles on both sides. We then identify a single, architecture-independent mechanism behind the conditions where diff-based generation does win: it is competitive on short, spatially localized edits, and its category-level wins concentrate in exactly the two task categories - refactoring and error-handling/edge-case fixes - with the lowest mean edit-step count in our dataset. We term this task locality and discuss its implications for when an edit-based training regime is and is not the right choice for a code-editing model.
Qwen2.5-Coder-0.5B is now available through local Ollama runtime. 32K context window listed. The latest series of Code-Specific Qwen models, with significant improvements in code generation, code reasoning, and code fixing.
Qwen published benchmark or leaderboard evidence for Qwen2.5-Coder-0.5B.