Qwen3 VL 8B Instruct Benchmark Update
Quality: 7.3/100 | Price: $0.31/M tokens | Output: 0 tok/s | MMLU: 0.686% | HumanEval: 0.332%
View sourceQwen
Qwen3-VL-8B-Instruct is a multimodal vision-language model from the Qwen3-VL series, built for high-fidelity understanding and reasoning across text, images, and video. It features improved multimodal fusion with Interleaved-MRoPE for long-horizon...
Running this yourself: desktop gpu should be enough.
OpenRouter
Price record: 2026-10-03. Source: openrouter.
No current subscription pricing is tracked for this model.
Confirm this specific model, usage limits, and billing terms with the provider. A subscription does not automatically include API credits.
No login needed to compare. Prices are in USD; provider charges are separate from AI Market Cap plans. Context length, caching, tools, taxes, and regional terms can change the final cost. Open weights do not mean free hosting.
49.3
Quality Score
---
Arena ELO
9B
Parameters
262K
Context
This measures the amount of verifiable public evidence we have, not how capable the model is. A missing field means it has not been verified yet, not that its value is zero.
18 of 22 public signals
Sign in to join the discussion
15.8M
Downloads
1.2K
Likes
Oct 2025
Released
5/5 signals
3/4 signals
5/5 signals
3/4 signals
2/4 signals
Gaps we are still tracking
Benchmarks
18
Open Source
1
Research
1
Recent launch, pricing, benchmark, and API signals linked to this model or its provider.
Quality: 7.3/100 | Price: $0.31/M tokens | Output: 0 tok/s | MMLU: 0.686% | HumanEval: 0.332%
View sourceQuality: 7.3/100 | Price: $0.31/M tokens | Output: 0 tok/s | MMLU: 0.686% | HumanEval: 0.332%
Quality: 7.3/100 | Price: $0.31/M tokens | Output: 0 tok/s | MMLU: 0.686% | HumanEval: 0.332%
Quality: 7.3/100 | Price: $0.31/M tokens | Output: 0 tok/s | MMLU: 0.686% | HumanEval: 0.332%
View sourceQuality: 7.3/100 | Price: $0.31/M tokens | Output: 0 tok/s | MMLU: 0.686% | HumanEval: 0.332%
View sourceQuality: 7.3/100 | Price: $0.31/M tokens | Output: 0 tok/s | MMLU: 0.686% | HumanEval: 0.332%
View sourceQuality: 7.3/100 | Price: $0.31/M tokens | Output: 0 tok/s | MMLU: 0.686% | HumanEval: 0.332%
Quality: 7.3/100 | Price: $0.31/M tokens | Output: 0 tok/s | MMLU: 0.686% | HumanEval: 0.332%
Quality: 7.3/100 | Price: $0.31/M tokens | Output: 0 tok/s | MMLU: 0.686% | HumanEval: 0.332%
Quality: 7.3/100 | Price: $0.31/M tokens | Output: 0 tok/s | MMLU: 0.686% | HumanEval: 0.332%
Quality: 7.3/100 | Price: $0.31/M tokens | Output: 0 tok/s | MMLU: 0.686% | HumanEval: 0.332%
Quality: 7.3/100 | Price: $0.31/M tokens | Output: 0 tok/s | MMLU: 0.686% | HumanEval: 0.332%
Quality: 7.3/100 | Price: $0.31/M tokens | Output: 0 tok/s | MMLU: 0.686% | HumanEval: 0.332%
Quality: 7.3/100 | Price: $0.31/M tokens | Output: 0 tok/s | MMLU: 0.686% | HumanEval: 0.332%
Quality: 7.3/100 | Price: $0.31/M tokens | Output: 0 tok/s | MMLU: 0.686% | HumanEval: 0.332%
Quality: 7.3/100 | Price: $0.31/M tokens | Output: 0 tok/s | MMLU: 0.686% | HumanEval: 0.332%
Quality: 7.3/100 | Price: $0.31/M tokens | Output: 0 tok/s | MMLU: 0.686% | HumanEval: 0.332%
Quality: 7.3/100 | Price: $0.31/M tokens | Output: 0 tok/s | MMLU: 0.686% | HumanEval: 0.332%
Quality: 7.3/100 | Price: $0.31/M tokens | Output: 0 tok/s | MMLU: 0.686% | HumanEval: 0.332%
Quality: 7.3/100 | Price: $0.31/M tokens | Output: 0 tok/s | MMLU: 0.686% | HumanEval: 0.332%
Quality: 7.3/100 | Price: $0.31/M tokens | Output: 0 tok/s | MMLU: 0.686% | HumanEval: 0.332%
Quality: 7.3/100 | Price: $0.31/M tokens | Output: 0 tok/s | MMLU: 0.686% | HumanEval: 0.332%
Quality: 7.3/100 | Price: $0.31/M tokens | Output: 0 tok/s | MMLU: 0.686% | HumanEval: 0.332%
Real-world embodied tasks, from everyday activities to professional procedures, require agents to act under physical constraints while tracking evolving object and task states. Tool use sits at the heart of such tasks, as many everyday and professional activities are tool-mediated. Understanding them requires reasoning about affordances, hand-tool-object geometry, procedural progress, and causal effects on target objects. Yet despite strong performance on perception-oriented video tasks such as captioning and general video QA, current multimodal video models remain limited in this form of tool-centric embodied reasoning. Progress in this direction has been limited by the lack of real-world egocentric data and diagnostic benchmarks. To address this gap, we introduce EgoTools, the first comprehensive suite for egocentric tool-use understanding. It consists of two complementary components: EgoTools-Data, a large-scale corpus of 100 hours of tool-centric egocentric recordings with synchronized audio, dense captions, reasoning-heavy narrations, and supplementary 3D information; and EgoTools-Bench, a diagnostic benchmark of 1,000 QA pairs across four tracks that cover tool-use understanding from perception and geometry to procedure and causal reasoning. Experimental results show that current models still struggle to ground tool use in visual evidence: Gemini-3.1-Pro achieves 66.9% overall accuracy but only 51.7% on Perception & Grounding. Beyond evaluation, we validate EgoTools-Data as a training resource. On the full 1,000-question benchmark, full supervised fine-tuning improves Qwen3-VL-8B-Instruct from 50.0% to 60.9%, under strict source-video separation. Together, these results establish EgoTools as a unified resource for both training and diagnostic evaluation of real-world egocentric tool-use understanding.
Qwen3 VL 8B Instruct is now available through local Ollama runtime. 40K context window listed. Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models.