Motif 3 Benchmark Update
Quality: 47.4/100 | Price: $0/M tokens | Output: 0 tok/s | HumanEval: 0.406%
View sourceMotif-Technologies
We introduce Motif 3, a decoder-only Mixture-of-Experts language model with 314 billion total parameters and 13. 2 billion parameters activated for each token. Each sparse MoE layer contains 384 routed experts, while only eight are selected for each token.
Running this yourself: likely needs a high-memory cloud gpu.
---
Quality Score
---
Arena ELO
315B
Parameters
262K
Context
This measures the amount of verifiable public evidence we have, not how capable the model is. A missing field means it has not been verified yet, not that its value is zero.
18 of 22 public signals
Sign in to join the discussion
4.4K
Downloads
125
Likes
Aug 2026
Released
5/5 signals
2/4 signals
3/5 signals
4/4 signals
4/4 signals
Parameters
314B
Training compute
1.1e24 FLOP
Dataset scale
Not reported
Base model
Not reported
Source-reported access: Open weights (unrestricted) · Confident confidence
Gaps we are still tracking
Benchmarks
14
Research
1
Recent launch, pricing, benchmark, and API signals linked to this model or its provider.
Quality: 47.4/100 | Price: $0/M tokens | Output: 0 tok/s | HumanEval: 0.406%
View sourceQuality: 47.4/100 | Price: $0/M tokens | Output: 0 tok/s | HumanEval: 0.406%
Quality: 47.4/100 | Price: $0/M tokens | Output: 0 tok/s | HumanEval: 0.406%
Quality: 47.4/100 | Price: $0/M tokens | Output: 0 tok/s | HumanEval: 0.406%
View sourceQuality: 47.4/100 | Price: $0/M tokens | Output: 0 tok/s | HumanEval: 0.406%
View sourceQuality: 47.4/100 | Price: $0/M tokens | Output: 0 tok/s | HumanEval: 0.406%
View sourceQuality: 47.4/100 | Price: $0/M tokens | Output: 0 tok/s | HumanEval: 0.406%
Quality: 47.4/100 | Price: $0/M tokens | Output: 0 tok/s | HumanEval: 0.406%
Quality: 47.4/100 | Price: $0/M tokens | Output: 0 tok/s | HumanEval: 0.406%
Quality: 47.4/100 | Price: $0/M tokens | Output: 0 tok/s | HumanEval: 0.406%
Quality: 47.4/100 | Price: $0/M tokens | Output: 0 tok/s | HumanEval: 0.406%
Quality: 47.4/100 | Price: $0/M tokens | Output: 0 tok/s | HumanEval: 0.406%
Quality: 47.4/100 | Price: $0/M tokens | Output: 0 tok/s | HumanEval: 0.406%
Quality: 47.4/100 | Price: $0/M tokens | Output: 0 tok/s | HumanEval: 0.406%
Quality: 47.4/100 | Price: $0/M tokens | Output: 0 tok/s | HumanEval: 0.406%
Quality: 47.4/100 | Price: $0/M tokens | Output: 0 tok/s | HumanEval: 0.406%
Quality: 47.4/100 | Price: $0/M tokens | Output: 0 tok/s | HumanEval: 0.406%
Quality: 47.4/100 | Price: $0/M tokens | Output: 0 tok/s | HumanEval: 0.406%
Quality: 47.4/100 | Price: $0/M tokens | Output: 0 tok/s | HumanEval: 0.406%
We introduce Motif 3, a decoder-only Mixture-of-Experts language model with 314 billion total parameters and 13.2 billion activated per token. Each sparse MoE layer contains 384 routed experts, with eight selected per token. This fine-grained sparsity provides substantial expert capacity while limiting computation. Motif 3 is built around Grouped Differential Latent Attention (GDLA), which integrates grouped differential attention with the compressed key-value representation of Multi-head Latent Attention. The architecture further incorporates modified manifold-constrained hyper-connections, Expert Specific PolyNorm activations, and multi-token prediction to improve optimization stability, expert specialization, and inference efficiency. We pretrain Motif 3 on approximately 12.5 trillion tokens spanning web documents, STEM, code, mathematics, multilingual content, and domain-specialized corpora. Expert-balancing and numerical-stabilization techniques support stable training at scale, while selective MXFP8 computation and communication, memory-efficient fused kernels, and window-aware context parallelism enable training with context lengths up to 256K tokens. Our post-training pipeline combines general supervised fine-tuning, six specialist teachers trained with reinforcement learning, a software-engineering teacher trained with supervised fine-tuning, and Multi-teacher On-Policy Distillation. The resulting unified model consolidates complementary capabilities in reasoning, coding, tool use, professional work, long-context understanding, calibrated abstention, and instruction following. Across a broad evaluation suite, Motif 3 demonstrates competitive performance against leading open weight models, including strong results on long-horizon agentic tasks, mathematical reasoning, scientific knowledge, and hallucination-sensitive evaluation.