large
taxonomy axis: scale_class
Decision rule: 10B ≤ p < 100B
Anchor: Llama 3.1 70B
Models in this category (11)
| Model | Org | Date | Block | Attention | Context | Total | Active |
|---|---|---|---|---|---|---|---|
| Mixtral | Mistral AI | 2024-01-08 | sparse-MoE | GQA | 32,768 | 47B | 13B |
| Jamba | AI21 Labs | 2024-03-28 | hybrid | GQA | 1M | 52B | 12B |
| Gemma 3 | Google DeepMind (Gemma Team) | 2025-03-25 | dense | sliding-window | 131,072 | 27B | 27B |
| Qwen3 | Qwen Team | 2025-05-15 | dense | GQA | 32,768 | 32B | 32B |
| Kimi Linear | Kimi Team (Moonshot AI) | 2025-11-01 | sparse-MoE | hybrid | 4,096 | 48B | 3B |
| OLMo 3 | Allen Institute for AI (Olmo Team) | 2025-12-15 | dense | hybrid | 8,192 | 32B | 32B |
| Nemotron 3 Nano | NVIDIA | 2025-12-23 | sparse-MoE | GQA | 524,288 | 31.6B | 3.2B |
| LongCat-Flash | Meituan LongCat Team | 2026-01-29 | sparse-MoE | not disclosed | 131,072 | 68.5B | 2.9B |
| Qwen3-Coder-Next | Qwen Team | 2026-02-28 | hybrid | hybrid | 262,144 | 80B | 3B |
| Mellum 2 | JetBrains (with Constructor University, Bremen) | 2026-05-29 | sparse-MoE | hybrid | 131,072 | 12B | 2.5B |
| Gemma 4 | Google DeepMind (Gemma Team) | 2026-07-24 | dense | hybrid | not disclosed | 31.25B | 31.25B |
What varies within this category
attention variant varies — GQA, hybrid, not disclosed, sliding-window. block type varies — dense, hybrid, sparse-MoE. trained context varies — 1000000, 131072, 262144, 32768, 4096, 524288, 8192, not disclosed. openness varies — open-weights, open-weights-open-data.