open-weights
taxonomy axis: openness
Decision rule: Weights released; data not released.
Anchor: Llama 3 (report states weights release; data composition disclosed but not released)
Models in this category (34)
| Model | Org | Date | Block | Attention | Context | Total | Active |
|---|---|---|---|---|---|---|---|
| Antares | Cisco Foundation AI | not disclosed | dense | GQA | 131,072 | 1B | 1B |
| Nanbeige4.2 | Nanbeige LLM Lab, Boss Zhipin | not disclosed | dense | not disclosed | 262,144 | 3B | 3B |
| Laguna | poolside | not disclosed | sparse-MoE | hybrid | 131,072 | 33.4B | 3B |
| Mixtral | Mistral AI | 2024-01-08 | sparse-MoE | GQA | 32,768 | 47B | 13B |
| Jamba | AI21 Labs | 2024-03-28 | hybrid | GQA | 1M | 52B | 12B |
| Llama 3 | Llama Team, AI @ Meta | 2024-07-23 | dense | GQA | 131,072 | 405B | 405B |
| Llama 3.1 | Meta (Llama Team, AI @ Meta) | 2024-07-31 | dense | GQA | 131,072 | 405B | 405B |
| DeepSeek-V3 | DeepSeek-AI | 2024-12-27 | sparse-MoE | MLA | 4,096 | 671B | 37B |
| DeepSeek-R1 | DeepSeek-AI | 2025-01-22 | sparse-MoE | 671B | 37B | ||
| xLSTM-7B | NXAI (NX-AI) | 2025-03-17 | dense | linear/state-space | 8,192 | 6.87B | 6.87B |
| Gemma 3 | Google DeepMind (Gemma Team) | 2025-03-25 | dense | sliding-window | 131,072 | 27B | 27B |
| Qwen3 | Qwen Team | 2025-05-15 | dense | GQA | 32,768 | 32B | 32B |
| Qwen3 | Qwen Team | 2025-05-15 | sparse-MoE | GQA | 32,768 | 235B | 22B |
| Kimi K2 | Kimi Team (Moonshot AI) | 2025-07-28 | sparse-MoE | MLA | 4,096 | 1.043T | 32.6B |
| gpt-oss | OpenAI | 2025-08-05 | sparse-MoE | hybrid | not disclosed | 116.83B | 5.13B |
| GLM-4.5 (ARC series) | Zhipu AI & Tsinghua University (GLM-4.5 Team) | 2025-08-08 | sparse-MoE | GQA | 131,072 | 106B | 12B |
| GLM-4.5 (ARC series) | Zhipu AI & Tsinghua University (GLM-4.5 Team) | 2025-08-08 | sparse-MoE | GQA | 131,072 | 355B | 32B |
| Kimi Linear | Kimi Team (Moonshot AI) | 2025-11-01 | sparse-MoE | hybrid | 4,096 | 48B | 3B |
| DeepSeek-V3.2 | DeepSeek-AI | 2025-12-02 | sparse-MoE | MLA | 131,072 | not disclosed | not disclosed |
| MiMo-V2 | Xiaomi (LLM-Core) | 2026-01-06 | sparse-MoE | hybrid | 262,144 | 309B | 15B |
| LongCat-Flash | Meituan LongCat Team | 2026-01-29 | sparse-MoE | not disclosed | 131,072 | 68.5B | 2.9B |
| Nanbeige4.1 | Nanbeige LLM Lab, Boss Zhipin | 2026-02-13 | dense | not disclosed | 262,144 | 3B | 3B |
| GLM-5 | Zhipu AI & Tsinghua University (GLM-5 Team) | 2026-02-17 | sparse-MoE | MLA | 200K | 744B | 40B |
| Trinity | Arcee AI | 2026-02-19 | sparse-MoE | hybrid | 262,144 | 400B | 13B |
| Qwen3-Coder-Next | Qwen Team | 2026-02-28 | hybrid | hybrid | 262,144 | 80B | 3B |
| Tiny Aya | Cohere Labs / Cohere | 2026-03-12 | dense | hybrid | 8,192 | 3.35B | 3.35B |
| DeepSeek-V4 | DeepSeek-AI | 2026-04-26 | sparse-MoE | hybrid | 1M | 1.6T | 49B |
| MiniMax-M2 series | MiniMax | 2026-05-26 | sparse-MoE | GQA | 192K | 229.9B | 9.8B |
| Mellum 2 | JetBrains (with Constructor University, Bremen) | 2026-05-29 | sparse-MoE | hybrid | 131,072 | 12B | 2.5B |
| MiniMax M3 | MiniMax (with authors from Peking University, NVIDIA, Zhejiang University, HUST, Nanjing University, Hangzhou Dianzi University) | 2026-06-11 | sparse-MoE | GQA | not disclosed | 109B | 6B |
| Ling 2.6 | Ling Team, Inclusion AI | 2026-06-13 | sparse-MoE | hybrid | 262,144 | 1T | not disclosed |
| Solar Open 2 | Upstage (Upstage Solar Team) | 2026-07-22 | sparse-MoE | hybrid | 1,048,576 | 250B | 15B |
| Gemma 4 | Google DeepMind (Gemma Team) | 2026-07-24 | dense | hybrid | not disclosed | 31.25B | 31.25B |
| Kimi K3 | Kimi Team (Moonshot AI) | 2026-07-27 | sparse-MoE | hybrid | 1M | 2.78T | 104.2B |
What varies within this category
attention variant varies — GQA, MLA, hybrid, linear/state-space, not disclosed, sliding-window. block type varies — dense, hybrid, sparse-MoE. trained context varies — 1000000, 1048576, 131072, 192000, 200000, 262144, 32768, 4096, 8192, not disclosed. scale class varies — frontier, large, medium, not disclosed.