LLM Architecture Atlas
One row per architecture, extracted from its technical report. Hover any cell for provenance. not disclosed means the report is silent — that is itself a finding.
MixtralJambaPhi-3Llama 3Llama 3.1Phi-4DeepSeek-V3OLMo 2DeepSeek-R1xLSTM-7BGemma 3Qwen3Qwen3Kimi K2gpt-ossGLM-4.5 (ARC series)GLM-4.5 (ARC series)Kimi LinearDeepSeek-V3.2OLMo 3Nemotron 3 NanoMiMo-V2LongCat-FlashNanbeige4.1GLM-5TrinityQwen3-Coder-NextTiny AyaNemotron 3 SuperQwen3.5DeepSeek-V4ZAYA1MiniMax-M2 seriesMellum 2Nemotron 3 Ultra (Nemotron 3 family)MiniMax M3Ling 2.6VibeThinkerSolar Open 2Gemma 4Kimi K3
2024-01-082026-07-27
| Model | Org | Date | Total | Active | Block | Attention | Pos enc | Context | Openness | Scale | Intent | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Jamba | AI21 Labs | 2024-03-28 | 52B | 12B | hybrid | GQA | NoPE | 1M | open-weights | large | frontier-generalist | |
| OLMo 2 | OLMo Team, Allen Institute for AI (Ai2) | 2024-12-31 | 7B | 7B | dense | MHA | RoPE | 4,096 | open-weights-open-data | medium | frontier-generalist | |
| OLMo 3 | Allen Institute for AI (Olmo Team) | 2025-12-15 | 32B | 32B | dense | hybrid | RoPE | 8,192 | open-weights-open-data | large | frontier-generalist | |
| Trinity | Arcee AI | 2026-02-19 | 400B | 13B | sparse-MoE | hybrid | RoPE | 262,144 | open-weights | frontier | frontier-generalist | |
| Antares | Cisco Foundation AI | 1B | 1B | dense | GQA | RoPE | 131,072 | open-weights | medium | domain-specific | ||
| Tiny Aya | Cohere Labs / Cohere | 2026-03-12 | 3.35B | 3.35B | dense | hybrid | other | 8,192 | open-weights | medium | on-device | |
| DeepSeek-R1 | DeepSeek-AI | 2025-01-22 | 671B | 37B | sparse-MoE | open-weights | frontier | reasoning-specialised | ||||
| DeepSeek-V3.2 | DeepSeek-AI | 2025-12-02 | sparse-MoE | MLA | 131,072 | open-weights | reasoning-specialised | |||||
| DeepSeek-V3 | DeepSeek-AI | 2024-12-27 | 671B | 37B | sparse-MoE | MLA | RoPE | 4,096 | open-weights | frontier | frontier-generalist | |
| DeepSeek-V4 | DeepSeek-AI | 2026-04-26 | 1.6T | 49B | sparse-MoE | hybrid | RoPE | 1M | open-weights | frontier | long-context | |
| Gemma 3 | Google DeepMind (Gemma Team) | 2025-03-25 | 27B | 27B | dense | sliding-window | RoPE | 131,072 | open-weights | large | on-device | |
| Gemma 4 | Google DeepMind (Gemma Team) | 2026-07-24 | 31.25B | 31.25B | dense | hybrid | RoPE | open-weights | large | on-device | ||
| Ling 2.6 | Ling Team, Inclusion AI | 2026-06-13 | 1T | sparse-MoE | hybrid | RoPE | 262,144 | open-weights | frontier | frontier-generalist | ||
| Mellum 2 | JetBrains (with Constructor University, Bremen) | 2026-05-29 | 12B | 2.5B | sparse-MoE | hybrid | RoPE | 131,072 | open-weights | large | domain-specific | |
| LongCat-Flash | Meituan LongCat Team | 2026-01-29 | 68.5B | 2.9B | sparse-MoE | RoPE | 131,072 | open-weights | large | long-context | ||
| Llama 3.1 | Meta (Llama Team, AI @ Meta) | 2024-07-31 | 405B | 405B | dense | GQA | RoPE | 131,072 | open-weights | frontier | frontier-generalist | |
| Llama 3 | Llama Team, AI @ Meta | 2024-07-23 | 405B | 405B | dense | GQA | RoPE | 131,072 | open-weights | frontier | frontier-generalist | |
| Phi-3 | Microsoft | 2024-04-22 | 3.8B | 3.8B | dense | MHA | 4,096 | undisclosed | medium | on-device | ||
| Phi-4 | Microsoft Research | 2024-12-12 | 14B | 14B | dense | MHA | 4,096 | undisclosed | medium | reasoning-specialised | ||
| MiniMax-M2 series | MiniMax | 2026-05-26 | 229.9B | 9.8B | sparse-MoE | GQA | RoPE | 192K | open-weights | frontier | frontier-generalist | |
| MiniMax M3 | MiniMax (with authors from Peking University, NVIDIA, Zhejiang University, HUST, Nanjing University, Hangzhou Dianzi University) | 2026-06-11 | 109B | 6B | sparse-MoE | GQA | RoPE | open-weights | frontier | long-context | ||
| Mixtral | Mistral AI | 2024-01-08 | 47B | 13B | sparse-MoE | GQA | 32,768 | open-weights | large | frontier-generalist | ||
| Kimi K2 | Kimi Team (Moonshot AI) | 2025-07-28 | 1.043T | 32.6B | sparse-MoE | MLA | RoPE | 4,096 | open-weights | frontier | frontier-generalist | |
| Kimi K3 | Kimi Team (Moonshot AI) | 2026-07-27 | 2.78T | 104.2B | sparse-MoE | hybrid | NoPE | 1M | open-weights | frontier | frontier-generalist | |
| Kimi Linear | Kimi Team (Moonshot AI) | 2025-11-01 | 48B | 3B | sparse-MoE | hybrid | NoPE | 4,096 | open-weights | large | frontier-generalist | |
| Nanbeige4.1 | Nanbeige LLM Lab, Boss Zhipin | 2026-02-13 | 3B | 3B | dense | 262,144 | open-weights | medium | frontier-generalist | |||
| Nanbeige4.2 | Nanbeige LLM Lab, Boss Zhipin | 3B | 3B | dense | 262,144 | open-weights | medium | frontier-generalist | ||||
| Nemotron 3 Nano | NVIDIA | 2025-12-23 | 31.6B | 3.2B | sparse-MoE | GQA | NoPE | 524,288 | open-weights-open-data | large | reasoning-specialised | |
| Nemotron 3 Super | NVIDIA | 2026-04-03 | 120.6B | 12.7B | sparse-MoE | GQA | NoPE | 1,048,576 | open-weights-open-data | frontier | reasoning-specialised | |
| Nemotron 3 Ultra (Nemotron 3 family) | NVIDIA | 2026-06-09 | 550B | 55B | sparse-MoE | hybrid | 1,048,576 | open-weights-open-data | frontier | frontier-generalist | ||
| xLSTM-7B | NXAI (NX-AI) | 2025-03-17 | 6.87B | 6.87B | dense | linear/state-space | NoPE | 8,192 | open-weights | medium | frontier-generalist | |
| GPT-2 | OpenAI | 1.542B | 1.542B | dense | 1,024 | undisclosed | medium | frontier-generalist | ||||
| gpt-oss | OpenAI | 2025-08-05 | 116.83B | 5.13B | sparse-MoE | hybrid | RoPE | open-weights | frontier | reasoning-specialised | ||
| Laguna | poolside | 33.4B | 3B | sparse-MoE | hybrid | RoPE | 131,072 | open-weights | medium | domain-specific | ||
| INTELLECT-3 | Prime Intellect (Prime Intellect, Inc.) | 106B | 12B | sparse-MoE | 65,536 | open-weights-open-data | frontier | reasoning-specialised | ||||
| Qwen3.5 | Qwen Team | 2026-04-21 | hybrid | hybrid | RoPE | 262,144 | undisclosed | frontier | frontier-generalist | |||
| Qwen3 | Qwen Team | 2025-05-15 | 32B | 32B | dense | GQA | RoPE | 32,768 | open-weights | large | frontier-generalist | |
| Qwen3 | Qwen Team | 2025-05-15 | 235B | 22B | sparse-MoE | GQA | RoPE | 32,768 | open-weights | frontier | frontier-generalist | |
| Qwen3-Coder-Next | Qwen Team | 2026-02-28 | 80B | 3B | hybrid | hybrid | 262,144 | open-weights | large | domain-specific | ||
| Solar Open 2 | Upstage (Upstage Solar Team) | 2026-07-22 | 250B | 15B | sparse-MoE | hybrid | NoPE | 1,048,576 | open-weights | frontier | frontier-generalist | |
| VibeThinker | 2026-06-15 | 3B | 3B | dense | 65,536 | undisclosed | medium | reasoning-specialised | ||||
| MiMo-V2 | Xiaomi (LLM-Core) | 2026-01-06 | 309B | 15B | sparse-MoE | hybrid | RoPE | 262,144 | open-weights | frontier | reasoning-specialised | |
| GLM-4.5 (ARC series) | Zhipu AI & Tsinghua University (GLM-4.5 Team) | 2025-08-08 | 106B | 12B | sparse-MoE | GQA | RoPE | 131,072 | open-weights | frontier | frontier-generalist | |
| GLM-4.5 (ARC series) | Zhipu AI & Tsinghua University (GLM-4.5 Team) | 2025-08-08 | 355B | 32B | sparse-MoE | GQA | RoPE | 131,072 | open-weights | frontier | frontier-generalist | |
| GLM-5 | Zhipu AI & Tsinghua University (GLM-5 Team) | 2026-02-17 | 744B | 40B | sparse-MoE | MLA | 200K | open-weights | frontier | frontier-generalist | ||
| ZAYA1 | Zyphra | 2026-05-06 | 8.4B | 760M | sparse-MoE | GQA | RoPE | 131,072 | undisclosed | medium | reasoning-specialised |