multimodal
taxonomy axis: modality
Decision rule: Report describes vision/audio/other modalities in or out.
Anchor: GPT-4V (if its report is ever in the atlas — else first multimodal report ingested)
Models in this category (5)
| Model | Org | Date | Block | Attention | Context | Total | Active |
|---|---|---|---|---|---|---|---|
| Gemma 3 | Google DeepMind (Gemma Team) | 2025-03-25 | dense | sliding-window | 131,072 | 27B | 27B |
| MiniMax-M2 series | MiniMax | 2026-05-26 | sparse-MoE | GQA | 192K | 229.9B | 9.8B |
| MiniMax M3 | MiniMax (with authors from Peking University, NVIDIA, Zhejiang University, HUST, Nanjing University, Hangzhou Dianzi University) | 2026-06-11 | sparse-MoE | GQA | not disclosed | 109B | 6B |
| Gemma 4 | Google DeepMind (Gemma Team) | 2026-07-24 | dense | hybrid | not disclosed | 31.25B | 31.25B |
| Kimi K3 | Kimi Team (Moonshot AI) | 2026-07-27 | sparse-MoE | hybrid | 1M | 2.78T | 104.2B |
What varies within this category
attention variant varies — GQA, hybrid, sliding-window. block type varies — dense, sparse-MoE. trained context varies — 1000000, 131072, 192000, not disclosed. scale class varies — frontier, large.