efficient-attention
taxonomy axis: sequence_mixing
Decision rule: Attention with restricted/computed patterns (sliding window, linear attention, sparse patterns) — but still attention.
Anchor: Mistral (sliding window layers)
Models in this category (9)
| Model | Org | Date | Block | Attention | Context | Total | Active |
|---|---|---|---|---|---|---|---|
| Gemma 3 | Google DeepMind (Gemma Team) | 2025-03-25 | dense | sliding-window | 131,072 | 27B | 27B |
| DeepSeek-V3.2 | DeepSeek-AI | 2025-12-02 | sparse-MoE | MLA | 131,072 | not disclosed | not disclosed |
| MiMo-V2 | Xiaomi (LLM-Core) | 2026-01-06 | sparse-MoE | hybrid | 262,144 | 309B | 15B |
| GLM-5 | Zhipu AI & Tsinghua University (GLM-5 Team) | 2026-02-17 | sparse-MoE | MLA | 200K | 744B | 40B |
| DeepSeek-V4 | DeepSeek-AI | 2026-04-26 | sparse-MoE | hybrid | 1M | 1.6T | 49B |
| ZAYA1 | Zyphra | 2026-05-06 | sparse-MoE | GQA | 131,072 | 8.4B | 760M |
| MiniMax M3 | MiniMax (with authors from Peking University, NVIDIA, Zhejiang University, HUST, Nanjing University, Hangzhou Dianzi University) | 2026-06-11 | sparse-MoE | GQA | not disclosed | 109B | 6B |
| Ling 2.6 | Ling Team, Inclusion AI | 2026-06-13 | sparse-MoE | hybrid | 262,144 | 1T | not disclosed |
| Gemma 4 | Google DeepMind (Gemma Team) | 2026-07-24 | dense | hybrid | not disclosed | 31.25B | 31.25B |
What varies within this category
attention variant varies — GQA, MLA, hybrid, sliding-window. block type varies — dense, sparse-MoE. trained context varies — 1000000, 131072, 200000, 262144, not disclosed. openness varies — open-weights, undisclosed. scale class varies — frontier, large, medium, not disclosed.