hybrid
taxonomy axis: sequence_mixing
Decision rule: Mixed attention + SSM/linear layers (or attention variants in a deliberate per-layer pattern).
Anchor: Jamba (1 attention : 7 Mamba layers)
Models in this category (14)
| Model | Org | Date | Block | Attention | Context | Total | Active |
|---|---|---|---|---|---|---|---|
| Laguna | poolside | not disclosed | sparse-MoE | hybrid | 131,072 | 33.4B | 3B |
| Jamba | AI21 Labs | 2024-03-28 | hybrid | GQA | 1M | 52B | 12B |
| gpt-oss | OpenAI | 2025-08-05 | sparse-MoE | hybrid | not disclosed | 116.83B | 5.13B |
| Kimi Linear | Kimi Team (Moonshot AI) | 2025-11-01 | sparse-MoE | hybrid | 4,096 | 48B | 3B |
| OLMo 3 | Allen Institute for AI (Olmo Team) | 2025-12-15 | dense | hybrid | 8,192 | 32B | 32B |
| Trinity | Arcee AI | 2026-02-19 | sparse-MoE | hybrid | 262,144 | 400B | 13B |
| Qwen3-Coder-Next | Qwen Team | 2026-02-28 | hybrid | hybrid | 262,144 | 80B | 3B |
| Tiny Aya | Cohere Labs / Cohere | 2026-03-12 | dense | hybrid | 8,192 | 3.35B | 3.35B |
| Nemotron 3 Super | NVIDIA | 2026-04-03 | sparse-MoE | GQA | 1,048,576 | 120.6B | 12.7B |
| Qwen3.5 | Qwen Team | 2026-04-21 | hybrid | hybrid | 262,144 | not disclosed | not disclosed |
| Mellum 2 | JetBrains (with Constructor University, Bremen) | 2026-05-29 | sparse-MoE | hybrid | 131,072 | 12B | 2.5B |
| Nemotron 3 Ultra (Nemotron 3 family) | NVIDIA | 2026-06-09 | sparse-MoE | hybrid | 1,048,576 | 550B | 55B |
| Solar Open 2 | Upstage (Upstage Solar Team) | 2026-07-22 | sparse-MoE | hybrid | 1,048,576 | 250B | 15B |
| Kimi K3 | Kimi Team (Moonshot AI) | 2026-07-27 | sparse-MoE | hybrid | 1M | 2.78T | 104.2B |
What varies within this category
attention variant varies — GQA, hybrid. block type varies — dense, hybrid, sparse-MoE. trained context varies — 1000000, 1048576, 131072, 262144, 4096, 8192, not disclosed. openness varies — open-weights, open-weights-open-data, undisclosed. scale class varies — frontier, large, medium.