open-weights-open-data
taxonomy axis: openness
Decision rule: Weights **and** training data released.
Anchor: OLMo
Models in this category (6)
| Model | Org | Date | Block | Attention | Context | Total | Active |
|---|---|---|---|---|---|---|---|
| INTELLECT-3 | Prime Intellect (Prime Intellect, Inc.) | not disclosed | sparse-MoE | not disclosed | 65,536 | 106B | 12B |
| OLMo 2 | OLMo Team, Allen Institute for AI (Ai2) | 2024-12-31 | dense | MHA | 4,096 | 7B | 7B |
| OLMo 3 | Allen Institute for AI (Olmo Team) | 2025-12-15 | dense | hybrid | 8,192 | 32B | 32B |
| Nemotron 3 Nano | NVIDIA | 2025-12-23 | sparse-MoE | GQA | 524,288 | 31.6B | 3.2B |
| Nemotron 3 Super | NVIDIA | 2026-04-03 | sparse-MoE | GQA | 1,048,576 | 120.6B | 12.7B |
| Nemotron 3 Ultra (Nemotron 3 family) | NVIDIA | 2026-06-09 | sparse-MoE | hybrid | 1,048,576 | 550B | 55B |
What varies within this category
attention variant varies — GQA, MHA, hybrid, not disclosed. block type varies — dense, sparse-MoE. trained context varies — 1048576, 4096, 524288, 65536, 8192. scale class varies — frontier, large, medium.