open-weights-open-data

taxonomy axis: openness

Decision rule: Weights **and** training data released.

Anchor: OLMo

Models in this category (6)

ModelOrgDateBlockAttentionContextTotalActive
INTELLECT-3Prime Intellect (Prime Intellect, Inc.)not disclosedsparse-MoEnot disclosed65,536106B12B
OLMo 2OLMo Team, Allen Institute for AI (Ai2)2024-12-31denseMHA4,0967B7B
OLMo 3Allen Institute for AI (Olmo Team)2025-12-15densehybrid8,19232B32B
Nemotron 3 NanoNVIDIA2025-12-23sparse-MoEGQA524,28831.6B3.2B
Nemotron 3 SuperNVIDIA2026-04-03sparse-MoEGQA1,048,576120.6B12.7B
Nemotron 3 Ultra (Nemotron 3 family)NVIDIA2026-06-09sparse-MoEhybrid1,048,576550B55B

What varies within this category

attention variant varies — GQA, MHA, hybrid, not disclosed. block type varies — dense, sparse-MoE. trained context varies — 1048576, 4096, 524288, 65536, 8192. scale class varies — frontier, large, medium.