reasoning-specialised

taxonomy axis: design_intent

Decision rule: Framed around reasoning/thinking (CoT, RLVR, test-time compute).

Anchor: o1 (if its report is ever in the atlas)

Models in this category (10)

ModelOrgDateBlockAttentionContextTotalActive
INTELLECT-3Prime Intellect (Prime Intellect, Inc.)not disclosedsparse-MoEnot disclosed65,536106B12B
Phi-4Microsoft Research2024-12-12denseMHA4,09614B14B
DeepSeek-R1DeepSeek-AI2025-01-22sparse-MoE671B37B
gpt-ossOpenAI2025-08-05sparse-MoEhybridnot disclosed116.83B5.13B
DeepSeek-V3.2DeepSeek-AI2025-12-02sparse-MoEMLA131,072not disclosednot disclosed
Nemotron 3 NanoNVIDIA2025-12-23sparse-MoEGQA524,28831.6B3.2B
MiMo-V2Xiaomi (LLM-Core)2026-01-06sparse-MoEhybrid262,144309B15B
Nemotron 3 SuperNVIDIA2026-04-03sparse-MoEGQA1,048,576120.6B12.7B
ZAYA1Zyphra2026-05-06sparse-MoEGQA131,0728.4B760M
VibeThinkerWeibo2026-06-15densenot disclosed65,5363B3B

What varies within this category

attention variant varies — GQA, MHA, MLA, hybrid, not disclosed. block type varies — dense, sparse-MoE. trained context varies — 1048576, 131072, 262144, 4096, 524288, 65536, not disclosed. openness varies — open-weights, open-weights-open-data, undisclosed. scale class varies — frontier, large, medium, not disclosed.