large

taxonomy axis: scale_class

Decision rule: 10B ≤ p < 100B

Anchor: Llama 3.1 70B

Models in this category (11)

ModelOrgDateBlockAttentionContextTotalActive
MixtralMistral AI2024-01-08sparse-MoEGQA32,76847B13B
JambaAI21 Labs2024-03-28hybridGQA1M52B12B
Gemma 3Google DeepMind (Gemma Team)2025-03-25densesliding-window131,07227B27B
Qwen3Qwen Team2025-05-15denseGQA32,76832B32B
Kimi LinearKimi Team (Moonshot AI)2025-11-01sparse-MoEhybrid4,09648B3B
OLMo 3Allen Institute for AI (Olmo Team)2025-12-15densehybrid8,19232B32B
Nemotron 3 NanoNVIDIA2025-12-23sparse-MoEGQA524,28831.6B3.2B
LongCat-FlashMeituan LongCat Team2026-01-29sparse-MoEnot disclosed131,07268.5B2.9B
Qwen3-Coder-NextQwen Team2026-02-28hybridhybrid262,14480B3B
Mellum 2JetBrains (with Constructor University, Bremen)2026-05-29sparse-MoEhybrid131,07212B2.5B
Gemma 4Google DeepMind (Gemma Team)2026-07-24densehybridnot disclosed31.25B31.25B

What varies within this category

attention variant varies — GQA, hybrid, not disclosed, sliding-window. block type varies — dense, hybrid, sparse-MoE. trained context varies — 1000000, 131072, 262144, 32768, 4096, 524288, 8192, not disclosed. openness varies — open-weights, open-weights-open-data.