full-attention

taxonomy axis: sequence_mixing

Decision rule: All layers use unrestricted attention over the whole context.

Anchor: Llama 3

Models in this category (18)

ModelOrgDateBlockAttentionContextTotalActive
AntaresCisco Foundation AInot discloseddenseGQA131,0721B1B
GPT-2OpenAInot discloseddense1,0241.542B1.542B
INTELLECT-3Prime Intellect (Prime Intellect, Inc.)not disclosedsparse-MoEnot disclosed65,536106B12B
MixtralMistral AI2024-01-08sparse-MoEGQA32,76847B13B
Phi-3Microsoft2024-04-22denseMHA4,0963.8B3.8B
Llama 3Llama Team, AI @ Meta2024-07-23denseGQA131,072405B405B
Llama 3.1Meta (Llama Team, AI @ Meta)2024-07-31denseGQA131,072405B405B
Phi-4Microsoft Research2024-12-12denseMHA4,09614B14B
DeepSeek-V3DeepSeek-AI2024-12-27sparse-MoEMLA4,096671B37B
OLMo 2OLMo Team, Allen Institute for AI (Ai2)2024-12-31denseMHA4,0967B7B
DeepSeek-R1DeepSeek-AI2025-01-22sparse-MoE671B37B
Qwen3Qwen Team2025-05-15denseGQA32,76832B32B
Qwen3Qwen Team2025-05-15sparse-MoEGQA32,768235B22B
Kimi K2Kimi Team (Moonshot AI)2025-07-28sparse-MoEMLA4,0961.043T32.6B
GLM-4.5 (ARC series)Zhipu AI & Tsinghua University (GLM-4.5 Team)2025-08-08sparse-MoEGQA131,072106B12B
GLM-4.5 (ARC series)Zhipu AI & Tsinghua University (GLM-4.5 Team)2025-08-08sparse-MoEGQA131,072355B32B
MiniMax-M2 seriesMiniMax2026-05-26sparse-MoEGQA192K229.9B9.8B
VibeThinkerWeibo2026-06-15densenot disclosed65,5363B3B

What varies within this category

attention variant varies — GQA, MHA, MLA, not disclosed. block type varies — dense, sparse-MoE. trained context varies — 1024, 131072, 192000, 32768, 4096, 65536, not disclosed. openness varies — open-weights, open-weights-open-data, undisclosed. scale class varies — frontier, large, medium.