Methodology
Every record is extracted from a single technical report and validated against a machine-readable schema (data/schema.json, mirror of schema.md). Every field carries a provenance tag: stated (explicit in the report), derived (computed — computation recorded), inferred (judgement — evidence recorded), unknown (report silent — rendered as not disclosed), or n/a (mechanism makes the field inapplicable). Numbers are never imported from other reports, memory, or the web into a record.
Records
| Slug | Family | Last analysed | Skill version | Source |
|---|---|---|---|---|
ai21-jamba | Jamba | 2026-08-08 | v0.3.1 | report ↗ |
allenai-olmo-2-7b | OLMo 2 | 2026-08-10 | v0.3.5 | report ↗ |
allenai-olmo-3 | OLMo 3 | 2026-08-10 | v0.3.5 | report ↗ |
arcee-ai-trinity-large-400b | Trinity | 2026-08-10 | v0.3.5 | report ↗ |
cisco-antares-1b | Antares | 2026-08-10 | v0.4.0 | report ↗ |
coherelabs-tiny-aya-3-35b | Tiny Aya | 2026-08-10 | v0.3.5 | report ↗ |
deepseek-r1 | DeepSeek-R1 | 2026-08-10 | v0.3.5 | report ↗ |
deepseek-v3 | DeepSeek-V3 | 2026-08-08 | v0.3.0 | report ↗ |
deepseek-v3-2 | DeepSeek-V3.2 | 2026-08-10 | v0.3.5 | report ↗ |
deepseek-v4 | DeepSeek-V4 | 2026-08-10 | v0.4.0 | report ↗ |
google-gemma-3 | Gemma 3 | 2026-08-10 | v0.3.5 | report ↗ |
google-gemma-4 | Gemma 4 | 2026-08-10 | v0.4.0 | report ↗ |
inclusion-ling-2-6-1t | Ling 2.6 | 2026-08-10 | v0.4.0 | report ↗ |
jetbrains-mellum2-thinking-12b-a2-5b | Mellum 2 | 2026-08-10 | v0.3.5 | report ↗ |
meituan-longcat-flash-lite-68-5b-a3b | LongCat-Flash | 2026-08-10 | v0.3.5 | report ↗ |
meta-llama-3 | Llama 3 | 2026-08-10 | v0.3.5 | report ↗ |
meta-llama-3.1 | Llama 3.1 | 2026-08-08 | v0.3.0 | report ↗ |
microsoft-phi-3 | Phi-3 | 2026-08-08 | v0.3.1 | report ↗ |
microsoft-phi-4 | Phi-4 | 2026-08-10 | v0.3.5 | report ↗ |
minimax-m2 | MiniMax-M2 series | 2026-08-10 | v0.4.0 | report ↗ |
minimax-m3-428b | MiniMax M3 | 2026-08-10 | v0.3.5 | report ↗ |
mistral-mixtral-8x7b | Mixtral | 2026-08-08 | v0.3.0 | report ↗ |
moonshot-kimi-k2 | Kimi K2 | 2026-08-10 | v0.3.5 | report ↗ |
moonshot-kimi-k3 | Kimi K3 | 2026-08-08 | v0.3.5 | report ↗ |
moonshot-kimi-linear-48b-a3b | Kimi Linear | 2026-08-10 | v0.3.5 | report ↗ |
nanbeige-4-1-3b | Nanbeige4.1 | 2026-08-10 | v0.3.5 | report ↗ |
nanbeige-4-2-3b | Nanbeige4.2 | 2026-08-10 | v0.4.0 | report ↗ |
nvidia-nemotron-3-nano-30b-a3b | Nemotron 3 Nano | 2026-08-10 | v0.3.5 | report ↗ |
nvidia-nemotron-3-super-120b-a12b | Nemotron 3 Super | 2026-08-10 | v0.3.5 | report ↗ |
nvidia-nemotron-3-ultra-550b-a55b | Nemotron 3 Ultra (Nemotron 3 family) | 2026-08-10 | v0.3.5 | report ↗ |
nxai-xlstm-7b | xLSTM-7B | 2026-08-10 | v0.3.5 | report ↗ |
openai-gpt-2-xl-1-5b | GPT-2 | 2026-08-10 | v0.3.5 | report ↗ |
openai-gpt-oss | gpt-oss | 2026-08-10 | v0.3.5 | report ↗ |
poolside-laguna-xs-2-33b | Laguna | 2026-08-10 | v0.4.0 | report ↗ |
prime-intellect-intellect-3 | INTELLECT-3 | 2026-08-10 | v0.3.5 | report ↗ |
qwen-qwen3-5 | Qwen3.5 | 2026-08-10 | v0.4.0 | report ↗ |
qwen-qwen3-dense | Qwen3 | 2026-08-10 | v0.3.5 | report ↗ |
qwen-qwen3-moe | Qwen3 | 2026-08-10 | v0.3.5 | report ↗ |
qwen-qwen3-next | Qwen3-Coder-Next | 2026-08-10 | v0.4.0 | report ↗ |
upstage-solar-open-2 | Solar Open 2 | 2026-08-10 | v0.4.0 | report ↗ |
weibo-vibethinker-3b | VibeThinker | 2026-08-10 | v0.3.5 | report ↗ |
xiaomi-mimo-v2-flash-309b | MiMo-V2 | 2026-08-10 | v0.3.5 | report ↗ |
zai-glm-4-5 | GLM-4.5 (ARC series) | 2026-08-10 | v0.3.5 | report ↗ |
zai-glm-4-5-air | GLM-4.5 (ARC series) | 2026-08-10 | v0.3.5 | report ↗ |
zai-glm-5 | GLM-5 | 2026-08-10 | v0.3.5 | report ↗ |
zyphra-zaya1-8b | ZAYA1 | 2026-08-10 | v0.3.5 | report ↗ |
Skill version: v0.4.0. Atlas generated 2026-08-11.
Download combined data (atlas.json)
Extraction schema
Extraction Schema
The single authority on what a record contains. data/schema.json (created in Prompt B) is the machine-readable mirror of this file and must stay in sync with it — regenerate whenever this changes, and log the change in changelog.md.
Record form
One file per architecture: data/architectures/<slug>.json. slug = lowercase kebab-case, <org>-<model> (e.g. meta-llama-3.1).
Every leaf field is an object:
{"value": <value|null>, "provenance": "stated|derived|inferred|unknown|n/a", "ref": "§2.1 / Table 3", "note": "computation (derived) or evidence (inferred)"}
unknown→value: null. Renders as not disclosed — never as a blank and never as zero.n/a→value: null,provenance: "n/a", withref. Distinct from unknown: the field does not apply to this mechanism. (Numeric-typed leaves accept null, so n/a is always expressible.)refrequired forstatedandn/a;noterequired forderivedandinferred.- Arrays and objects carry the same tagging on every leaf; container-level tags are not allowed.
- Precision requirement: two independent analyses of the same report must produce identical values. When in doubt, the tag is
unknown.
1. Record meta (pipeline bookkeeping, not report content)
| Field | Type | Provenance | Definition |
|---|---|---|---|
slug | string | stated (pipeline) | <org>-<model> kebab-case. |
analysed_date | string ISO | stated (pipeline) | Day the analysis was completed. |
skill_version | string | stated (pipeline) | changelog.md version at analysis time (e.g. v0.2.0). |
source.url | string | stated (pipeline) | Report URL. |
source.fetched_date | string ISO | stated (pipeline) | Retrieval date. |
source.stored_path | string | stated (pipeline) | Path under data/sources/<slug>/. |
2. Identity
| Field | Type | Provenance rule | Definition |
|---|---|---|---|
identity.family | string | stated | Model family name as the report calls it. |
identity.variants | array[string] | stated | Disclosed sizes/names, e.g. ["8B","70B","405B"]. |
identity.org | string | stated | Organisation. |
identity.release_date | ISO date or null | stated if the report prints a date (ref = date line/abstract); else inferred from arXiv v1 submission metadata (note = arXiv ID + API); never from announcement/marketing dates | Report publication date, not model availability date. Unknown only if neither applies. |
identity.report_url | string | stated (pipeline) | Canonical report URL. |
identity.license | string | stated if the report names it; else unknown | License name exactly as given. |
identity.open_weights | boolean | stated if the report says weights are released; else unknown | Strictly report-based. World knowledge does not fill this field; it may appear in prose labelled as secondary. |
3. Scale (reference variant = the report's flagship)
Per-variant configs go in scale.variants[]; the top-level scale fields always describe the reference variant — the variant the report leads with as its flagship (title/abstract framing decides; e.g. phi-3-mini leads "Phi-3: ... Locally on Your Phone", so mini is the reference even though medium is larger). If the report leads with no single variant, the reference is the largest disclosed. If the report covers one size only, variants is [].
| Field | Type | Provenance rule | Definition |
|---|---|---|---|
scale.total_params | number (B) | stated (config table) or derived (sum of disclosed components, note the computation) | Reference variant. |
scale.active_params_per_token | number (B) | derived for dense (= total, note it); MoE: stated or derived from routing (note: attention + shared experts + per-token routed experts + embeddings) | Parameters touched per token. |
scale.layers | int | stated | Transformer layer count. |
scale.hidden_dim | int | stated | Model (embedding) dimension. |
scale.ffn_inner_dim | int | stated, or derived (note computation) | FFN hidden dimension. |
scale.ffn_ratio | number | derived: ffn_inner_dim / hidden_dim (note the computation). MoE: per-expert ratio with note. | FFN expansion ratio. |
scale.attention_heads_q | int | stated | Query heads. |
scale.attention_heads_kv | int | stated; n/a for MLA (ref = MLA section) | KV heads (GQA/MQA). For MQA = 1. |
scale.head_dim | int | derived: hidden_dim / attention_heads_q (note computation); stated if the report gives it directly; MLA: the latent dimension, stated | Head dimension. |
scale.vocab_size | int | stated (tokenizer/config section) | Vocabulary size. Single source of truth — tokenizer group references this, no duplicate. |
scale.embedding_tied | boolean | stated only; else unknown | Input/output embedding tying. Config-file knowledge is secondary and does not fill this field. |
scale.variants[] | array | per-element stated | {name, total_params, layers, hidden_dim, ffn_inner_dim, attention_heads_q, attention_heads_kv, context_length} for each disclosed size. |
4. Core block design
| Field | Type | Provenance rule | Definition |
|---|---|---|---|
block.block_type | dense \ | sparse-MoE \ | hybrid |
block.moe.expert_count | int | stated; n/a if dense | Total experts (excluding shared). |
block.moe.experts_per_token | int | stated | Routed experts activated per token (top-k). |
block.moe.shared_experts | int | stated; 0 if none | Always-active experts. |
block.moe.routing | string | stated | Router function (e.g. sigmoid gating, softmax top-k). |
block.moe.load_balancing | string | stated; unknown if silent | Aux loss / bias / none / other. |
block.moe.expert_granularity | string | stated or inferred (note evidence) | e.g. fine-grained, grouped. |
block.attention_variant | MHA \ | MQA \ | GQA \ |
block.attention_layer_pattern | string | stated | Per-layer pattern for hybrids (e.g. "layers 1–3 full, 4–61 sliding window"); "uniform" otherwise. |
block.depth_mixing | sequential-residual \ | attention-residuals \ | hyper-connections |
block.position_encoding.method | RoPE \ | ALiBi \ | NoPE \ |
block.position_encoding.rope_base | number | stated if the report gives it; else unknown | RoPE base frequency. |
block.position_encoding.partial_rope | boolean | stated; else per absence rule | RoPE applied to a subset of dims. |
block.position_encoding.extension | `{method: YaRN\ | NTK\ | PI\ |
block.normalization.type | string | stated | e.g. RMSNorm, LayerNorm, QK-norm variants. |
block.normalization.placement | pre \ | post \ | mixed |
block.normalization.qk_norm | boolean | stated; else per absence rule | QK-normalisation. |
block.activation | string | stated | e.g. SwiGLU, GELU. |
block.stability.attention_sinks | boolean | stated; else per absence rule | Designed sink tokens only — emergent sinks go in prose. |
block.stability.softcapping | number or boolean | stated; else per absence rule | Logit softcapping value. |
block.stability.other | array[string] | stated or inferred (note evidence) | Other stability tricks. |
5. Context
| Field | Type | Provenance rule | Definition |
|---|---|---|---|
context.trained_length | int (tokens) | stated | Trained context length. |
context.deployed_length | int (tokens) | stated | Length served/extended to. |
context.extension_method | string | stated if disclosed; "none" if the report says training was at the deployed length; else unknown | Extension method (may mirror position_encoding.extension). |
6. Tokenizer
| Field | Type | Provenance rule | Definition |
|---|---|---|---|
tokenizer.algorithm | string | stated | BPE / SentencePiece / tiktoken-style / other. |
tokenizer.notes | string | stated | Notable design choices. Vocab size lives in scale.vocab_size (no duplicate). |
7. Training
| Field | Type | Provenance rule | Definition |
|---|---|---|---|
training.tokens | number (T) | stated; else unknown | Pretraining tokens. |
training.data_composition | string | stated (paraphrase); unknown if silent | Disclosed mixture — absence is a finding. |
training.curriculum | string | stated; "none disclosed" if silent | Staging/annealing. |
training.optimizer | string | stated | |
training.lr_schedule | string | stated | |
training.batch_schedule | string | stated | |
training.precision | string | stated | BF16/FP8/other. |
training.parallelism | string | stated | TP/PP/EP/CP strategy. |
training.hardware | string | stated | |
training.compute | string | stated; else unknown | Disclosed FLOPs or GPU-hours. |
8. Post-training
| Field | Type | Provenance rule | Definition |
|---|---|---|---|
post_training.sft | boolean + note | stated or inferred (note evidence) | Supervised fine-tuning. |
post_training.preference_optimization | RLHF \ | DPO \ | GRPO \ |
post_training.reasoning_training | string | stated or inferred (note evidence) | CoT SFT, RLVR, test-time compute, etc. |
post_training.distillation | string | stated or inferred (note evidence) | Distilled from which model; "none" if report says trained from scratch. Teacher-generated training data (e.g. a large model generating SFT data for smaller siblings) is not distillation — note it in prose, not here. |
9. Modality
| Field | Type | Provenance rule | Definition |
|---|---|---|---|
modality.type | text-only \ | multimodal | stated if the report describes modalities; else inferred (note evidence, e.g. evaluation tasks imply vision) |
modality.attachment | native \ | adapter | stated; n/a if text-only |
10. Inference efficiency
| Field | Type | Provenance rule | Definition |
|---|---|---|---|
efficiency.kv_cache | string | derived from attention_variant (note the consequence, e.g. "GQA: KV cache ∝ 8 heads"; "MLA: low-rank latent KV"); stated if the report discusses it | KV-cache design consequences. |
efficiency.quantization | string | stated; "none disclosed" if silent | Shipped quantisation formats. |
efficiency.speculative_dedup | string | stated; "none disclosed" if silent | Speculative decoding, MTP, self-speculation. |
efficiency.serving | string | stated; "none disclosed" if silent | Disclosed serving optimisations. |
11. Evaluation
| Field | Type | Provenance rule | Definition |
|---|---|---|---|
evaluation.benchmarks[] | array of {name, value, ref} | stated only | name = benchmark as the report names it; value = the report's own number, digits as printed. Never import numbers from other sources, never compare across reports here. |
12. Author-claimed contributions
| Field | Type | Provenance rule | Definition |
|---|---|---|---|
contributions.claimed[] | array of {text, ref} | stated (paraphrased, attributed) | What the report says is novel. |
contributions.assessment | string | inferred (note evidence: comparison against atlas entries) | Your judgement of actual novelty relative to the atlas — kept separate from claimed. |
13. Taxonomy (one value per axis; each with provenance)
| Field | Values | Provenance rule |
|---|---|---|
taxonomy.compute_structure | dense / sparse-MoE / hybrid | derived from block per-layer pattern (majority rule, taxonomy.md) |
taxonomy.sequence_mixing | full-attention / efficient-attention / SSM / hybrid | derived from block.attention_layer_pattern |
taxonomy.modality | text-only / multimodal | derived from modality.type |
taxonomy.openness | open-weights-open-data / open-weights / open-data / closed / undisclosed | derived from identity.open_weights + training.data_composition disclosure |
taxonomy.scale_class | frontier / large / medium / small | derived from scale.total_params (reference variant) |
taxonomy.design_intent | frontier-generalist / reasoning-specialised / on-device / long-context / domain-specific | inferred (note evidence: report framing in abstract/intro) |
14. Prose (authored narrative, per record)
Authored analysis for the model page — not report facts. Every leaf is tagged inferred with note "authored by atlas analyst", and the prose text itself honours non-negotiable 2 (keeps stated / inferred distinctions explicit inside the narrative). Authored at ingestion time; influence_out is the one field written retroactively (when a descendant arrives).
| Field | Type | Definition |
|---|---|---|
prose.design_overview | leaf (string) | 150–300 word design overview: what the block looks like and why it matters. |
prose.lineage_in[] | array of {mechanism, origin} | Mechanisms borrowed from prior architectures — only when report-stated or unambiguous (methodology §6). mechanism = named mechanism ("MLA from DeepSeek-V2"); origin = the prior model. |
prose.influence_out[] | array of {model_slug, mechanism} | Descendants that borrow from this record; maintained when a descendant is ingested (never authored in advance). |
prose.notable_omissions | leaf (value = array of strings) | What this report omits that peers disclose — the absence findings. |
Taxonomy
Taxonomy
Classification is multi-axis: a model gets exactly one value per axis, never one bucket. For every axis: possible values, decision rule, anchor example. If a report genuinely fits no value, that is a signal to evolve the taxonomy (changelog procedure + backfill), not to force a fit.
Format contract (machine-parsed): build.py parses this file — the axis key comes from the ## Axis N — name heading, the values/rules/anchors from the first table in each section. Keep this structure; extend tables, don't restyle them.
Axis 1 — compute_structure
What the parameter footprint looks like per token.
| Value | Decision rule | Anchor |
|---|---|---|
dense | Every transformer layer's FFN is fully active per token; no routing anywhere. | Llama 3 (report describes no MoE; absence rule → inferred) |
sparse-MoE | MoE layers dominate (≥80% of transformer layers are routed FFNs). | Mixtral 8x7B |
hybrid | Neither pattern reaches 80% of layers (e.g. a substantial mix of dense and MoE layers). | None in atlas yet — first qualifying report sets the anchor |
Rule: count routed-FFN vs dense-FFN layers across the whole model; majority with ≥80% share wins, otherwise hybrid. The per-layer record lives in block.block_type / block.attention_layer_pattern.
Axis 2 — sequence_mixing
How the model mixes information across sequence positions.
| Value | Decision rule | Anchor |
|---|---|---|
full-attention | All layers use unrestricted attention over the whole context. | Llama 3 |
efficient-attention | Attention with restricted/computed patterns (sliding window, linear attention, sparse patterns) — but still attention. | Mistral (sliding window layers) |
SSM | State-space layers (Mamba-style) do the mixing; no attention in the block. | Mamba |
hybrid | Mixed attention + SSM/linear layers (or attention variants in a deliberate per-layer pattern). | Jamba (1 attention : 7 Mamba layers) |
Rule: same ≥80% majority rule over layers; deliberate alternation that falls below the threshold is hybrid (record the pattern in block.attention_layer_pattern).
Axis 3 — modality
| Value | Decision rule | Anchor |
|---|---|---|
text-only | Report describes no non-text input or output. | Llama 3.1 |
multimodal | Report describes vision/audio/other modalities in or out. | GPT-4V (if its report is ever in the atlas — else first multimodal report ingested) |
Axis 4 — openness
What the organisation actually releases, per the report.
| Value | Decision rule | Anchor |
|---|---|---|
open-weights-open-data | Weights and training data released. | OLMo |
open-weights | Weights released; data not released. | Llama 3 (report states weights release; data composition disclosed but not released) |
open-data | Data released; weights not. | (rare — first qualifying report sets the anchor) |
closed | Report says the model is not released. | GPT-4 (if its report is ever in the atlas) |
undisclosed | Report is silent on release. Absence is itself the classification. | First qualifying report sets the anchor |
Note: undisclosed exists so report silence never forces a guess. World knowledge does not override a silent report (non-negotiable 3).
Axis 5 — scale_class
By total parameters of the reference variant (largest disclosed).
| Value | Rule (total params) | Anchor |
|---|---|---|
frontier | ≥ 100B | Llama 3.1 405B |
large | 10B ≤ p < 100B | Llama 3.1 70B |
medium | 1B ≤ p < 10B | Llama 3.1 8B |
small | < 1B | Qwen2-0.5B (if ingested) |
Axis 6 — design_intent
What the report frames the model for. Inferred from the report's own framing (abstract, intro, evaluation emphasis) — never from marketing outside the report.
| Value | Decision rule | Anchor |
|---|---|---|
frontier-generalist | Framed as a general-purpose foundation model; broad benchmark coverage. | Llama 3.1 |
reasoning-specialised | Framed around reasoning/thinking (CoT, RLVR, test-time compute). | o1 (if its report is ever in the atlas) |
on-device | Framed for edge/deployment efficiency constraints. | Phi-3-mini |
long-context | Framed primarily around context length. | Gemini 1.5 (if in atlas) |
domain-specific | Framed for a specific domain (code, biology, law…). | DeepSeek-Coder (if in atlas) |
Rule: if several intents appear, take the one the report leads with; note the others in contributions.assessment or prose. If the report leads with none of these, evolve the axis.
Comparison methodology
Methodology — comparison and normalisation conventions
Rules that keep cross-model comparisons honest.
1. MoE vs dense — always both totals
Never compare total parameters of an MoE model against a dense model without also showing active parameters. Always present both (scale.total_params and scale.active_params_per_token); comparison views show the pair side by side.
2. Unit normalisation
Normalise before comparing:
- Token counts in T (trillions).
- Parameter counts in B (billions).
- Context lengths in tokens.
- Dates as ISO-8601 (YYYY-MM-DD).
Report values are stored as printed (verbatim digits); derived values are rounded to one decimal place. The stored value is never rounded; rounding happens only in rendered views.
3. Compare designs, not marketing
A comparison row must be a schema field, so every cell has provenance. Prose comparison is allowed only after the full schema is extracted (non-negotiable 1). Cells show provenance on hover; unknown renders as "not disclosed" — never as blank, never as zero.
4. Absence of disclosure is information
When a report omits a field that peers disclose, the comparison shows "not disclosed" and the omission is noted in prose ("the report does not disclose its data mixture"). Silence is a finding, not a gap to fill.
5. Benchmarks are never compared head-to-head across reports
Different harnesses, prompts, and contamination levels make cross-report numbers incomparable. The atlas records which benchmarks each report chose to emphasise and that report's own numbers only. At most, note the choice pattern (e.g. "this report leads with reasoning benchmarks, unlike X"). Same-report, same-harness comparisons (e.g. variant sizes within one report) are fine.
6. Lineage
Record which prior architectures a design borrows from — with the specific borrowed mechanism named (e.g. "MLA from DeepSeek-V2") — only when the report states it or the mechanism is unambiguous (mechanism + origin both identifiable). Lineage in → the new record's prose. Influence out → update older records' "influence" sections when a descendant arrives.
7. Reference variant
One record per family. Scale fields describe the largest disclosed variant; per-variant configs live in scale.variants[]. Comparisons use the reference variant unless a variant-level comparison is explicitly requested. If a report covers a single size, variants is empty.
8. Backfill discipline
When the schema or taxonomy evolves, every existing record is re-read from its stored source and the new field filled (unknown where silent) in the same change — no field may exist for some models only because of ingestion order. The changelog entry for the change lists every backfilled record.
9. Secondary sources
Secondary sources (other reports, model cards, config files, web) are labelled as such in prose and never fill schema fields (non-negotiable 3). They may be cited in contributions.assessment or prose as context, explicitly marked.
10. Closest relatives
"Closest relatives in the atlas" is computed mechanically from shared schema values: count matching non-unknown, non-n/a values across the fixed comparable field set — block.block_type, block.attention_variant, block.position_encoding.method, block.normalization.type, block.activation, modality.type, taxonomy.openness, taxonomy.scale_class, taxonomy.sequence_mixing. Tie-break by absolute difference in context.trained_length, then by scale.total_params. The field set and rule are recorded here so the computation is reproducible and versioned with the skill.
11. Numbers
All numbers on a model page come from that model's record. The report's own digits are stored as printed; derived values show their computation in the provenance note.
Glossary
Glossary — terminology map
Reports name the same mechanism differently. Add a new alias here BEFORE filling the schema. Format: alias → canonical schema term (path), note if needed. Canonical terms are defined in schema.md.
Attention
MQA,multi-query attention,shared KV heads→block.attention_variant: MQA(one KV head shared by all query heads)GQA,grouped-query attention,grouped KV heads→block.attention_variant: GQA(a few KV heads shared per query group)MLA,multi-head latent attention,low-rank KV compression,latent KV cache→block.attention_variant: MLA(KV compressed into a low-rank latent;attention_heads_kvisn/a)sliding window attention,SWA,local attention→block.attention_variant: sliding-windowblocksparse attention,block-sparse attention→block.attention_layer_pattern(efficient-attention pattern: per-head sparsity patterns over the KV cache; phi-3-small alternates blocksparse and dense attention layers)flash attention→ not a variant; an implementation detail (prose only)KV cache compression,KV sharing→efficiency.kv_cacheKDA,Kimi Delta Attention,delta-rule attention,delta rule recurrence→block.attention_variant: linear/state-space(recurrent delta-rule linear attention; chunkwise-parallel; in hybrid patterns record the mix inattention_layer_pattern)Gated MLA,gated multi-head latent attention→block.attention_variant: MLA(MLA with an input-dependent full-rank output gate)Attention Residuals,AttnRes,attention over layers,layer-wise attention→block.depth_mixing: attention-residuals(per-layer learned attention over prior layer outputs; Block variant: attention over summed block representations)
Mixture of experts
MoE,mixture-of-experts,sparse expert model,routed FFN,SMoE,sparse mixture of experts→block.block_type: sparse-MoE/moefieldstop-k routing,expert routing,router,router network,gating network,gating vector→block.moe.routingsparse parameter count→scale.total_paramsactive parameter count→scale.active_params_per_tokenauxiliary loss,aux loss,load balancing loss,balancing loss→block.moe.load_balancingshared expert,always-on expert→block.moe.shared_expertsactivated parameters,active parameters,per-token parameters→scale.active_params_per_tokenfine-grained experts,expert splitting→block.moe.expert_granularityLatentMoE,latent experts,routed experts in a compact latent space→block.moe.expert_granularity(routed experts operate at reduced width; shared experts keep full width)Quantile Balancing,QB,quantile-based bias update→block.moe.load_balancing(auxiliary-loss-free bias set from the router-score margin quantile matching target load)auxiliary-loss-free load balancing,bias-based balancing,bias update speed,sequence-wise balance loss,aux loss→block.moe.load_balancingnode-limited routing,device-limited routing,restricted routing→block.moe.routing(routing constraint limiting tokens to ≤M nodes)redundant expert,expert duplication,dynamic redundancy→efficiency.servingSparseMixer,sparse backpropagation→block.moe.routing(MoE router training via sparse backpropagation [LGC23, LDL+23])Jamba block→block.block_type: hybrid(report's name for the repeated unit of attention/Mamba layers + MLP/MoE)attention-to-Mamba ratio,a:m ratio,ratio of attention-to-Mamba layers→block.attention_layer_pattern
Positional encoding
RoPE,rotary,rotary position embeddings,rotary embeddings→block.position_encoding.method: RoPE;rope_basefor the base frequencyALiBi,attention with linear biases→block.position_encoding.method: ALiBiNoPE,no positional encoding→block.position_encoding.method: NoPEno positional information,no explicit positional encoding,without explicit positional information→block.position_encoding.method: NoPE(report states positional embeddings/RoPE are not necessary)YaRN,YaRN-NTK,YaRN scaling→block.position_encoding.extension.method: YaRNNTK-aware scaling,NTK scaling,dynamic NTK→block.position_encoding.extension.method: NTKlinear scaling,position interpolation,PI→block.position_encoding.extension.method: PILongRope,long-rope,LongRoPE,long rope→block.position_encoding.extension.method: other(RoPE-frequency-rescaling context extension; 4K to 128K)context window,context length,sequence length,max sequence length→context.*
Normalisation and activation
RMSNorm→block.normalization.type: RMSNormLayerNorm,LN→block.normalization.type: LayerNormpre-norm,pre-LN,PreNorm→block.normalization.placement: prepost-norm,post-LN→block.normalization.placement: postQK-norm,query-key normalisation,QK layernorm→block.normalization.qk_normSwiGLU,GLU,SiLU-gated linear unit,gated activation→block.activation: SwiGLUSiTU-GLU,sigmoid tanh unit GLU,tanh-softcapped GLU→block.activation(record exactly as stated; bounded SwiGLU variant with β·tanh soft-caps)GEGLU→block.activation: GEGLU(gated linear unit with GELU nonlinearity)softcapping,logit capping,logit softcap,attention logit capping→block.stability.softcappingattention sink,sink token,streaming attention sink→block.stability.attention_sinks(designed sinks only)
Training and precision
bfloat16,BF16,bf16 mixed precision→training.precision: BF16FP8,FP8 mixed precision,8-bit training→training.precision: FP8AdamW,Adam→training.optimizer(record exactly as stated)annealing,data annealing,cooldown,stage 2→training.curriculumtokens seen,training tokens,corpus size→training.tokensexpert parallelism,EP→training.parallelismpipeline parallelism,PP→training.parallelismtensor parallelism,TP→training.parallelismcontext parallelism,CP→training.parallelismsequence parallelism→training.parallelismFSDP,fully sharded data parallel→training.parallelismDualPipe,bidirectional pipeline scheduling,computation-communication overlap→training.parallelismE4M3,E5M2,E5M6,fine-grained quantization,tile-wise quantization,block-wise quantization,online quantization→training.precisionFIM,fill-in-the-middle,Prefix-Suffix-Middle,PSM,document packing→training.data_compositionMaximal Update Parametrization,muP,µP→ training methodology (zero-shot hyperparameter transfer from a small proxy model; prose — no dedicated field)data optimal regime,data-optimal regime→training.data_composition(data-quality-focused training regime at a fixed model scale, as opposed to compute-optimal scaling)
Post-training
RLHF,PPO,reinforcement learning from human feedback→post_training.preference_optimization: RLHFDPO,direct preference optimisation→post_training.preference_optimization: DPOGRPO,group relative policy optimisation→post_training.preference_optimization: GRPORLAIF,RL from AI feedback→post_training.preference_optimization: other(note)chain-of-thought,CoT,reasoning traces,thinking tokens,thinking mode→post_training.reasoning_trainingtest-time compute,self-consistency,best-of-n→post_training.reasoning_trainingRLVR,RL with verifiable rewards→post_training.reasoning_trainingdistillation,KD,knowledge distillation,teacher model→post_training.distillationMOPD,multi-teacher on-policy distillation,on-policy distillation,OPD→post_training.distillationgenerative reward model,GRM,agentic reward model→post_training.preference_optimization(note the reward structure; e.g. tournament-style group reward with binary comparisons)rejection sampling(used to curate SFT data) →post_training.sft
Inference
speculative decoding,speculative sampling,self-speculative,draft model→efficiency.speculative_dedupEAGLE,EAGLE-3,MTP draft,draft from MTP layer→efficiency.speculative_dedupMTP,multi-token prediction,multi-token forecasting→efficiency.speculative_dedup(and training objective → prose)quantisation,quantization,INT4,INT8,W8A8,AWQ,GPTQ,MXFP4,MXFP8,MXFP4 QAT→efficiency.quantization(record shipped formats as stated)KV cache,key-value cache→efficiency.kv_cacheIBGDA,GPUDirect Async→efficiency.servingmixed context window approach→context.extension_method(used with long-rope for the phi-3.5 series' 4K to 128K mid-training)
Attention (2026-08-10 batch — gallery ingestion)
DSA,DeepSeek Sparse Attention,lightning indexer,fine-grained token selection,keep routing,keep sampling mask→block.attention_variant/attention_layer_pattern(content-based top-k sparse attention over the KV cache; DSA appears as the sparse stage inside CSA)CSA,compressed sparse attention,HCA,heavy compression attention→block.attention_layer_pattern(KV compressed every m tokens then sparsely/densely attended; DeepSeek V4)CCA,CCGQA,compressed convolutional attention→block.attention_variant: hybrid(attention in a compressed latent space with short conv + grouped head-wise conv preconditioners and GQA-style KV sharing; Zyphra ZAYA1)banded-window attention,off-by-one attention,learned softmax-denominator bias→block.attention_layer_pattern/block.stability.other(GPT-OSS: window/dense alternation and per-head softmax-denominator bias)mHC,manifold-constrained hyper-connections,Hyper-Connections,connection matrix→block.depth_mixing: hyper-connections(learned connection matrix over the residual stream; doubly-stochastic constraint in DeepSeek V4)sliding-window,SWA,local attention→block.attention_variant: sliding-window(alias for the gallery batch)attention sink bias→block.stability.attention_sinks(learnable sink in the softmax denominator; gpt-oss lineage)
Mixture of experts (2026-08-10 batch)
SMEBU,sequence auxiliary loss→block.moe.load_balancing(sequence-level expert-balancing loss; Arcee Trinity)zero-experts,dynamic activation,sparsity scaling→block.moe(experts that can be skipped per token, making active params input-dependent; LongCat-Flash-Lite)n-gram embedding,over-encoding,N-gram Cache,embedding amplification→scale.total_params/ prose (hashed n-gram lookup tables as a parameter block outside the FFN; LongCat-Flash-Lite)LatentMoE→block.moe.expert_granularity(routed experts in a compact latent space; Nemotron 3 Super, Kimi K3's Stable LatentMoE)Sqrt(Softplus) affinity,hash routing,expert bias,post-top-k score normalization→block.moe.routing(router variants as stated)GenRM,group relative length control→post_training.preference_optimization(reward-model/RM structures in RLHF; Nemotron 3)
Post-training (2026-08-10 batch)
OlmoRL,active sampling,token-level loss,clip-higher,truncated importance sampling→post_training.reasoning_training(GRPO + DAPO/Dr.GRPO-style RLVR algorithm; OLMo 3)CISPO,clipped importance sampling policy optimization→post_training.preference_optimization: other(MiniMax M2, Laguna)MGPO,max-entropy guided policy optimization→post_training.preference_optimization: other(GRPO variant with max-entropy prompt weighting; VibeThinker)pivotal token search,PTS→post_training.preference_optimization(DPO-pair construction from pivotal tokens; Phi-4)midtraining,mid-training,continued pretraining stage→training.curriculum(stage between pretraining and post-training)MOPD→post_training.distillation(already listed; also used by Upstage Solar Open 2 and cited by Kimi K3)Long2Short RL,length-reward redistribution,offline self-distillation→post_training.reasoning_training(VibeThinker)Looped Transformer,layer reuse,two-pass stack reuse→block.depth_mixing/ prose (same stack processed twice; Nanbeige 4.2)
Training (2026-08-10 batch)
Muon,MuonClip,QK-Clip,hybrid Newton-Schulz,WSD→training.optimizer/block.normalization.qk_norm(Muon optimizer family; MuonClip is per-head weight clipping — NOT QK-norm; WSD = warmup-stable-decay LR)checkpoint souping,model souping→training.curriculummicro-annealing,microanneal→training.curriculum(cheap data-source ablations via short anneal runs)
Misc (2026-08-10 batch)
o200k_harmony→scale.vocab_size/tokenizer.algorithm(GPT-OSS tokenizer name)cl100k-derived→tokenizer.algorithm(OLMo 2/3 tokenizer lineage)MM Claw→ prose only (internal multimodal benchmark; not a modality pipeline — MiniMax M2)
2026-08-10 batch 2 — found-report ingestion
Lightning Attention→block.attention_variant: linear/state-space(transpose-free linear attention; TransNormerLLM lineage; Ling 2.6)TransMLA→block.attention_variant: MLA(transplanting GQA checkpoints to MLA via QK-norm absorption + partial-RoPE decoupling; Ling 2.6)GDN,gated delta net,Hybrid-Attention MoE→block.attention_variant: linear/state-space(delta-rule gated linear attention inside a hybrid MoE; Qwen3.5)pp-RoPE→block.position_encoding.partial_rope(partial-rotary RoPE, p=0.25 on global layers; Gemma 4)keys-as-values reuse,values=keys→block.stability.other(global-attention layer reuse of keys as values; Gemma 4)KV cache sharing,cross-layer KV cache→efficiency.kv_cache(shared KV tensors across layer groups — distinct from GQA; Gemma 4 E-series)MTP drafter→efficiency.speculative_dedup(multi-token prediction used as draft model)80A3→scale.total_params/scale.active_params_per_token(report shorthand for 80B total / 3B active; Qwen3-Coder-Next)einsums→scale.total_params(non-embedding matmul parameter accounting; Gemma 4 Table 1)Looped Transformer,layer reuse,two-pass stack reuse→block.depth_mixing/ prose (same stack processed twice; Nanbeige 4.2)
tied embeddings,weight tying,shared input-output embeddings,tied word embeddings→scale.embedding_tiedopen weights,open-source weights,weights released,public checkpoints→identity.open_weightstokenizer,vocabulary,vocab→scale.vocab_size,tokenizer.algorithmtiktoken,byte-level BPE,sentencepiece,BBPE→tokenizer.algorithm(record as stated)modality,vision-language,VLM,multi-modal→modality.*adapter,connector,projector,vision encoder,ViT→modality.attachment
Changelog
Changelog
Every schema/taxonomy change, dated, with the trigger. Current version: v0.4.0.
2026-08-08 — v0.3.5 — depth-mixing field (trigger: Kimi K3 ingestion — Attention Residuals)
block.depth_mixingadded (schema.md §4, schema.jsonleafEnumDepthMixing):sequential-residual|attention-residuals. Triggered by Kimi K3's Attention Residuals (AttnRes) — per-layer learned attention over prior layer outputs — which no existing field could represent; without a field it would have lived only in prose.- Backfilled all 5 existing records:
depth_mixing→sequential-residual(inferred, absence rule — their stored reports describe the block in detail and never mention attention over prior layers). - Kimi K3 record:
depth_mixing→attention-residuals(stated, §2.2; Block variant: 8 blocks × 12 layers + partial final block).
2026-08-08 — v0.3.4 — ingest pipeline: arXiv PDF fallback (trigger: report 2607.24653 has no HTML version)
ingest.py_arxiv_source_url: arXiv abs URLs now resolve HTML → ar5iv → PDF via export.arxiv.org/pdf/<id> instead of erroring when no HTML exists. PDFs extract via pymupdf (installed) or pypdf.- SKILL.md pitfall updated: PDF is last-resort extraction; flag PDF layout artifacts (math/tables garble) in prose, and rely on the step-7 grep-verify against source.txt.
- No schema/taxonomy fields changed — no backfill needed.
2026-08-08 — v0.3.3 — release-date convention (trigger: first seed batch — all five reports are undated in extracted text; Llama record initially carried an untraceable stated date)
identity.release_daterule extended (schema.md §2, SKILL.md conventions): stated when the report prints a date; otherwise inferred from arXiv v1 submission metadata (note records the arXiv ID + API as evidence); announcement/marketing dates are never used. Strict "report text only" dating left the timeline strip empty (zero dated models) and forced world-knowledge dates to masquerade as stated.- Backfilled all 5 records: release_date → inferred with arXiv-ID evidence (Llama 3.1 corrected from an untraceable stated 2024-07-23 — the announcement date — to the arXiv v1 date 2024-07-31).
2026-08-08 — v0.3.2 — reference-variant rule refined (trigger: Phi-3 ingestion)
- schema.md §3 + SKILL.md conventions: reference variant is now the variant the report leads with as its flagship (title/abstract framing decides; example: phi-3-mini), falling back to largest disclosed when the report leads with none. The v0.2.0 rule ("largest disclosed") misclassifies reports whose flagship is not the largest (Phi-3: mini 3.8B leads; small 7B / medium 14B are variants).
- No backfill needed — the Phi-3 record (analysed under v0.3.1) already uses the correct reference; its skill_version field records the version it was extracted under.
2026-08-08 — v0.3.1 — schema.md/data.schema.json alignment (trigger: first ingestion batch — Llama 3.1, DeepSeek-V3, Mixtral)
Clarifications so the human schema and the machine schema can never disagree:
block.block_typevalues corrected todense | sparse-MoE | hybridin schema.md (wasdense | moe | hybrid; the machine enum and taxonomy.md already usedsparse-MoE). Two independent extractors hit the mismatch.n/aleaf shape clarified in the record form:value: null+provenance: "n/a"+ref(numeric-typed leaves accept null, so n/a is always expressible).prose.notable_omissionsclarified as a leaf whose value is an array of strings.
2026-08-08 — v0.3.0 — prose group (trigger: Prompt B, site bootstrap)
- schema.md §14 added:
prosegroup (design_overview, lineage_in[], influence_out[], notable_omissions) — authored narrative lives inside the record so one JSON file per architecture remains the single source of truth for both specs and prose. All leaves tagged inferred ("authored by atlas analyst"). - taxonomy.md: added machine-parse format contract (build.py reads axis names from
## Axis N — nameheadings and the first table of each section). - No backfill needed — no records exist yet.
2026-08-08 — v0.2.0 — dry-run validation fixes (trigger: Llama 3.1 validation fill, dry-runs/llama-3.1.md)
Fixes applied after dry-running the schema on a well-known architecture. No backfill needed — no records exist yet.
- Absence rule added (schema.md §4 + SKILL.md): for a mechanism the report doesn't mention — if the report's architecture description is detailed enough that the mechanism would appear were it used, absence →
inferred"not used" with that evidence; otherwise →unknown. Triggered by qk_norm, softcapping, partial_rope, embedding_tied on Llama 3.1. - Designed vs emergent distinction (schema.md
block.stability.attention_sinks, SKILL.md): only explicitly designed mechanisms fill schema fields; emergent behaviour (e.g. emergent attention sinks) goes in prose. Llama 3.1 certainly has emergent sinks; the report designs none. - Reference-variant convention (schema.md §3, SKILL.md conventions, methodology.md §7): families with multiple disclosed sizes get one record; scale fields describe the largest variant; per-variant configs in
scale.variants[]. Triggered by Llama 3.1's 8B/70B/405B. n/aprovenance tag (schema.md record form, SKILL.md): for fields inapplicable to a mechanism (e.g.attention_heads_kvunder MLA), distinct fromunknown(not disclosed).undisclosedopenness value (taxonomy.md Axis 4): report silence on release is itself a classification; never force a guess from world knowledge.- Release date convention (schema.md §2, SKILL.md): report publication date, not model availability date; ref required.
- Active params for dense models (schema.md §3, SKILL.md conventions):
active_params_per_token=total_params, tagged derived. - Distillation nuance (schema.md §8): teacher-generated training data is not distillation; noted in prose, not the field. Triggered by Llama 3.1 using 405B to generate SFT data for 8B/70B.
- Benchmark value rule (schema.md §11): values stored digits-as-printed; no rounding of report numbers.
- Closest-relatives rule (methodology.md §10): fixed comparable field set + tie-breaks, so the computation is reproducible.
2026-08-08 — v0.1.0 — initial schema and taxonomy (trigger: Prompt A, README.md)
- schema.md v0.1: record form (per-leaf provenance objects), groups 1–13 as in the prompt.
- taxonomy.md v0.1: six axes with values, decision rules, anchors.
- methodology.md v0.1: comparison conventions.
- glossary.md v0.1: alias table (~60 entries).
2026-08-10 — v0.4.0 — depth_mixing gains hyper-connections (trigger: DeepSeek V4 ingestion)
- DeepSeek V4's mHC (Manifold-Constrained Hyper-Connections) is a learned residual-matrix mechanism over the residual stream — neither plain sequential residuals nor Attention Residuals. Extended
block.depth_mixingenum tohyper-connections. - No backfill needed: no existing record uses the mechanism; the K3 Attention Residuals value is unaffected.
- data/schema.json leafEnumDepthMixing updated in sync.