nvidia/nemo-automodel-model-onboarding
Guide for onboarding new model architectures into NeMo AutoModel, including architecture discovery, implementation patterns, registration, and validation.
npx skills add https://github.com/NVIDIA/skills --skill nemo-automodel-model-onboarding
This skill guides implementation of new model architectures in NeMo AutoModel. Follow the five phases in order.
<!-- NVSkills signature refresh requested after PR #2998 (2026-07-31). -->
When answering an onboarding question, keep the response in this order:
config.json.components/models/<name>/.For conceptual onboarding questions, answer from this skill without opening the
pattern files unless the user asks you to edit code. Mention pattern filenames
as references, then give the direct checklist.
Use direct action verbs: classify the model, name the files, map the weights,
register the class, and add tests. Do not discuss distributed strategy,
launcher configuration, or general recipe authoring unless the user explicitly
connects it to onboarding a new architecture.
Use these compact answer patterns for common questions:
architectures contains aForCausalLM class and expert fields such as num_local_experts,
n_routed_experts, or num_experts_per_tok are absent. Create
components/models/<name>/model.py, state_dict_adapter.py, __init__.py,
and optional config.py, register MODEL_ARCH_MAPPING in
_transformers/registry.py, add example YAML, and add tiny-config unit tests
plus layer-equivalence tests for rewritten layers.
config.json, referencemoe-patterns.md, map router tensors separately, preserve routed-expert
index order, map routed experts, shared experts, and gate/up/down projections,
add adapter key-map tests and tiny-config numerical equivalence tests, and do
not rely only on from_pretrained() or silent tensor reshapes.
vision_config, text_config, anda ForConditionalGeneration architecture are present. Reference
vlm-patterns.md and existing VLM implementations such as mistral4,
kimivl, or kimi_k25_vl; check text backbone, vision tower, projector,
processor assumptions, text and vision state_dict_adapter.py mappings,
registry registration, and tiny image-text tests before full checkpoints.
Do not treat VLM onboarding as a pure causal-LM path or skip processor/image
tests.
For MoE state-dict and VLM questions, apply the checklists in Sections 2.4 and 2.5.
Use this skill only when the user is adding or modifying model architecture support: model files, custom layers, state-dict adapters, Hugging Face config mapping, registry entries, or model capability flags.
Do not use this skill for standalone training recipe YAML questions about optimizers, datasets, schedulers, validation datasets, or trainer wiring unless they are explicitly part of onboarding a new model architecture. Those recipe questions belong to the nemo-automodel-recipe-development skill.
In-scope examples:
Out-of-scope examples:
Before writing code, gather information about the target model.
Download the model's config.json from the HuggingFace Hub (or use AutoConfig.from_pretrained). Key fields to extract:
architectures -- determines the class name and registration key (e.g., "LlamaForCausalLM", "Qwen3MoeForCausalLM", "Mistral3ForConditionalGeneration")model_type -- used for custom config registration in _CUSTOM_CONFIG_REGISTRATIONS if HF does not have a built-in config classhidden_size, intermediate_size, num_hidden_layers, num_attention_heads, num_key_value_heads -- sizingvocab_size -- needed for tiny test configstie_word_embeddings -- the saved setting in each supported checkpoint; do not infer it from a bare config constructorhidden_act -- activation function (e.g., "silu" for SwiGLU)| Type | Indicators | Pattern file |
|------|-----------|-------------|
| Dense LLM | ForCausalLM in architectures, no expert fields | llm-patterns.md |
| MoE LLM | n_routed_experts, num_local_experts, num_experts_per_tok in config | moe-patterns.md |
| VLM | ForConditionalGeneration in architectures, has vision_config + text_config | vlm-patterns.md |
Look in components/models/ for architectures with similar attention or MLP patterns:
components/models/
llama/ # Standard GQA + SwiGLU (CombinedQKV + CombinedGateUpMLP)
qwen2/ # Same as Llama but with attention bias + QKV bias
baichuan/ # ALiBi attention variant
deepseek_v3/ # MLA attention + MoE (DeepSeek-style grouped experts)
mistral4/ # MLA + MoE + VLM (Pixtral vision)
kimivl/ # DeepSeek-V3 backbone + MoonVit vision
kimi_k25_vl/ # Updated KimiVL with different projector
qwen3_moe/ # Qwen3 with MoE layers
nemotron_v3/ # Hybrid mamba-attention
Check whether the model needs:
AutoConfig cannot parse the model's config.json (check auto_map field)For unit tests, create a tiny config. Target: ~1M parameters or less.
# Example tiny config for a Llama-like model:
tiny_config = LlamaConfig(
hidden_size=64,
intermediate_size=128,
num_hidden_layers=2,
num_attention_heads=4,
num_key_value_heads=2,
vocab_size=256,
max_position_embeddings=128,
)
components/models/<name>/
__init__.py
model.py
state_dict_adapter.py
config.py # Only if HF config is insufficient
layers.py # Only for MoE / MLA / other non-standard layers
rope_utils.py # Only for custom RoPE
Implement files in dependency order:
PretrainedConfig subclassForCausalLM (or ForConditionalGeneration) classSee the pattern files for detailed implementation guidance:
Every registered model class with a causal lm_head must:
tie_word_embeddings_support: TieSupport as BOTH, TIED_ONLY, orUNTIED_ONLY.
reject_unsupported_tie_word_embeddings(type(self), config) at the topof __init__, using the original config before unwrapping text_config or
thinker_config.
Only classes with no causal LM head may be explicitly exempted from the registry
test.
Choose the policy from the implementation and the actual supported checkpoint
configs, not from a bare config constructor:
BOTH: tied and untied configurations are both supported.TIED_ONLY: only a tied configuration is supported.UNTIED_ONLY: only an untied configuration is supported.Runtime helpers must treat TIED_ONLY and UNTIED_ONLY as authoritative and
only resolve a per-checkpoint config flag for BOTH. All current BOTH VLMs
honor the outer tie_word_embeddings flag, so do not add a model-specific
resolver until a supported BOTH model actually requires another config path.
For BOTH and TIED_ONLY, always declare _tied_weights_keys and implement
tie_weights() with the actual lm_head and input-embedding FQNs. Do not rely
on inherited Hugging Face tying, and re-tie after any language-model swap.
Add policy-specific tests:
BOTH: tied aliases; untied does not alias.TIED_ONLY: tied aliases; untied is rejected.UNTIED_ONLY: weights stay separate; tied is rejected.Do not tie architectures with intentionally separate heads, asymmetric vocab
sizes, or stages that do not own both tensors.
For from_pretrained, the checkpoint's saved tie_word_embeddings value is
authoritative, even for BOTH. The NeMoAuto* bridge rejects flips in either
direction. A model-owned from_pretrained that bypasses that bridge must call
`reject_tie_word_embeddings_flip(checkpoint_config, requested_config,
model_class_name)`.
For MoE models, do not stop at generic loading. The adapter must explicitly map:
Add tests that assert expected key mappings and run numerical equivalence with tiny configs before trying full checkpoints.
Do not use these shortcuts:
from_pretrained().and NeMo layouts require it and a test proves the conversion is reversible.
For VLMs, confirm the Hugging Face config has vision_config and text_config
and that architectures points to a conditional-generation class. Start from
the closest VLM pattern file, usually vlm-patterns.md, and
compare existing implementations such as mistral4, kimivl, or
kimi_k25_vl.
The implementation should explicitly cover:
state_dict_adapter.py.ForConditionalGeneration class in _transformers/registry.py.Add the model to MODEL_ARCH_MAPPING in _transformers/registry.py:
# In _transformers/registry.py
MODEL_ARCH_MAPPING = OrderedDict([
# ... existing entries ...
(
"NewModelForCausalLM",
("nemo_automodel.components.models.new_model.model", "NewModelForCausalLM"),
),
])
If the model has a custom config class with auto_map in its config.json, also register in _CUSTOM_CONFIG_REGISTRATIONS:
_CUSTOM_CONFIG_REGISTRATIONS: Dict[str, Tuple[str, str]] = {
# ... existing entries ...
"new_model": ("nemo_automodel.components.models.new_model.configuration", "NewModelConfig"),
}
Every class registered in MODEL_ARCH_MAPPING must declare parallelism
capabilities, either with a static nested ModelCapabilities dataclass or a
variant-aware get_capabilities(cls, config) method. Pick exactly one pattern.
Capabilities should reflect recipe YAMLs that have been validated end to end.
If the model has precision-sensitive parameters such as Mamba A_log /
dt_bias, MoE sigmoid gate bias, attention-sink bias, or per-head scale,
declare _keep_in_fp32_modules_strict so sharding keeps those params in fp32
compute. See capabilities-and-precision.md
for examples, variant dispatch rules, and frozen-submodule dtype guidance.
This phase is only for adding a minimal example config that proves the newly
onboarded architecture can load and run. Use nemo-automodel-recipe-development for general
recipe authoring or existing recipe modifications.
Create an example config under examples/llm_finetune/<name>/ (or examples/vlm_finetune/<name>/):
model:
_target_: nemo_automodel.NeMoAutoModelForCausalLM.from_pretrained
pretrained_model_name_or_path: <org>/<model-name>
trainer:
max_steps: 100
gradient_clip_val: 1.0
accumulate_grad_batches: 1
# ... data, optimizer config ...
Test that the model loads from a HuggingFace checkpoint:
from nemo_automodel import NeMoAutoModelForCausalLM
model = NeMoAutoModelForCausalLM.from_pretrained("<org>/<model-name>")
Before using full-size models, verify with a tiny config (1-2 layers, small hidden dim) to catch shape mismatches early.
Create tests/unit_tests/models/<name>/ and cover the checks below before
loading full checkpoints:
from_hf -> to_hf preserves mapped names,shapes, dtypes, and values.
RoPE, or MoE layer. Use the model dtype from config, identical seeded weights,
identical inputs, and dtype-appropriate torch.allclose tolerances.
Edit the appropriate file in docs/model-coverage/:
docs/model-coverage/llm/index.mddocs/model-coverage/vlm/index.mdAdd a row with the model name, supported features (TP, PP, FSDP, LoRA, QLoRA), and any limitations.
After implementation and unit tests are complete, run the full parity-testing
workflow to verify that the new model produces numerically equivalent results to
the reference HuggingFace implementation.
Run three levels of comparison:
into the NeMo AutoModel layout, export it back, and verify that all mapped
tensors match the reference names, shapes, dtypes, and values within the
expected tolerance.
RoPE, and MoE components against the HuggingFace implementation with fixed
seeds and identical dtype.
on the same tokenized input and compare logits, hidden states, and loss.
Do not skip this phase. A model that passes unit tests can still diverge from HF
due to subtle weight-conversion bugs, backend differences, or RoPE mismatches
that only surface in a full parity comparison.
| File | Purpose |
|------|---------|
| _transformers/registry.py | MODEL_ARCH_MAPPING and _CUSTOM_CONFIG_REGISTRATIONS |
| components/models/common/__init__.py | Exports CombinedQKVAttentionMixin, CombinedGateUpMLP, BackendConfig, HFCheckpointingMixin, etc. |
| components/models/common/combined_projection/combined_qkv.py | CombinedQKVAttentionMixin with setup_qkv_projection() and compute_qkv() |
| components/models/common/combined_projection/combined_mlp.py | CombinedGateUpMLP with interleaved gate/up layout |
| components/models/common/combined_projection/state_dict_adapter.py | CombinedProjectionStateDictAdapter base class |
| components/models/common/hf_checkpointing_mixin.py | HFCheckpointingMixin for save/load |
| components/models/common/utils.py | BackendConfig, initialize_rms_norm_module, initialize_linear_module, get_rope_config |
| components/moe/config.py | MoEConfig dataclass |
| components/moe/fsdp_mixin.py | MoEFSDPSyncMixin for distributed expert handling |
| components/moe/layers.py | MoE layer, MLP (dense) for MoE blocks |
| components/moe/experts.py | GroupedExperts, GroupedExpertsDeepEP, GroupedExpertsTE |
config.json from HuggingFacecomponents/models/<name>/ directoryHFCheckpointingMixinMODEL_ARCH_MAPPING in _transformers/registry.py_CUSTOM_CONFIG_REGISTRATIONS (if applicable)ModelCapabilities nested dataclass (static) OR get_capabilities(cls, config) classmethod (variant dispatch, e.g. ERNIE-4.5 MoE vs dense) — never both, never neitherTieSupport and called the constructor guard for every class with a causal lm_head (or added an explicit no-head exemption) -- see §2.3_tied_weights_keys and tie_weights() for BOTH / TIED_ONLY, plus policy-specific alias and rejection tests -- see §2.3from_pretrained that bypasses the NeMoAuto* bridge against checkpoint flips -- see §2.3NeMoAutoModelForCausalLM.from_pretrained()_keep_in_fp32_modules_strict for every intrinsically-fp32 param (SSM A_log/dt_bias, Mamba D when reference-fp32, MoE gate bias, attention-sink bias, scale, …) — see §2.7ModelClass = <Name>ForCausalLM at module bottomTake nvidia/nemo-automodel-model-onboarding from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.