mcpbeat Sign in

Qad Agent Skill

>- Run explicitly requested ModelOpt Quantization-Aware Distillation (QAD) on Slurm through Megatron Bridge to recover a measured BF16-to-PTQ accuracy gap. Use only when the user explicitly asks for QAD, including its topology, data preparation, Slurm launch, resume, checkpoint export, or recovery decisions.

1k tokens
context cost
the whole folder, loaded on every use
1
files
instructions only
0
copies elsewhere
how many repositories repackaged it
3381
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/NVIDIA/Model-Optimizer --skill qad

The instruction itself

5 sections, as written by the author

ModelOpt Quantization-Aware Distillation

QAD is expensive. Run it only when the user explicitly authorizes QAD for the

target model or run. A Day-0, PTQ, evaluation, comparison, or recipe-search

request alone is not authorization to start QAD.

Follow the supported workflow

Before constructing commands, read:

  • examples/megatron_bridge/README.md, especially PTQ, data preparation, QAD,

export, and Slurm usage

  • examples/megatron_bridge/{quantize.py,distill.py} via --help
  • skills/common/{environment-setup,workspace-management,slurm-setup}.md; also

skills/common/remote-execution.md for remote Slurm

Treat the example README and --help output as authoritative for mutable flags,

commands, containers, and checkpoint formats. This skill supports Slurm only.

Execute in this order

  • Confirm the gap. Reuse only validated, comparable BF16/PTQ results and

the exact benchmark configuration from preceding evaluation or recipe

search; run missing, invalid, or non-comparable baselines. Confirm the target

benchmarks and their context-length needs. Stop if the PTQ gap to BF16 is

already below 1%.

  • Reproduce PTQ and verify compatibility. In the target runtime, require

AutoBridge.can_handle() for the target model and PTQ through quantize.py

to succeed while preserving the exact preceding PTQ config or recipe:

format, layer selection, calibration data/count, sequence length, and seed.

A changed quantization setting is a new PTQ candidate and must be evaluated

before QAD. In the master-rank .quant_summary.txt, require finite positive

amax for enabled static quantizers; accept dynamic/format-defined None

only when the recipe intends it. Treat the summary as rank-local under model

parallelism.

  • Choose topology explicitly. Derive the smallest fitting node count and

TP/PP/CP/EP from student and teacher architecture, the chosen sequence length,

and available GPU memory. Prefer CP before TP for small long-context models;

keep EP=1 for dense models and ETP=1 because the current distill.py

workflow does not support expert tensor parallelism. For MoE require:

  • DP = world_size / (TP * PP * CP)
  • EDP = world_size / (EP * PP)
  • integral DP/EDP, num_experts % EP == 0, and

GBS % (MBS * DP) == 0

  • Prepare the full capped dataset once. Use suitable user-provided data, or

copy examples/megatron_bridge/data/nemotron-cascade-2-blend.yaml as the

default. Set the target tokenizer and workspace path, then materialize the

randomly sampled subset before training. Pack the chosen sequence length;

Megatron's 99,1,0 split creates the 1% validation holdout from the same

data.

  • Run and monitor QAD. Run one QAD training job at a time and fold startup

validation into it; do not submit separate GPU preflight jobs or split at

recovery iterations. Let training continue while evaluating saved

checkpoints, and cancel it when a stop condition below is met.

Default training policy

| Setting | Default |

| --- | --- |

| Sequence length | 32768; adjust for target benchmarks |

| Peak / minimum LR | 1e-5 / 1e-6 |

| LR schedule | cosine |

| Training cap | 1000 iterations |

| Global batch size | 512 |

| Dataset | nvidia/Nemotron-Cascade-2-SFT-Data by default |

| Materialized token budget | 17.3B at 32K; cover the full cap at the chosen length |

| Training validation | every 25 iterations; deterministic 1% holdout; 2 batches |

| Checkpoint interval | 50 iterations |

| Loss logging | every 10 iterations |

| Recovery benchmark | 150, then every 100 iterations while training runs |

| Slurm duration exit | 220 minutes for a 4-hour allocation |

Run policy

  • Keep train_iters=1000 and leave exit_interval unset.
  • From initial step timing, submit only enough sequential jobs to reach

checkpoint 150; never submit through iteration 1000 upfront. At each recovery

checkpoint, submit to the next only after its targeted evaluation and any

triggered full suite, and only if the BF16 gap remains at least 1% and

recovery has neither plateaued nor regressed.

  • Give all training jobs the same run-specific job name and

--dependency=singleton; record job IDs and, on any stop, cancel pending jobs

before the active job.

  • Cancel on non-finite loss, repeated skipped iterations, or a sustained spike.

At iteration 50, require the loss aggregate to be lower than at iteration 10.

  • At each recovery checkpoint, first evaluate the one to three benchmarks with

the largest PTQ drops. Run the remaining original PTQ suite at that checkpoint

only after recovery beyond run noise.

  • Cancel when the full-suite gap to BF16 is below 1%, benchmark recovery

regresses beyond run noise, or benchmark recovery and loss both plateau.

  • After a duration exit, resume the latest QAD checkpoint in the same output

directory with unchanged prepared data paths/cache, seed, topology, optimizer,

scheduler, iteration, and consumed-sample state; do not restart from PTQ.

  • Report the PTQ recipe, data sample, Slurm topology, loss/state, checkpoints,

and comparable BF16/PTQ/QAD results.

Other skills for the same job

different authors, same section of the catalogue
Doc Coauthoring
by anthropics
vendor ×10

Guide users through a structured workflow for co-authoring documentation. Use when user wants to write documentation, proposals, technical specs, decision docs, or similar structured content. This workflow helps users efficiently transfer context, refine content through iteration, and verify the doc works for readers. Trigger when user mentions writing docs, creating proposals, drafting specs, or similar documentation tasks.

4k tokens
File Organizer
by frostant
×10

Intelligently organizes your files and folders across your computer by understanding context, finding duplicates, suggesting better structures, and automating cleanup tasks. Reduces cognitive load and keeps your digital workspace tidy without manual effort.

3k tokens
Domain Name Brainstormer
by frostant
×8

Generates creative domain name ideas for your project and checks availability across multiple TLDs (.com, .io, .dev, .ai, etc.). Saves hours of brainstorming and manual checking.

1k tokens
Brainstorming
by ZhanlinCui
×4

You MUST use this before any creative work - creating features, building components, adding functionality, or modifying behavior. Explores user intent, requirements and design before implementation.

626 tokens
Planning With Files
by ZhanlinCui
×3

Implements Manus-style file-based planning for complex tasks. Creates task_plan.md, findings.md, and progress.md. Use when starting complex multi-step tasks, research projects, or any task requiring >5 tool calls.

9k tokens scripts
Scientific Brainstorming
by christophacham
×3

Creative research ideation and exploration. Use for open-ended brainstorming sessions, exploring interdisciplinary connections, challenging assumptions, or identifying research gaps. Best for early-stage research planning when you do not have specific observations yet. For formulating testable hypotheses from data use hypothesis-generation.

5k tokens
GitHub Project Management
by ComeOnOliver
×3

Comprehensive GitHub project management with swarm-coordinated issue tracking, project board automation, and sprint planning

14k tokens
Grill Me
by ComeOnOliver
×3

Interview the user relentlessly about a plan or design until reaching shared understanding, resolving each branch of the decision tree. Use when user wants to stress-test a plan, get grilled on their design, or mentions "grill me".

3k tokens

How to use it

Copy the folder

Take nvidia/qad from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.