xuzhougeng/scgpt
> Embed and annotate single-cell expression data with scGPT, a foundation model (1) Producing cell embeddings from an AnnData for clustering/integration, (2) Zero-shot or fine-tuned cell-type annotation, (3) Gene-level representation for perturbation/GRN tasks. For probabilistic single-cell models (scVI etc.), use the scvi-tools library.
npx skills add https://github.com/xuzhougeng/wisp-science --skill scgpt
| Requirement | Minimum | Recommended |
| ----------- | ------- | ----------- |
| Python | 3.10+ | 3.11 |
| CUDA | 12.1+ | 12.4+ |
| GPU VRAM | 16 GB | 24 GB+ |
scGPT checkpoints are raw directories (args.json, best_model.pt,
vocab.json) — not Hugging Face hub repos. Point at the directory, not an HF
repo id.
from scgpt.tokenizer.gene_tokenizer import GeneVocab
gv = GeneVocab.from_file("/path/to/scgpt-human/vocab.json")
print(len(gv)) # 60697 for the released human checkpoint
import anndata as ad
from scgpt.tasks import embed_data
adata = ad.read_h5ad("dataset.h5ad") # var must contain a gene-name column
emb = embed_data(
adata,
model_dir="/path/to/scgpt-human",
gene_col="feature_name",
use_fast_transformer=False, # see Gotchas
)
# emb is an AnnData with .obsm["X_scGPT"]
embed_data returns an AnnData whose .obsm["X_scGPT"] is the per-cell
embedding (n_cells × emb_dim, 512 by default). Downstream: feed to
scanpy.pp.neighbors / scanpy.tl.umap.
Needs ≥24 GB VRAM and the released human checkpoint (~200 MB:
args.json, best_model.pt, vocab.json). Use a selected and probed
ssh:<alias> context and load remote-compute-ssh. Confirm the environment
and checkpoint with bounded read-only discovery, then write a self-contained
runs/scgpt_embed.py and submit it with run_in_context:
{
"context_id": "ssh:gpu-box",
"title": "scGPT embedding for 50k cells",
"command": "source ~/miniforge3/etc/profile.d/conda.sh && conda activate scgpt && python scgpt_embed.py --input dataset.h5ad --model-dir /srv/models/scgpt-human --output /home/me/wisp-results/scgpt/embedded.h5ad",
"timeout_secs": 1800,
"input_paths": ["runs/scgpt_embed.py", "data/dataset.h5ad"],
"output_specs": [
{
"glob": "ssh://gpu-box/home/me/wisp-results/scgpt/embedded.h5ad",
"kind": "h5ad",
"residency": "remote"
}
]
}
Replace every context and remote path with discovered values. For large data
already on the server, use an absolute remote path instead of staging it. Call
monitor_run once to wait, get_run once for a snapshot, or cancel_run to
stop. If flash-attn is unavailable in that environment, set
use_fast_transformer=False.
use_fast_transformer default is True but resolves to a FlashAttentionpath that may not import in every env. Pass use_fast_transformer=False
unless you've confirmed flash_attn loads cleanly.
torchtext.vocab.Vocab; inenvironments without torchtext a pure-Python shim provides Vocab —
functionally identical for GeneVocab, but if you hit
AttributeError: 'Vocab' object has no attribute …, you're on a stale shim.
gene_col to the column in adata.var that holds symbols.
| Symptom | Fix |
| ------------------------------------------------- | ------------------------------------------------ |
| flash_attn is not installed warning at import | Harmless; pass use_fast_transformer=False |
| 'Vocab' object has no attribute 'vocab' | Env has an old torchtext shim — update the env |
| Nearly all genes dropped | Wrong gene_col; check adata.var.columns |
| "scgpt not in manifest" / env-detection misses scGPT | The baked env manifest lists the distribution as scGPT (and flash_attn), pip's canonical casing — normalize manifest keys before lookup: name.lower().replace('-', '_') |
Next: cluster/annotate the embedding with the scanpy library
(sc.pp.neighbors → sc.tl.leiden / sc.tl.umap), or compare to an
scvi-tools latent space on the same data.
Take xuzhougeng/scgpt from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.