> Score, embed, and generate DNA sequences with Evo 2, a long-context genomic (1) Computing per-nucleotide or per-sequence likelihoods for variant effect scoring, (2) Embedding genomic windows for downstream classification, (3) Generating DNA conditioned on a prefix, (4) Scoring regulatory or coding regions across species.
npx skills add https://github.com/xuzhougeng/wisp-science --skill evo2
| Requirement | Minimum | Recommended |
| ----------- | ------- | ---------------- |
| Python | 3.11 | 3.12 (<3.13) |
| CUDA | 12.1+ | 12.4+ |
| GPU VRAM | 24 GB (7B bf16) | 80 GB (40B) |
| RAM | 32 GB | 128 GB |
pip install evo2
# Weights pulled from Hugging Face on first model load.
from evo2 import Evo2
model = Evo2("evo2_7b") # or "evo2_40b" — see model table
seqs = ["ATCG" * 50, "GGGCTTAA" * 25]
ll = model.score_sequences(seqs) # → list[float], mean per-token log-likelihood
print(ll)
out = model.generate(
prompt_seqs=["ATGAAAGCT"],
n_tokens=256,
temperature=0.7,
)
print(out.sequences[0])
| Name | Params | Context | VRAM (bf16) | Notes |
| ----------- | ------ | ------- | ----------- | ---------------------------------- |
| evo2_7b | 7 B | 1 M nt | ~22 GB | Default; fits on a single 24 GB+ GPU |
| evo2_40b | 40 B | 1 M nt | ~78 GB | H100 80 GB or multi-GPU |
| evo2_1b_base | 1 B | 8 K nt | ~6 GB | FP8 path requires sm_89+ (H100) |
score_sequences returns a list[float] (or np.ndarray) of mean log-likelihoods,
one per input sequence. More negative ⇒ less likely under the model. For variant
effect, compute Δll = ll_alt - ll_ref over a fixed window.
generate returns a GenerationOutput with .sequences (list[str]), .logits
(list[Tensor]), and .logprobs_mean (list[float]) — always populated, no flag required.
Need a DNA model?
│
├─ Per-base/per-sequence likelihood, generation → Evo 2 ✓
├─ Predict experimental tracks (expression, accessibility) → borzoi
└─ Protein, not DNA → fair-esm2 / esmfold2
7B/40B inference is GPU-bound (≥24 GB / 80 GB VRAM). Use a selected and
probed ssh:<alias> context and load remote-compute-ssh. Confirm that the
environment imports Evo 2 and that the desired weights are cached. Submit a
self-contained scoring script through one run_in_context call:
{
"context_id": "ssh:gpu-box",
"title": "Evo 2 variant scoring",
"command": "source ~/miniforge3/etc/profile.d/conda.sh && conda activate evo2 && HF_HOME=/srv/model-cache HF_HUB_OFFLINE=1 python score_evo2.py --output /home/me/wisp-results/evo2/scores.json",
"timeout_secs": 1800,
"input_paths": ["runs/score_evo2.py"],
"output_specs": [
{
"glob": "ssh://gpu-box/home/me/wisp-results/evo2/scores.json",
"kind": "json",
"residency": "remote"
}
]
}
Replace context, environment, cache, and output paths with discovered values.
Call monitor_run once to wait, get_run once for a snapshot, or cancel_run
to stop. Set HF_HUB_OFFLINE=1 only after confirming the cache is complete, so
the loader does not try to write refs/ into a read-only mount. Weight footprint:
~15 GB (7B), ~80 GB (40B).
| Task | 7B on H100 | Notes |
| --------------------------- | ---------- | --------------------------- |
| Model load (cached) | ~5-7 min | First call hydrates weights |
| score_sequences, 200×200bp| ~10-20 s | After load |
| generate, 1×512 nt | ~15 s | |
| Symptom | Cause | Fix |
| ------------------------------------ | ------------------------------ | ------------------------------------------ |
| Transformer Engine not installed | No FP8 — falls back to bf16 | Informational only on non-H100; ignore |
| OOM on load | 40B on <80 GB GPU | Use evo2_7b or shard with device_map |
| HF tries to write refs/main | HF_HOME points at RO mount | Set HF_HUB_OFFLINE=1 |
| dtype mismatch in score_sequences| Passing tensors not strings | Pass list[str]; the API tokenises for you |
Next: pair with borzoi to predict track-level effects of the same
variants.
Create new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, edit, or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy.
Access NCBI GEO for gene expression/genomics data. Search/download microarray and RNA-seq datasets (GSE, GSM, GPL), retrieve SOFT/Matrix files, for transcriptomics and expression analysis.
Bayesian modeling with PyMC. Build hierarchical models, MCMC (NUTS), variational inference, LOO/WAIC comparison, posterior checks, for probabilistic programming and inference.
Multi-objective optimization framework. NSGA-II, NSGA-III, MOEA/D, Pareto fronts, constraint handling, benchmarks (ZDT, DTLZ), for engineering design and optimization problems.
Statistical modeling toolkit. OLS, GLM, logistic, ARIMA, time series, hypothesis tests, diagnostics, AIC/BIC, for rigorous statistical inference and econometric analysis.
Add unsigned integer (uint) type support to PyTorch operators by updating AT_DISPATCH macros. Use when adding support for uint16, uint32, uint64 types to operators, kernels, or when user mentions enabling unsigned types, barebones unsigned types, or uint support.
Convert PyTorch AT_DISPATCH macros to AT_DISPATCH_V2 format in ATen C++ code. Use when porting AT_DISPATCH_ALL_TYPES_AND*, AT_DISPATCH_FLOATING_TYPES*, or other dispatch macros to the new v2 API. For ATen kernel files, CUDA kernels, and native operator implementations.
Write docstrings for PyTorch functions and methods following PyTorch conventions. Use when writing or updating docstrings in PyTorch code.
Take xuzhougeng/evo2 from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.
The instructions reference pip.
Without those the skill loads but fails at the first command.