Install and verify cuPyNumeric for Python — requirements, commands, verification. Source builds are out of scope.
npx skills add https://github.com/NVIDIA/skills --skill cupynumeric-install
Use this skill to install cuPyNumeric for *use* from Python and to verify the install actually works (including GPU usage). Apply it whenever a user wants cuPyNumeric running via conda or pip. Do not use it to build from source (to modify or contribute) — that is out of scope.
pip install, conda install, or any installer. Print the command; let the user run it.--version checks are fine.Confirm these system requirements before recommending any install:
Follow these steps in order: confirm the prerequisites, ask the scoping questions, install via the chosen path, then verify.
conda --version and pip --version. Prefer conda (upstream-recommended); fall back to pip.nvidia-smi / nvcc --version.If neither conda nor pip is available, install one. Provide the command and the docs link; do not run it.
curl -L -O "https://github.com/conda-forge/miniforge/releases/latest/download/Miniforge3-$(uname)-$(uname -m).sh"
bash "Miniforge3-$(uname)-$(uname -m).sh"
Docs: https://github.com/conda-forge/miniforge
Install Python from your OS package manager (apt/dnf/brew) or https://www.python.org/downloads/. If pip is missing on an existing Python: python -m ensurepip --upgrade.
After installing, open a new shell so the binary is on PATH.
conda create -n cupynumeric -c conda-forge -c legate cupynumeric
conda activate cupynumeric
Into an existing env: conda install -c conda-forge -c legate cupynumeric.
conda auto-selects the GPU vs CPU variant from whether nvidia-smi works at install time. To override that, see below.
Set CONDA_OVERRIDE_CUDA only when no GPU is visible at install time (e.g. building a container for a GPU host). Use the runtime host's CUDA version:
CONDA_OVERRIDE_CUDA="12.2" conda install -c conda-forge -c legate cupynumeric
conda install -c conda-forge -c legate-nightly cupynumeric
python -m venv .venv
source .venv/bin/activate
pip install nvidia-cupynumeric
Run a self-contained script through the legate launcher — no repo checkout needed.
TMP=$(mktemp -d)
cat > "$TMP/smoke.py" <<'EOF'
import cupynumeric as np
a = np.arange(10)
b = np.ones((4, 4))
print("sum:", a.sum()) # expect 45
print("matmul:", (b @ b).sum()) # expect 64.0
EOF
legate "$TMP/smoke.py"
rm -rf "$TMP"
Expect sum: 45 and matmul: 64.0. If legate is missing, the env is not activated — see Troubleshooting.
A passing smoke test does not prove GPU usage — a CPU-variant install on a GPU box produces correct results too. Run both steps.
1. Force a GPU launch. legate --gpus N requests N GPUs; fails fast if no GPU is visible or the CPU variant is installed.
TMP=$(mktemp -d)
cat > "$TMP/check.py" <<'EOF'
import cupynumeric as np
print(np.ones((4096, 4096)).sum())
EOF
legate --gpus 1 "$TMP/check.py"
rm -rf "$TMP"
Expect 16777216.0. If you see CUDA driver, libcudart, or no GPUs available, the CPU variant is installed; reinstall with CONDA_OVERRIDE_CUDA.
2. Confirm the GPU was touched. Run a deadline-bounded matmul loop alongside nvidia-smi, all from one shell — no second-terminal race:
TMPDIR_GPU=$(mktemp -d)
SCRIPT="$TMPDIR_GPU/cupynumeric_gpu_check.py"
cat > "$SCRIPT" <<'EOF'
import cupynumeric as np, time
a = np.ones((10000, 10000))
deadline = time.time() + 20
iters = 0
while time.time() < deadline:
b = a @ a
_ = float(b.sum()) # force sync so the matmul actually runs
iters += 1
print("iters:", iters)
EOF
legate --gpus 1 "$SCRIPT" &
WORKLOAD=$!
sleep 5 # buffer for Legate startup
for _ in $(seq 10); do # 10 samples at 1s — covers slow startup
nvidia-smi --query-gpu=utilization.gpu,memory.used --format=csv,noheader
sleep 1
done
wait "$WORKLOAD"
rm -rf "$TMPDIR_GPU"
Expect memory.used in the GiB range across most samples and non-trivial
utilization.gpu in several. If both stay at baseline across every sample, the
GPU variant is not installed — check conda list cupynumeric for *_gpu (not
*_cpu).
See verification_examples.md for multi-GPU checks, CPU fallback, container, and troubleshooting.
pip uninstall nvidia-cupynumeric or conda remove cupynumeric first.legate launcher for multi-GPU / multi-rank runs. Plain python runs single-process: legate --gpus 2 script.py.CONDA_OVERRIDE_CUDA. conda otherwise auto-selects the CPU or GPU variant from nvidia-smi at install time.conda --version ≥ 24.1. Older releases silently break variant selection.ModuleNotFoundError: No module named 'cupynumeric' → Run which python and pip list | grep cupynumeric (or conda list | grep cupynumeric) from the same shell to find the env mismatch.ImportError mentioning CUDA / libcudart → Reinstall with CONDA_OVERRIDE_CUDA="<your-cuda-version>"; the CPU variant is on a GPU box, or CUDA versions are mismatched.legate: command not found → Activate the env, then run which legate to confirm.Guide for creating high-quality MCP (Model Context Protocol) servers that enable LLMs to interact with external services through well-designed tools. Use when building MCP servers to integrate external APIs or services, whether in Python (FastMCP) or Node/TypeScript (MCP SDK).
Automatically creates user-facing changelogs from git commits by analyzing commit history, categorizing changes, and transforming technical commits into clear, customer-friendly release notes. Turns hours of manual changelog writing into minutes of automated generation.
Use when implementation is complete, all tests pass, and you need to decide how to integrate the work - guides completion of development work by presenting structured options for merge, PR, or cleanup
Guide for creating high-quality MCP (Model Context Protocol) servers that enable LLMs to interact with external services through well-designed tools. Use when building MCP servers to integrate external APIs or services, whether in Python (FastMCP) or Node/TypeScript (MCP SDK).
React Native and Expo best practices for building performant mobile apps. Use when building React Native components, optimizing list performance, implementing animations, or working with native modules. Triggers on tasks involving React Native, Expo, mobile performance, or native platform APIs.
React and Next.js performance optimization guidelines from Vercel Engineering. This skill should be used when writing, reviewing, or refactoring React/Next.js code to ensure optimal performance patterns. Triggers on tasks involving React components, Next.js pages, data fetching, bundle optimization, or performance improvements.
Next.js best practices - file conventions, RSC boundaries, data patterns, async APIs, metadata, error handling, route handlers, image/font optimization, bundling
Use when starting feature work that needs isolation from current workspace or before executing implementation plans - creates isolated git worktrees with smart directory selection and safety verification
Take nvidia/cupynumeric-install from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.
The instructions reference pip.
Without those the skill loads but fails at the first command.