nvidia/sglang-setup
Deploy an SGLang inference server on an NVIDIA DGX Station GB300 with the cu130 container, RadixAttention prefix caching, and structured JSON output support. Use when the user asks to serve a model with SGLang, start an SGLang endpoint, or needs structured-output inference on DGX Station.
npx skills add https://github.com/NVIDIA/dgx-spark-playbooks --skill sglang-setup
Deploy an SGLang inference server on DGX Station with validated configuration.
nvidia-smi --query-gpu=index,name --format=csv,noheader
Identify the device index for the GB300 (typically device 1). Use this index for --gpus below. Do NOT use --gpus all — mixed coherency will cause CUDA failures.
Qwen/Qwen3-8B — small, fast, good for testingQwen/Qwen3-32B — medium, good balancemeta-llama/Llama-3.1-70B-Instruct — large general-purpose-e HF_TOKEN="...". docker pull lmsysorg/sglang:latest-cu130
docker run -d \
--name sglang-server \
--gpus '"device=<GB300_INDEX>"' \
--ipc host \
--ulimit memlock=-1 \
--ulimit stack=67108864 \
-p 30000:30000 \
-e HF_TOKEN="<TOKEN>" \
-v "$HOME/.cache/huggingface/hub:/root/.cache/huggingface/hub" \
lmsysorg/sglang:latest-cu130 \
sglang serve --model-path "<MODEL>" \
--host 0.0.0.0 \
--port 30000 \
--context-length 32768 \
--mem-fraction-static 0.85
Container version: Use lmsysorg/sglang:latest-cu130. The cu130 tag is required for Blackwell SM103 support.
First launch downloads the model and compiles kernels. This takes extra time — subsequent starts are faster.
docker logs -f sglang-server
curl http://localhost:30000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "<MODEL>",
"messages": [{"role": "user", "content": "Hello"}],
"max_tokens": 64
}'
docker stop sglang-server && docker rm sglang-serverdocker logs sglang-server 2>&1 | grep "cached-token" | tail -5response_format.json_schema in API requests for guaranteed valid JSON.--chunked-prefill-size 8192 to break long prefills into chunks, reducing time-to-first-token.| Parameter | Default | Agent workloads | Throughput workloads |
|-----------|---------|-----------------|---------------------|
| --context-length | 32768 | 32768-65536 | 8192-16384 |
| --mem-fraction-static | 0.85 | 0.80-0.85 | 0.85-0.88 |
| --chunked-prefill-size | off | 4096-8192 | 8192 |
| --enable-metrics | off | Optional | Recommended |
curl http://localhost:30000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "<MODEL>",
"messages": [{"role": "user", "content": "List three programming languages."}],
"max_tokens": 512,
"response_format": {
"type": "json_schema",
"json_schema": {
"name": "languages",
"schema": {
"type": "object",
"properties": {
"languages": {
"type": "array",
"items": {
"type": "object",
"properties": {
"name": {"type": "string"},
"primary_use": {"type": "string"}
},
"required": ["name", "primary_use"]
}
}
},
"required": ["languages"]
}
}
}
}'
Take nvidia/sglang-setup from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.
The instructions reference docker.
Without those the skill loads but fails at the first command.