Develop, debug, and optimize SGLang LLM serving engine. Use when the user mentions SGLang, sglang, srt, sgl-kernel, LLM serving, model inference, KV cache, attention backend, FlashInfer backend, MLA, MoE routing, MoE dispatch, expert parallelism SGLang, speculative decoding, disaggregated serving, TP/PP/EP, radix cache, continuous batching, chunked prefill, CUDA graph SGLang, model loading, quantization FP8/GPTQ/AWQ, JIT kernel, triton kernel SGLang, DeepSeek serving, EPLB (expert load balancing), HiCache, launch_server, sglang Engine API, LoRA inference, torch.compile SGLang, or asks about serving LLMs with SGLang. Also use when the user wants to add a new model to SGLang, add a new attention backend, debug SGLang serving issues, or optimize SGLang throughput/latency.
npx skills add https://github.com/slowlyC/agent-gpu-skills --skill sglang-skill
SGLang 源码位于此 skill 安装目录下的 repos/sglang/。
实际路径取决于所用工具:
~/.cursor/skills/sglang-skill/repos/sglang/~/.claude/skills/sglang-skill/repos/sglang/~/.codex/skills/sglang-skill/repos/sglang/SGLANG_REPO: 下文示例用 ~/.cursor/skills/sglang-skill/repos/sglang/ 作占位符,替换为实际路径。
如果该路径不存在,在项目目录下运行 bash update-repos.sh sglang。
SGLANG_REPO/python/sglang/srt/
├── layers/
│ ├── attention/ # Attention backends
│ │ ├── flashinfer_backend.py # FlashInfer (默认)
│ │ ├── flashinfer_mla_backend.py # FlashInfer MLA (DeepSeek)
│ │ ├── cutlass_mla_backend.py # CUTLASS MLA
│ │ ├── flashattention_backend.py # FlashAttention
│ │ ├── triton_backend.py # Triton attention
│ │ ├── flashmla_backend.py # FlashMLA
│ │ ├── nsa_backend.py # Native Sparse Attention
│ │ ├── tbo_backend.py # TBO
│ │ ├── fla/ # Flash Linear Attention
│ │ ├── triton_ops/ # Triton attention ops
│ │ └── wave_ops/ # Wave attention ops
│ ├── moe/ # MoE routing and dispatch
│ ├── quantization/ # FP8, GPTQ, AWQ, Marlin, etc.
│ ├── deep_gemm_wrapper/ # DeepGEMM 集成
│ └── utils/
├── models/ # 模型实现 (LLaMA, DeepSeek, Qwen, etc.)
│ └── deepseek_common/ # DeepSeek V2/V3 共享组件
├── managers/ # Scheduler, TokenizerManager, Detokenizer
├── mem_cache/ # KV cache, Radix cache
├── model_executor/ # 模型执行器, forward batch
├── model_loader/ # 模型加载, 权重映射
├── entrypoints/ # 启动入口: Engine, OpenAI API server
├── speculative/ # Speculative decoding
├── disaggregation/ # Disaggregated prefill/decode
├── distributed/ # TP/PP/EP 分布式
├── compilation/ # CUDA Graph, Torch.compile
├── configs/ # 模型配置
├── lora/ # LoRA 推理
├── eplb/ # Expert-level load balancing
├── hardware_backend/ # 硬件适配 (CUDA, ROCm, XPU)
└── utils/ # 工具函数
SGLANG_REPO/python/sglang/jit_kernel/
├── flash_attention/ # Flash Attention 自定义实现
├── flash_attention_v4.py # Flash Attention v4
├── cutedsl_gdn.py # CuTeDSL GDN kernel
├── concat_mla.py # MLA concat kernel
├── norm.py # Normalization kernels
├── rope.py # RoPE position encoding
├── pos_enc.py # Position encoding
├── per_tensor_quant_fp8.py # FP8 量化
├── kvcache.py # KV cache kernels
├── hicache.py # HiCache kernels
├── gptq_marlin.py # GPTQ Marlin kernel
├── cuda_wait_value.py # CUDA sync primitives
└── diffusion/ # Diffusion model kernels
SGLANG_REPO/sgl-kernel/
├── csrc/
│ ├── attention/ # Custom attention CUDA kernels
│ ├── cutlass_extensions/ # CUTLASS GEMM extensions
│ ├── gemm/ # GEMM kernels
│ ├── moe/ # MoE dispatch/combine kernels
│ ├── quantization/ # Quantization CUDA kernels
│ ├── allreduce/ # AllReduce CUDA kernels
│ ├── speculative/ # Speculative decoding kernels
│ ├── kvcacheio/ # KV cache I/O
│ ├── mamba/ # Mamba SSM kernels
│ ├── memory/ # Memory management
│ └── grammar/ # Grammar-guided generation
├── include/ # C++ headers
├── python/ # Python bindings
├── tests/ # Kernel tests
└── benchmark/ # Kernel benchmarks
SGLANG_REPO/python/sglang/lang/ # SGLang 前端 DSL
SGLANG_REPO/examples/ # 使用示例
SGLANG_REPO/benchmark/ # 性能基准
SGLANG_REPO/test/ # 测试套件
SGLANG_REPO/docs/ # 文档
用 Grep 工具搜索,不要整文件加载。
SGLANG_REPO="$HOME/.cursor/skills/sglang-skill/repos/sglang"
# 查找 attention backend 注册
rg "register\|Backend" $SGLANG_REPO/python/sglang/srt/layers/attention/attention_registry.py
# 查找 FlashInfer MLA 实现
rg "forward\|mla" $SGLANG_REPO/python/sglang/srt/layers/attention/flashinfer_mla_backend.py
# 查找 CUTLASS MLA
rg "cutlass\|mla" $SGLANG_REPO/python/sglang/srt/layers/attention/cutlass_mla_backend.py
# 查找 attention 通用接口
rg "class.*Backend\|def forward" $SGLANG_REPO/python/sglang/srt/layers/attention/base_attn_backend.py
# Scheduler 核心逻辑
rg "class Scheduler\|def get_next_batch" $SGLANG_REPO/python/sglang/srt/managers/
# Continuous batching 和 chunked prefill
rg "chunk\|prefill\|extend" $SGLANG_REPO/python/sglang/srt/managers/
# CUDA Graph
rg "cuda_graph\|CudaGraph" $SGLANG_REPO/python/sglang/srt/compilation/
# Radix cache 实现
rg "RadixCache\|radix" $SGLANG_REPO/python/sglang/srt/mem_cache/
# KV cache 管理
rg "class.*Pool\|allocate\|free" $SGLANG_REPO/python/sglang/srt/mem_cache/
# HiCache (hierarchical cache)
rg "HiCache\|hicache" $SGLANG_REPO/python/sglang/srt/mem_cache/
# 查找特定模型实现
rg "class.*ForCausalLM" $SGLANG_REPO/python/sglang/srt/models/
# DeepSeek V2/V3 实现
rg "DeepSeek\|MLA\|MoE" $SGLANG_REPO/python/sglang/srt/models/deepseek_v2.py
# 模型加载和权重映射
rg "load_weight\|weight_map" $SGLANG_REPO/python/sglang/srt/model_loader/
# MoE routing
rg "TopK\|router\|expert" $SGLANG_REPO/python/sglang/srt/layers/moe/
# MoE CUDA kernels
rg "moe" $SGLANG_REPO/sgl-kernel/csrc/moe/
# FP8 量化
rg "fp8\|float8" $SGLANG_REPO/python/sglang/srt/layers/quantization/
# GPTQ/AWQ/Marlin
rg "gptq\|awq\|marlin" $SGLANG_REPO/python/sglang/srt/layers/quantization/
rg "speculative\|draft\|verify" $SGLANG_REPO/python/sglang/srt/speculative/
# TP/PP/EP
rg "tensor_parallel\|pipeline_parallel\|expert_parallel" $SGLANG_REPO/python/sglang/srt/distributed/
# Disaggregated serving
rg "disagg\|prefill_worker\|decode_worker" $SGLANG_REPO/python/sglang/srt/disaggregation/
| Need | Source | Path |
|------|--------|------|
| Attention backend 接口 | SRT layers | srt/layers/attention/base_attn_backend.py |
| FlashInfer attention | SRT layers | srt/layers/attention/flashinfer_backend.py |
| MLA (DeepSeek) | SRT layers | srt/layers/attention/*mla*.py |
| MoE routing/dispatch | SRT layers | srt/layers/moe/ |
| 量化 (FP8/GPTQ/AWQ) | SRT layers | srt/layers/quantization/ |
| Scheduler | SRT managers | srt/managers/ |
| KV cache / Radix cache | SRT mem_cache | srt/mem_cache/ |
| 模型实现 | SRT models | srt/models/ |
| DeepSeek V2/V3 | SRT models | srt/models/deepseek_v2.py, deepseek_common/ |
| Speculative decoding | SRT speculative | srt/speculative/ |
| Disaggregated serving | SRT disagg | srt/disaggregation/ |
| TP/PP/EP 分布式 | SRT distributed | srt/distributed/ |
| CUDA Graph | SRT compilation | srt/compilation/ |
| 模型加载 | SRT model_loader | srt/model_loader/ |
| 启动入口 | SRT entrypoints | srt/entrypoints/ |
| JIT Triton kernels | jit_kernel | jit_kernel/ |
| Custom CUDA kernels | sgl-kernel | sgl-kernel/csrc/ |
| CUTLASS extensions | sgl-kernel | sgl-kernel/csrc/cutlass_extensions/ |
| 前端 DSL | lang | python/sglang/lang/ |
| 使用示例 | examples | examples/ |
base_attn_backend.py 中的 AttnBackendforward() 方法attention_registry.py 注册flashinfer_backend.py 作为模板srt/models/ 创建模型文件ForCausalLM 类load_weights() 方法srt/models/llama.py 作为模板srt/layers/quantization/ 添加量化模块fp8_kernel.py 或 gptq.py# 启动 OpenAI 兼容 API server
python -m sglang.launch_server --model-path meta-llama/Meta-Llama-3-8B-Instruct --tp 1
# 使用 Engine API (Python)
from sglang import Engine
engine = Engine(model_path="meta-llama/Meta-Llama-3-8B-Instruct")
# Profiling
python -m sglang.launch_server --model-path ... --enable-torch-compile
nsys profile -o report python -m sglang.launch_server ...
# 在 cursor-gpu-skills 项目目录下
bash update-repos.sh sglang
Convert PyTorch AT_DISPATCH macros to AT_DISPATCH_V2 format in ATen C++ code. Use when porting AT_DISPATCH_ALL_TYPES_AND*, AT_DISPATCH_FLOATING_TYPES*, or other dispatch macros to the new v2 API. For ATen kernel files, CUDA kernels, and native operator implementations.
Write docstrings for PyTorch functions and methods following PyTorch conventions. Use when writing or updating docstrings in PyTorch code.
Statistical models library for Python. Use when you need specific model classes (OLS, GLM, mixed models, ARIMA) with detailed diagnostics, residuals, and inference. Best for econometrics, time series, rigorous inference with coefficient tables. For guided statistical test selection with APA reporting use statistical-analysis.
Answer questions about the AI SDK and help build AI-powered features. Use when developers: (1) Ask about AI SDK functions like generateText, streamText, ToolLoopAgent, embed, or tools, (2) Want to build AI agents, chatbots, RAG systems, or text generation features, (3) Have questions about AI providers (OpenAI, Anthropic, Google, etc.), streaming, tool calling, structured output, or embeddings, (4) Use React hooks like useChat or useCompletion. Triggers on: "AI SDK", "Vercel AI SDK", "generateText", "streamText", "add AI to my app", "build an agent", "tool calling", "structured output", "useChat".
Create an llms.txt file from scratch based on repository structure following the llms.txt specification at https://llmstxt.org/
Use when working directly with the `esm` Python SDK, ESM3 or ESMC model IDs, Forge/Biohub inference clients, or ESMFold2 folding workflows.
Modal is a serverless cloud platform for running Python on demand, including on-demand GPUs. Use when deploying or serving AI/ML models, running GPU-accelerated workloads (training, fine-tuning, inference), serving web endpoints, scheduling batch jobs, or scaling Python code to cloud containers with the Modal SDK.
Use Therapeutics Data Commons through the PyTDC Python package for registry discovery, approved dataset access, task-aware splits, evaluator metrics, benchmark groups, and bounded molecular-oracle workflows.
Take slowlyc/sglang-skill from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.