| Operating SOP for DSPy (Stanford NLP) — the declarative framework for "programming, not prompting" language models. "BootstrapFewShot", "GEPA", "Signatures + Modules", "teleprompter", "auto-tune prompts for a different LM", or whenever a brittle hand-crafted prompt pipeline needs to be turned into a *compiled*, measurable, swappable program. Do NOT activate for one-shot prompt tweaks, no-metric exploratory work, or pipelines where prompts must remain human-authored verbatim — use raw prompting or LangChain templates instead.
npx skills add https://github.com/agentsope/SkillAlchemy --skill agentsop-dspy
> *"DSPy isn't a prompt-optimization agent framework. It's the LLM compiler for the shortest, cleanest code."*
> — Eito Miyamura [eito.substack.com/p/dspy-the-most-misunderstood-agent]
>
> *"Prompts are effectively the weights of an LLM application."*
> — Core philosophy [arxiv.org/abs/2310.03714]
Activate this skill when any of the following triggers are present in the user's intent or codebase:
| Trigger | Signal |
|---|---|
| Imports / mentions | import dspy, dspy.Signature, dspy.ChainOfThought, dspy.ReAct, Predict, MIPROv2, BootstrapFewShot, GEPA, teleprompter, compile( on an LM program |
| Tasks | "auto-tune this prompt", "I want to swap GPT-4 for a smaller model without re-engineering prompts", "I have 50/200/1000 labeled examples — optimize this", "compile a pipeline for our metric", "distill GPT-4 into Llama-3-8B" |
| Symptoms | Hand-written prompts grow past ~50 lines; brittleness on model swap; the team manually tunes few-shot examples; a metric exists but isn't being used to drive prompt design |
| Cross-skill bridges | LangGraph node calls an LLM and needs better prompts → wrap the node body in a DSPy module. LlamaIndex retriever feeds a reranker → DSPy-compile the reranker against a labeled set |
Do NOT activate when:
client.messages.create.DSPy's full name is Declarative Self-improving Python. The three primitives form a PyTorch-like compile chain [arxiv.org/abs/2310.03714]:
┌─────────────┐ ┌──────────┐ ┌──────────────┐ ┌─────────┐
│ Signature │ → │ Module │ → │ Teleprompter │ → │ Compile │
│ (what) │ │ (how) │ │ (optimizer) │ │ (tune) │
└─────────────┘ └──────────┘ └──────────────┘ └─────────┘
I/O spec Predict/CoT/ MIPROv2/GEPA/ Bake demos
field names ReAct/PoT BootstrapFewShot + instructions
= semantic = strategy = search algorithm into JSON
Three mental shifts the agent must internalize:
program.json, not a .txt prompt [dspy.ai/tutorials/saving/].question -> answer is not the same as query -> response. DSPy uses the *field names* as the only natural-language hint the optimizer has about intent before it sees data. Name them like you'd name function parameters in well-written code [dspy.ai/learn/programming/signatures/].num_trials × |trainset| × |program LM calls| [dspy.ai/faqs/].The PyTorch analogy is load-bearing. Signatures ≈ nn.Module.forward() shape contract. Modules ≈ nn.Linear / nn.Transformer. Teleprompters ≈ torch.optim.Adam. compile() ≈ training loop. save()/load() ≈ checkpoint.
The DSPy team is explicit about a three-stage gate [dspy.ai/learn/]:
> "It's unproductive to launch optimization runs using a poorly designed program or a bad metric."
Do not skip stages. Each stage has an exit criterion.
"question -> answer"); upgrade to a class-based dspy.Signature with InputField(desc=...) / OutputField(desc=...) when types matter or fields need disambiguation.dspy.ChainOfThought. Use dspy.Predict for trivial classification, dspy.ReAct only when tools are needed, dspy.ProgramOfThought for arithmetic-heavy tasks [dspy.ai/learn/programming/modules/].dspy.Module, instantiate sub-modules in __init__, call them in forward(). No special DSL.dspy.inspect_history(n=3).Exit criterion: the un-optimized program produces *plausible* outputs on 5+ examples. Not great — plausible.
def metric(example, pred, trace=None) -> float|bool. Start with exact-match; only escalate to LLM-as-judge when the task demands it (open-ended generation, multi-criteria).dspy.Evaluate(devset=dev, metric=metric, num_threads=16) and record a baseline score.Exit criterion: baseline score is stable across two runs (cache-free) AND the metric agrees with human judgment on 10 spot-checks.
10. Start auto="light". Only escalate to "medium"/"heavy" if dev-set gains flatten and budget allows.
11. Save the compiled program: compiled.save("v1.json") for state, or compiled.save("./v1/", save_program=True) for whole-program (preferred for production with metadata) [dspy.ai/tutorials/saving/].
12. Deploy via FastAPI (dspy.asyncify) or MLflow (mlflow.dspy.log_model) [dspy.ai/tutorials/deployment/].
Exit criterion: compiled program beats baseline on a *held-out* test set (not the val set used in optimization) by ≥ task-relevant delta.
Loop to Stage 1 if optimization plateaus. Per the docs: "Is your task well-defined? Do you need more data? Should your evaluation metric change?" — these are the questions to re-ask, not "should I try a different optimizer?" [dspy.ai/learn/optimization/overview/].
| Trigger | Action | Output | Evidence |
|---|---|---|---|
| ≤10 labeled examples | BootstrapFewShot(metric=m, max_bootstrapped_demos=4, max_rounds=1) | Compiled program with self-generated demos | [dspy.ai/learn/optimization/optimizers/] |
| 30–50 examples | BootstrapFewShotWithRandomSearch | Best-of-N candidate programs | [dspy.ai/learn/optimization/optimizers/] |
| 200+ examples, willing to spend compute | MIPROv2(metric=m, auto="light") then escalate | Jointly-tuned instructions + few-shot demos via Bayesian optimization | [dspy.ai/api/optimizers/MIPROv2/] |
| Need zero-shot prompts (no demos in final) | MIPROv2(..., max_bootstrapped_demos=0, max_labeled_demos=0) | Instruction-only optimization | [dspy.ai/learn/optimization/optimizers/] |
| Have textual error feedback (test diffs, schema violations, judge rationales) | dspy.GEPA(metric=m_with_feedback) | Reflection-evolved prompts; sample-efficient | [dspy.ai/tutorials/gepa_ai_program/], [arxiv.org/abs/2507.19457] |
| Already optimized with MIPROv2 / want to ship a smaller model | Chain into BootstrapFinetune(student=small_lm, teacher=optimized) | Finetuned weights (not just prompts) | [dspy.ai/api/optimizers/BootstrapFinetune/] |
| Just want labeled demos in prompt (no search) | LabeledFewShot(k=8) | Trivial — fastest, cheapest, weakest | [dspy.ai/cheatsheet/] |
| Trigger | Action | Why |
|---|---|---|
| Simple input → output | dspy.Predict(Sig) | Lowest overhead |
| Reasoning helps | dspy.ChainOfThought(Sig) | Default choice per docs |
| Math / counting / parsing | dspy.ProgramOfThought(Sig) | Code execution grounds the answer |
| Tools (search, calc, API) | dspy.ReAct(Sig, tools=[...]) | Built-in tool loop |
| Ensemble for hard cases | dspy.MultiChainComparison or dspy.majority | Vote across N CoT samples |
| Trigger | Action | Caveat |
|---|---|---|
| Exact answer expected | lambda ex, pred: ex.answer.lower() == pred.answer.lower() | Cheap, deterministic |
| Open-ended generation | LLM-as-judge with dspy.ChainOfThought(JudgeSig) | Watch for self-preference bias, recency bias, score-ID bias [arxiv.org/pdf/2509.26072] |
| Multi-criteria (factuality + tone + length) | Sub-judge each dim, return bool during optimization (trace is not None) and float during evaluation | Documented pattern [dspy.ai/learn/evaluation/metrics/] |
| Have rich error context | Return dspy.Prediction(score=..., feedback="missing field X") and use GEPA | Textual feedback is GEPA's superpower [dspy.ai/api/optimizers/GEPA/overview/] |
| Trigger | Action | Reference |
|---|---|---|
| Before any MIPROv2 call | Estimate: auto="light" ≈ a few $; auto="heavy" on 1000+ examples can hit tens of $ | [dspy.ai/faqs/] |
| Budget tight | Use a cheap optimizer LM (e.g. gpt-4o-mini) to optimize prompts for a more expensive task LM — community-reported parity [github.com/stanfordnlp/dspy/issues/1596] |
| Compile stuck mid-trial | Check issue #1970 pattern; reduce minibatch_size or kill and restart with smaller num_trials |
| Need reproducibility | dspy.configure(track_usage=True) + log program.get_lm_usage() |
困境 (Dilemma): User has a 3-stage RAG pipeline. Hand-tuned prompts already hit 72% on dev. MIPROv2 auto="heavy" would cost ~$40 and 4 hours. Worth it?
约束 (Constraints):
决策步骤 (Decision steps):
auto="light" (~$2) typically yields 10–30%+ on hand-tuned baselines per the paper's GPT-3.5/Llama2 results (25%/65% lift over standard few-shot) [arxiv.org/abs/2310.03714].auto="light" first as a cheap signal. The docs explicitly recommend "start with moderate values, observe behavior, and scale up only if you see clear gains" [github.com/stanfordnlp/dspy issue #1596].light gives <2% lift, do not escalate to heavy. Instead, revisit Stage 1: is the signature ambiguous? Is the program structure (3 stages) actually right?light gives 5–10% lift, run medium. Only escalate to heavy if data ≥ 300 *and* you have a held-out test set distinct from val.结果 (Outcome): Typical: light exposes whether more compute helps. Often the answer is "no — fix the program/metric first."
可提取的操作 (Extractable operation): Never start compilation at auto="heavy". Always probe with light and use a cheap optimizer LM.
困境: Compiled program for GPT-4o works at 85%. Need to switch to Llama-3-8B for cost. Re-use the GPT-4o-compiled program.json or recompile?
约束:
决策步骤:
BootstrapFinetune as a follow-on: optimize prompts on the big model, then distill into a 1B–7B student. Typical setup: student=Llama-3.2-1B-Instruct, teacher=gpt-4o-mini [dspy.ai/api/optimizers/BootstrapFinetune/].program.gpt4o.json and program.llama8b.json checked in; A/B in production.结果: Recompiled programs typically recover 70–90% of the larger-model performance at 1/10–1/50 the per-call cost. The "transfer without recompile" path is reliably worse.
可提取的操作: Treat the compiled program as a (program × LM) pair. Changing the LM invalidates the artifact — recompile.
困境: Open-ended customer-support response task. No exact-match metric possible. LLM-as-judge "feels right" but the team worries the judge will be biased toward verbose, hedged outputs.
约束:
决策步骤:
dspy.Predict(Assess) call with a single yes/no question (factual? on-topic? concise? non-hedging?). Documented pattern [dspy.ai/learn/evaluation/metrics/].trace is not None to return bool during compile, float during eval — same metric function, two modes. Avoids the optimizer overfitting to score noise.dspy.GEPA instead of MIPROv2 — GEPA leverages text feedback for faster, more sample-efficient convergence [dspy.ai/api/optimizers/GEPA/overview/, arxiv.org/abs/2507.19457].结果: Multi-dimension metric with explicit length penalty + GEPA's textual feedback typically beats single-judge + MIPROv2 by 10–13% on AIME-style benchmarks [arxiv.org/abs/2507.19457] and is the empirically robust path.
可提取的操作: Never compile against a metric you haven't human-validated on ≥ 20 spot-checks. Decompose multi-criteria metrics. Prefer GEPA when you can express textual feedback.
困境: MIPROv2 compile stuck mid-trial (no progress logs for 30 min). Reported pattern in issue #1970 [github.com/stanfordnlp/dspy/issues/1970]. Abort and restart, or wait?
约束:
决策步骤:
dspy.inspect_history(n=3) — does the last LM call show truncation or rate-limit error?max_bootstrapped_demos and max_labeled_demos (default 4 each); the docs explicitly cite this as the #1 context-length fix [dspy.ai/faqs/].num_threads in the underlying Evaluate; add retry/backoff in the LM client.minibatch_size (default 35; try 16) and smaller num_trials. Hanging is a known failure mode without graceful resume.student retains best demo candidates — check compiled._predictors state.可提取的操作: Compile is not atomic. Treat long hangs as failure. The cost of restart < cost of indefinite wait.
auto="heavy". Always probe with light first [Case A].10. Forgetting cache=False in Lambda / stateless deploys. Caches default to a writable dir and break in serverless [dspy.ai/faqs/].
dspy.streamify from 2.6.0+ but it's newer than the rest of the stack — verify your version [dspy.ai/tutorials/deployment/].DSPy is not the same layer as LangChain / LlamaIndex / LangGraph. It sits *underneath* them as a compiler for the individual LM calls inside those orchestration layers [langwatch.ai/blog/best-ai-agent-frameworks-in-2025-...].
┌──────────────────────────────────────────────┐
│ Orchestration: LangGraph, CrewAI │ ← graphs, agents, state
├──────────────────────────────────────────────┤
│ Retrieval: LlamaIndex │ ← ingestion, indexing
├──────────────────────────────────────────────┤
│ Compiler: DSPy │ ← signatures, modules, compile
├──────────────────────────────────────────────┤
│ Generation: Guidance, LMQL, Outlines │ ← single-call grammar control
├──────────────────────────────────────────────┤
│ Inference: vLLM, llama.cpp, Anthropic │ ← serving
└──────────────────────────────────────────────┘
Predict(context, question -> answer) synthesizes. JetBlue's chatbot uses exactly this split — retrieval quality + answer quality as separate metrics, DSPy optimizes both [databricks.com/blog/optimizing-databricks-llm-pipelines-dspy].OutputField already pushes the LM toward structure but doesn't guarantee grammar conformance [dspy.ai/faqs/].import dspy
# 1. Signature
class BasicQA(dspy.Signature):
"""Answer questions with short factoid answers."""
question: str = dspy.InputField()
answer: str = dspy.OutputField(desc="often between 1 and 5 words")
# 2. Module
qa = dspy.ChainOfThought(BasicQA)
# 3. Metric
def metric(ex, pred, trace=None):
return ex.answer.lower() in pred.answer.lower()
# 4. Compile
from dspy.teleprompt import MIPROv2
optimizer = MIPROv2(metric=metric, auto="light")
compiled = optimizer.compile(qa, trainset=trainset)
# 5. Save / load
compiled.save("v1.json")
Have a metric? ─── No ──► Stop. Build a metric first. (Or skip DSPy.)
│
Yes
│
Have ≥ 30 examples? ─── No ──► Stop. Collect more data, or use LabeledFewShot(k=8) as floor.
│
Yes
│
Have textual error feedback? ─── Yes ──► dspy.GEPA
│
No
│
≤ 10 examples? ──► BootstrapFewShot
30–50? ──► BootstrapFewShotWithRandomSearch
50–200? ──► MIPROv2(auto="light", max_bootstrapped_demos=4)
200+? ──► MIPROv2(auto="light" → "medium" if gains; "heavy" only if 300+ and budget)
Need to ship small model? ──► chain BootstrapFinetune after MIPROv2
program.jsonAfter compiled.save("v1.json"), the file is plain JSON. Per-predictor it contains [dspy.ai/tutorials/saving/]:
{
"predictor_name": {
"signature_instructions": "Given the context, answer the question with a short factoid...",
"signature_prefix": "Answer:",
"extended_signature_instructions": "...",
"demos": [
{"question": "...", "reasoning": "...", "answer": "..."},
...
],
"signature": {
"instructions": "...",
"fields": [{"prefix": "Question:", "description": "..."}, ...]
}
}
}
What changes when you compile:
What does NOT change between LMs (so you can read across artifacts):
What DOES change between LMs (so you can't reuse):
dspy.Assert vs dspy.SuggestFor self-refining pipelines [dspy.ai/learn/programming/7-assertions/, arxiv.org/pdf/2312.13382]:
# Hard: halts after max retries with dspy.AssertionError
dspy.Assert(len(pred.answer) < 100, "Answer must be < 100 chars")
# Soft: retries with feedback in prompt, logs failure, continues
dspy.Suggest(is_valid_json(pred.output), "Output must be valid JSON")
When a constraint fails, DSPy backtracks to the previous module and re-runs with the error message injected into the prompt. This is *self-refinement at inference time* — distinct from compile-time optimization.
Use Assert during development (catch bugs hard). Use Suggest in production (degrade gracefully).
This skill should be used when the user asks to "create a hook", "add a PreToolUse/PostToolUse/Stop hook", "validate tool use", "implement prompt-based hooks", "use ${CLAUDE_PLUGIN_ROOT}", "set up event-driven automation", "block dangerous commands", or mentions hook events (PreToolUse, PostToolUse, Stop, SubagentStop, SessionStart, SessionEnd, UserPromptSubmit, PreCompact, Notification). Provides comprehensive guidance for creating and implementing Claude Code plugin hooks with focus on advanced prompt-based hooks API.
This skill should be used when the user asks to "create a hook", "add a PreToolUse/PostToolUse/Stop hook", "validate tool use", "implement prompt-based hooks", "use ${CLAUDE_PLUGIN_ROOT}", "set up event-driven automation", "block dangerous commands", or mentions hook events (PreToolUse, PostToolUse, Stop, SubagentStop, SessionStart, SessionEnd, UserPromptSubmit, PreCompact, Notification). Provides comprehensive guidance for creating and implementing Claude Code plugin hooks with focus on advanced prompt-based hooks API.
Build agentic applications with GitHub Copilot SDK. Use when embedding AI agents in apps, creating custom tools, implementing streaming responses, managing sessions, connecting to MCP servers, or creating custom agents. Triggers on Copilot SDK, GitHub SDK, agentic app, embed Copilot, programmable agent, MCP server, custom agent.
Coding Agent Session Search - unified CLI/TUI to index and search local coding agent history from Claude Code, Codex, Gemini, Cursor, Aider, ChatGPT, Pi-Agent, Factory, and more. Purpose-built for AI agent consumption with robot mode.
Destructive Command Guard - High-performance Rust hook for Claude Code that blocks dangerous commands before execution. SIMD-accelerated, modular pack system, whitelist-first architecture. Essential safety layer for agent workflows.
Makepad UI development skills for Rust apps: setup, patterns, shaders, packaging, and troubleshooting.
Secure environment variable management ensuring secrets are never exposed in Claude sessions, terminals, logs, or git commits
Prompt for generating an AGENTS.md file for a repository
Take agentsope/agentsop-dspy from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.