omidzamani/dspy-production-deployment
Use for deploying DSPy with save/load, configure_cache, restrict_pickle, track_usage, async execution, streaming, and production runtime controls.
npx skills add https://github.com/OmidZamani/dspy-skills --skill dspy-production-deployment
Prepare a DSPy program for repeatable, observable, scalable, and safer production execution.
DSPy enables memory and disk caches by default. Disk cache deserialization uses pickle unless restricted. Enable the allowlist mode in production:
import dspy
dspy.configure_cache(restrict_pickle=True)
Register trusted custom cache types only when needed:
dspy.configure_cache(
restrict_pickle=True,
safe_types=[MyResult, Metadata],
)
Disable a cache layer explicitly when a deployment cannot persist data or requires fresh model responses:
dspy.configure_cache(
enable_disk_cache=False,
enable_memory_cache=True,
)
Prefer state-only JSON for readable, safer artifacts:
compiled.save("./artifacts/program.json", save_program=False)
loaded = MyProgram()
loaded.load("./artifacts/program.json")
Use whole-program save only for trusted artifacts. It uses cloudpickle:
compiled.save("./artifacts/program/", save_program=True)
loaded = dspy.load("./artifacts/program/")
Keep the DSPy major version compatible when loading saved programs.
dspy.configure(
lm=dspy.LM("openai/gpt-4o-mini"),
track_usage=True,
)
prediction = program(question="What is DSPy?")
print(prediction.get_lm_usage())
Cached calls return no new token usage.
Most built-in modules support acall():
import asyncio
async def main():
prediction = await program.acall(question="What is DSPy?")
print(prediction.answer)
asyncio.run(main())
Implement aforward() for custom async modules. Use dspy.asyncify(program) only when adapting a synchronous callable is the right boundary.
import asyncio
import dspy
stream_program = dspy.streamify(
dspy.Predict("question -> answer"),
stream_listeners=[
dspy.streaming.StreamListener(signature_field_name="answer"),
],
)
async def main():
async for chunk in stream_program(question="Explain DSPy briefly."):
print(chunk)
asyncio.run(main())
For looped modules such as ReAct, set allow_reuse=True on listeners for repeated fields. Cache hits yield the final Prediction without replaying token chunks.
restrict_pickle=True.Take omidzamani/dspy-production-deployment from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.