Use when the user wants their Claude agent to self-improve from past usage, asks about a nightly/offline 'sleep' or 'dream' cycle, memory/skill consolidation, or says things like 'make my agent better the more I use it', 'review my past sessions', 'learn my preferences', 'consolidate what you learned', 'run the sleep cycle', or wants to schedule background self-optimization. Drives the skillopt_sleep engine: harvest past sessions -> mine recurring tasks -> replay through a selected backend -> consolidate validated CLAUDE.md/SKILL.md behind a held-out gate.
npx skills add https://github.com/microsoft/SkillOpt --skill skillopt-sleep
SkillOpt-Sleep gives the user's agent a sleep cycle. On demand or on a
nightly schedule, it reviews real past Claude Code sessions, re-runs recurring
tasks through the selected backend, and consolidates what it
learns into memory (CLAUDE.md) and skills (SKILL.md). With the
default validation gate enabled, it keeps only changes that improve a held-out
score. Live files change only through explicit adoption or a user-requested
--auto-adopt. It aims to improve this user's recurring work, while making
each accepted proposal measurable on the run's held-out tasks,
with no model-weight training. It is the deployment-time analogue of training:
short-term experience → long-term competence.
It synthesizes three ideas:
edits; accepted only through a held-out gate; rejected edits are recorded in
the run report for review.
inside protected learned blocks; the input is never mutated, and output is
reviewed before adoption.
Trigger when the user wants any of:
CLAUDE.md or a managed skill~/.claude/projects/*/<session>.jsonl + ~/.claude/history.jsonl (READ-ONLY) → session digests.TaskRecords (recurring intents + outcome labels + checkable refs where possible).skill+memory → (hard, soft) scores.
proposed_CLAUDE.md and/orproposed_SKILL.md, plus report.md, report.json, manifest.json, and
diagnostics.json into <project>/.skillopt-sleep/staging/<timestamp>/.
Nothing live changes. A rejected run still has a report but no proposed
live-file replacement.
Prefer the /skillopt-sleep command. Under the hood it calls the bundled runner:
"${CLAUDE_PLUGIN_ROOT}/scripts/sleep.sh" status # what's happened
"${CLAUDE_PLUGIN_ROOT}/scripts/sleep.sh" dry-run --project "$(pwd)" # no-staging preview
"${CLAUDE_PLUGIN_ROOT}/scripts/sleep.sh" run --project "$(pwd)" # full cycle, stages a proposal
"${CLAUDE_PLUGIN_ROOT}/scripts/sleep.sh" adopt --project "$(pwd)" # apply staged proposal (with backup)
mock (deterministic, no API spend) — good for trying the plumbing.--backend claude or --backend codex to spend the user's real budgetfor model-driven optimization. A held-out gain is run-specific evidence, not
a guarantee of broader improvement; results depend on the tasks, model, and
checks.
--scope all harvests every Claudeproject into the current run's configured targets.
the data-boundary rules below before using one with sensitive sessions.
"${CLAUDE_PLUGIN_ROOT}/scripts/sleep.sh" schedule --project "$(pwd)" --hour 3 --minute 17
"${CLAUDE_PLUGIN_ROOT}/scripts/sleep.sh" unschedule --project "$(pwd)"
Installs a nightly cron entry. unschedule --all removes every managed entry.
| Flag | Default | Description |
|------|---------|-------------|
| --project PATH | cwd | Project directory to evolve |
| --scope all\|invoked | invoked | Harvest scope |
| --backend mock\|claude\|codex\|copilot\|handoff\|azure_openai | mock | Backend (mock = no provider calls) |
| --model NAME | backend default | Override the model used for replay |
| --source claude\|codex\|auto | claude | Transcript source |
| --lookback-hours N | 72 | Harvest window |
| --max-sessions N | derived | Cap harvested sessions; defaults to 3 × max tasks (120 with current defaults) |
| --max-tasks N | 40 | Cap mined tasks |
| --target-skill-path PATH | ~/.claude/skills/skillopt-sleep-learned/SKILL.md | Explicit SKILL.md to evolve |
| --tasks-file PATH | — | Reviewed TaskRecord JSON (skip harvest) |
| --progress | off | Print phase progress to stderr |
| --auto-adopt | off | Auto-adopt if gate passes |
| --edit-budget N | 4 | Max bounded edits per night |
| --preferences TEXT | empty | Add house rules to the optimizer's reflection prior |
| --json | off | Machine-readable JSON output |
The CLI also has source/runtime path overrides (--claude-home, --codex-home,
and --codex-path) and action-specific flags. Use
python -m skillopt_sleep <action> --help as the authoritative surface.
~/.skillopt-sleep/config.json)Beyond the CLI flags, advanced behavior is controlled via config:
preferences — free-text house rules injected into the optimizer's reflect step (e.g. "Always use async/await", "Answers in \boxed{}").gate_mode — on (default, validation-gated) or off (greedy, accept all edits).gate_metric — hard, soft, or mixed (default). Controls how the held-out gate scores.dream_rollouts — >1 enables multi-rollout contrastive reflection per task.recall_k — >0 recalls K similar past tasks into the dream (long-term memory).evolve_memory / evolve_skill — independently toggle CLAUDE.md vs SKILL.md consolidation.The sleep cycle can consolidate both:
With the default gate enabled, both are evaluated by the same held-out score.
Set evolve_memory: false to consolidate only skills, or evolve_skill: false
for only memory.
CLAUDE.md / SKILL.md as part of this skill.Let the engine's explicit adopt or user-requested --auto-adopt path apply
the staging manifest and back up existing live files first.
mock replay has no side effects.selected provider for mining, replay, judging, and reflection. The Claude
transcript path is not guaranteed to remove every secret before those calls.
Review provider policy and session contents first. For sensitive data, use
mock or run harvest --output <file>, inspect/redact the JSON, set
"reviewed": true, and replay it with --tasks-file; real backends refuse an
unreviewed task file.
exact proposed edits before suggesting adoption. Evidence before adoption.
python -m skillopt_sleep.experiments.run_experiment --persona researcher --json
— a deterministic synthetic demo of held-out lift and gate rejection. It
validates the mechanism, not effectiveness on the user's own tasks.
# deterministic synthetic demo (no API): score rises and the gate blocks a regression
python -m skillopt_sleep.experiments.run_experiment --persona researcher --assert-improves
python -m skillopt_sleep.experiments.run_experiment --persona programmer --assert-improves
See the SkillOpt-Sleep documentation
for recorded results, limitations, and the supported integration surface.
Complete development kit for Microsoft 365 Copilot declarative agents with three comprehensive workflows (basic, advanced, validation), TypeSpec support, and Microsoft 365 Agents Toolkit integration
Format and structurally validate local treatment-plan documentation after clinical decisions have already been supplied and verified by authorized licensed professionals. Use for source traceability, clinician-authored intervention records, goals and checkpoints, shared-decision records, reconciliation handoffs, and release gates—not for clinical decision-making.
> provider/change budget/修改卖家/修改预算/draft/草稿/我的任务/my tasks/what am I working on/关闭/取消任务/决策列表/decision list/指定服务商/browse (sender.role = COUNTERPARTY, not you); (3) literal "Read the okx-ai skill" (or legacy "Read the okx-agent-task skill") in the envelope.
Automate payer review of prior authorization (PA) requests. This skill should be used when users say "Review this PA request", "Process prior authorization for [procedure]", "Assess medical necessity", "Generate PA decision", or when processing clinical documentation for coverage policy validation and authorization decisions.
Expert in designing and building autonomous AI agents. Masters tool use, memory systems, planning strategies, and multi-agent orchestration.
Autonomous agents are AI systems that can independently decompose goals, plan actions, execute tools, and self-correct without constant human guidance. The challenge isn't making them capable - it's making them reliable. Every extra decision multiplies failure probability.
Orchestrates design workflows by routing work through brainstorming, multi-agent review, and execution readiness in the correct order.
Structured persuasion for tech leads, PMs, and founders—not activity logs. Five scenarios (kickoff, status update, wrap-up, investor pitch, solution selling) on one 5-part framework (Hook→Context→Proposal→Evidence→Ask). AI prompts for missing materials and audience context; pre-submit checklist. Claude Code plugin; Cursor, Codex, and chat via prompts.
Take microsoft/skillopt-sleep from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.