> Estimate the dollars saved by routing eligible Claude Code work to a cheaper model family, using the Agent Monitor pricing engine. Re-prices each model's token mix at the target family's rates and quantifies the delta. Uses /api/pricing (rates), /api/pricing/cost (current per-model spend), /api/sessions, and /api/analytics. Use when hunting for cost cuts or comparing model tiers.
npx skills add https://github.com/hoangsonww/Claude-Code-Agent-Monitor --skill model-savings
Quantify how much spend you would recover by moving eligible work to a cheaper model.
The user provides: $ARGUMENTS
This is the routing question — e.g. "Opus → Sonnet", "move simple work to Haiku",
or empty (analyze every premium model against the next tier down). If no target family
is named, default to proposing the next-cheaper tier per model and say so.
| Endpoint | Returns |
|----------|---------|
| GET /api/pricing | { pricing: [{ model_pattern, display_name, input_per_mtok, output_per_mtok, cache_read_per_mtok, cache_write_per_mtok }] } — the rate card for every family |
| GET /api/pricing/cost | { total_cost, breakdown: [{ model, input_tokens, output_tokens, cache_read_tokens, cache_write_tokens, cost, matched_rule }] } — current spend and the exact token mix per model |
| GET /api/sessions?limit=200 | Sessions with model, inline cost, and metadata (turn_count, thinking_blocks) — used to judge which work is *eligible* to downshift |
| GET /api/analytics | agent_types, tool_usage, total_subagents — corroborate which task types are low-complexity and safe to route cheaper |
For each candidate model in the cost breakdown, re-price its exact token mix at the target family's rates:
cost_at_target = (input_tokens / 1M) × target.input_per_mtok
+ (output_tokens / 1M) × target.output_per_mtok
+ (cache_read_tokens / 1M) × target.cache_read_per_mtok
+ (cache_write_tokens/ 1M) × target.cache_write_per_mtok
savings = current_model_cost − cost_at_target
Pull target.*_per_mtok from /api/pricing (longest model_pattern match wins). Default rates ($/Mtok in/out/cacheRead/cacheWrite): Opus $5/$25/$0.50/$6.25, Sonnet $3/$15/$0.30/$3.75, Haiku $1/$5/$0.10/$1.25.
Re-pricing the full token mix is the *theoretical ceiling*. Scope it to eligible work:
metadata.turn_count small) and simple subagent/tool work are safe to downshift.Table from /api/pricing/cost: each model, its 4 token counts, and current cost. Note its share of total_cost.
For each candidate, show cost_at_target and savings (absolute $ and %). Make the target rate card explicit.
Apply the eligibility rule and recompute savings over just the downshiftable token mix. Show how many sessions / what share of tokens qualified.
Rank routing moves by eligible monthly savings (descending), top 5. For each: source → target, the token mix moved, estimated $ saved, and a confidence level (high/medium/low) based on how clearly the work is low-complexity.
Cheaper models may need more turns or produce more output — note that realized savings can be lower than the static re-price, and that quality-sensitive work should stay on the premium tier.
Markdown tables. Currency as USD to 4 decimal places; token counts with thousands separators; rates as $/Mtok. Always present both the ceiling (full re-price) and the eligible-only estimate so the number is honest.
Create new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, edit, or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy.
Guide for creating effective skills. This skill should be used when users want to create a new skill (or update an existing skill) that extends Claude's capabilities with specialized knowledge, workflows, or tool integrations.
Guide for creating effective skills. This skill should be used when users want to create a new skill (or update an existing skill) that extends Claude's capabilities with specialized knowledge, workflows, or tool integrations.
Replace with description of the skill and when Claude should use it.
Use when facing 2+ independent tasks that can be worked on without shared state or sequential dependencies
This skill should be used when the user wants to "create a skill", "add a skill to plugin", "write a new skill", "improve skill description", "organize skill content", or needs guidance on skill structure, progressive disclosure, or skill development best practices for Claude Code plugins.
Helps users discover and install agent skills when they ask questions like "how do I do X", "find a skill for X", "is there a skill that can...", or express interest in extending capabilities. This skill should be used when the user is looking for functionality that might exist as an installable skill.
Use when creating new skills, editing existing skills, or verifying skills work before deployment
Take hoangsonww/model-savings from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.