athola/graduated-implementation
Ramps implementation ambition a notch only after the prior increment is understood. Use when building a feature you must understand, not just ship.
npx skills add https://github.com/athola/claude-night-market --skill graduated-implementation
> Start with the smallest slice you can fully understand. Earn the
> next notch by proving you understood the last one. Ambition that
> outruns understanding is how a fluent diff becomes an unverifiable
> one.
The sibling skill imbue:assisted-mastery fades scaffolding as
competence grows. This skill ramps the other axis: the ambition of
the next increment. They are the two directions of one move, the
graduated practice that turned novices into experts long before
agents existed. Not "ban the tool," but "couple the next challenge
to demonstrated competence on the last one."
The learning sciences give the move a number. Wilson et al. (2019,
*Nature Communications* 10:4646) derive the optimal training point
for a learner at roughly 85% success: hard enough to learn from,
not so hard that the signal is noise. The same band is what
Vygotsky's zone of proximal development, Ericsson's edge of
ability, and Csikszentmihalyi's flow channel all gesture at.
Bloom's mastery learning (advance a unit at >=90% on a fresh
check), Bayesian Knowledge Tracing (advance at p(mastery) >= 0.95),
and competence-based curriculum learning (Platanios et al. 2019,
only attempt tasks within the current competence) are the same
rule at different resolutions.
The danger this guards against is specific. An agent that one-shots
a large change is maximally helpful to throughput and quietly
corrosive to verification: you cannot review what you did not watch
get built, and automation bias means you will trust it precisely
when it is wrong (Perry et al. 2023). Aviation named the endpoint
"children of the magenta": ramp the operator's autonomy faster than
their retained understanding and they can no longer hand-fly or
override the automation when it misbehaves.
Do not design the whole system up front. Pick the smallest slice
that is a real, end-to-end step and stop there. The default rung is
about 40 added lines: a change a human can read and explain in one
sitting. The bound is the point, not a nuisance: it keeps
understanding in pace with output. The guard_scope_ramp.py hook
makes this concrete by flagging an increment that jumps past the
current rung.
The next increment may be more ambitious only after the prior one's
understanding is demonstrated and recorded. The check is sized to
blast radius, the advancement gate:
slice has green tests and a recorded tradeoff (what was chosen,
what was rejected, why).
crypto): ramp only when the human explains the prior diff
unaided. This is the magenta hand-fly check. If they cannot
explain it, the rung drops rather than rises.
Recording the demonstration mints a ramp token (`touch
.imbue/ramp-ok`), which the hook consumes to widen the rung one
notch. You ramp by proving you understood the last slice, not by
writing more. Each notch is appended to the
ramp ledger so a reviewer can later audit
that the demonstration was real, not rubber-stamped.
Advancing too fast is one failure; never advancing is the other.
shrink the increment, re-scaffold. Do not ramp.
notch.
faster. Drilling a mastered skill is over-practice, the boredom
failure that gets spaced-repetition decks abandoned (Cen &
Koedinger 2007).
the human will maintain or be accountable for it.
throwaway.
Skip it for a single bounded edit, a trivial reversible change, or
generated and vendored code. Forcing a ramp ritual on a typo fix is
ceremony, and ceremony trains people to ignore the gate.
imbue:proof-of-work)
| Thought | Reality |
|---------|---------|
| "I'll just build the whole thing, then review" | You cannot review what you did not watch get built. Start with one slice. |
| "Tests pass, so it is understood" | Completion is not understanding. Duolingo streaks prove a cheap signal decouples from skill. |
| "I can self-certify I get it" | The producer may not grade its own readiness. Demonstrate it, record it. |
| "Bigger increments are faster" | Faster to write, slower to verify, and the verification is the point. |
| "The rung is slowing me down" | On work you must own, staying in the 85% band is the fast path to durable skill. |
imbue:assisted-mastery: fades scaffolding as competence grows;this skill ramps challenge. Two directions, one axis.
imbue:proof-of-work: the evidence half of the low-stakes gate.imbue:scope-guard: bounds the branch; this bounds theincrement within it.
leyline:risk-classification: the stakes tier that selects whichgate (evidence vs explanation) applies.
leyline:decision-journal: the durable home for the recordedtradeoff that mints a ramp token.
The empirical basis for the 85% band, the failure modes, and the
cross-domain gate design is preserved in
research-basis.md.
start rung, not the whole design.
recorded demonstration of the prior increment (a tradeoff
entry; for high-stakes paths, the human explaining the diff
unaided).
explanation gate, not defaulted silently.
triggered a hold and a smaller next slice, not a ramp.
Take athola/graduated-implementation from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.