mcpbeat Sign in

Foreman Debug Agent Skill

Headless root-cause debugging loop for a Foreman worker whose tests, build, or acceptance check are failing — especially on a retry. Find the root cause before changing anything, fix at the source with a regression test, and never thrash on symptom patches. Used inside a foreman-tdd build session; emits no summary of its own.

1k tokens
context cost
the whole folder, loaded on every use
1
files
instructions only
0
copies elsewhere
how many repositories repackaged it
427
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/VisionForge-OU/foreman --skill foreman-debug

The instruction itself

9 sections, as written by the author

foreman-debug

(Adapted from obra/superpowers systematic-debugging (MIT) — see NOTICE. Made

headless: removed the "discuss with your human partner" hand-offs — under Foreman

there is no live human, so a genuine architectural dead end becomes a FOREMAN-SUMMARY

escalate from the surrounding foreman-tdd run, not a question. Folded the regression

step into Foreman's existing foreman-test + evidence contract.)

You are invoked inside a foreman-tdd build session when something is failing: a

red test that should be green, the project test/lint/typecheck command, the issue's

acceptance_check, or a distilled failure report from a prior attempt. You run

headless — do not ask questions. Find the root cause, fix it at the source, and

hand control back to foreman-tdd. You emit no FOREMAN-SUMMARY of your own; the

foreman-tdd run owns the single summary block.

The Iron Law

NO FIX WITHOUT ROOT-CAUSE INVESTIGATION FIRST

A symptom patch that makes the red go away without explaining *why* it was red is a

failure — it will bounce at Foreman's merge gate or resurface on the next slice.

The four phases (complete each before the next)

1. Root cause

  • Read the failure completely. The exact assertion, the stack trace, the file and

line, the ERROR lines in the foreman-test log on disk. The message often *is*

the answer. If a distilled failure report from a prior attempt is in your context,

treat its "why it was rejected" as the starting hypothesis, not noise to re-discover.

  • Reproduce deterministically. Run the single failing test through foreman-test

(use --fast while iterating). If it is flaky, that *is* the bug — chase the

nondeterminism (ordering, time, shared state), don't paper over it.

  • Check what changed. git diff the slice against the integration branch. The

regression almost always lives in the diff.

  • Trace the bad value to its source. Where does the wrong value originate? What

passed it in? Keep walking *up* the call stack until you reach the origin. Fix

there, not at the symptom.

2. Pattern

Find working code that does the same thing elsewhere in the repo. List every

difference between it and the broken path, however small — "that can't matter" is how

root causes hide.

3. Hypothesis

State one specific hypothesis: "the root cause is X because Y." Make the smallest

change that tests it. One variable at a time. If it doesn't hold, form a *new*

hypothesis — do not stack a second fix on top of an unproven first.

4. Fix at the source

  • Lock it with a failing test first. Add (or keep) a test that fails *because of

this root cause* and will pass once it's fixed — exactly the red-green discipline

foreman-tdd already uses. A fix with no test that proves it does not count.

  • One change. Address the root cause only — no "while I'm here" refactors.
  • Verify through foreman-test. The targeted test passes AND the full suite

stays green. Read the output; do not assume.

When 3+ fixes have failed

If three distinct fixes each fail or each surfaces a new problem somewhere else, the

issue is architectural, not a bug — the slice's seam is wrong. Stop patching.

Hand back to foreman-tdd with a clear note for its FOREMAN-SUMMARY: set `escalate:

true` with a one-line statement of the structural problem (e.g. "ISS-012 assumes a

synchronous store but the queue is async — the seam can't hold"). A wrong architecture

is Foreman's human's call, not another guess.

Red flags — stop and return to phase 1

  • "Quick fix now, understand it later."
  • "Just try changing X and see."
  • Bundling several changes, then running tests.
  • Proposing a fix before you traced the bad value to its origin.
  • A fourth fix attempt after three failures (→ escalate the architecture instead).

Other skills for the same job

different authors, same section of the catalogue
Webapp Testing
by anthropics
vendor ×12

Toolkit for interacting with and testing local web applications using Playwright. Supports verifying frontend functionality, debugging UI behavior, capturing browser screenshots, and viewing browser logs.

6k tokens scripts
Finishing A Development Branch
by ZhanlinCui
×7

Use when implementation is complete, all tests pass, and you need to decide how to integrate the work - guides completion of development work by presenting structured options for merge, PR, or cleanup

1k tokens
Test Driven Development
by w95
×7

Use when implementing any feature or bugfix, before writing implementation code

2k tokens
Systematic Debugging
by ratacat
×7

Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes

10k tokens scripts
Verification Before Completion
by ZhanlinCui
×6

Use when about to claim work is complete, fixed, or passing, before committing or creating PRs - requires running verification commands and confirming output before making any success claims; evidence before assertions always

1k tokens
Backtest Expert
by BaggaT236
×3

Expert guidance for systematic backtesting of trading strategies. Use when developing, testing, stress-testing, or validating quantitative trading strategies. Covers "beating ideas to death" methodology, parameter robustness testing, slippage modeling, bias prevention, and interpreting backtest results. Applicable when user asks about backtesting, strategy validation, robustness testing, avoiding overfitting, or systematic trading development.

15k tokens scripts
Adaptyv
by christophacham
×3

Cloud laboratory platform for automated protein testing and validation. Use when designing proteins and needing experimental validation including binding assays, expression testing, thermostability measurements, enzyme activity assays, or protein sequence optimization. Also use for submitting experiments via API, tracking experiment status, downloading results, optimizing protein sequences for better expression using computational tools (NetSolP, SoluProt, SolubleMPNN, ESM), or managing protein design workflows with wet-lab validation.

16k tokens
Aeon
by christophacham
×3

This skill should be used for time series machine learning tasks including classification, regression, clustering, forecasting, anomaly detection, segmentation, and similarity search. Use when working with temporal data, sequential patterns, or time-indexed observations requiring specialized algorithms beyond standard ML approaches. Particularly suited for univariate and multivariate time series analysis with scikit-learn compatible APIs.

19k tokens

How to use it

Copy the folder

Take visionforge-ou/foreman-debug from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.