mcpbeat Sign in

Harness Engineering Agent Skill

Adopt repository-level harness engineering for coding agents. Use when a user wants to prevent repeated AI coding-agent mistakes by turning failures into durable instructions, drift checks, regression tests, failure memory, and adoption reports tailored to the target repository.

2k tokens
context cost
the whole folder, loaded on every use
1
files
instructions only
0
copies elsewhere
how many repositories repackaged it
37394
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/github/awesome-copilot --skill harness-engineering

The instruction itself

13 sections, as written by the author

Harness Engineering

Harness engineering turns repeated coding-agent mistakes into durable

repository artifacts:

Harness = Instructions + Constraints + Feedback + Memory + Evaluation + Governance

Use this skill when the user asks to:

  • make a repository more reliable for GitHub Copilot or other coding agents
  • add durable agent instructions, repository rules, or guardrails
  • prevent repeated AI coding-agent mistakes
  • record known failure paths and the checks that prevent recurrence
  • add lightweight drift checks for project rules
  • review, refresh, or update an existing agent harness

Do not use this skill for ordinary feature implementation unless the user asks

to improve the repository's agent operating environment.

Core Principles

  • Treat the target repository as the source of truth.
  • Inspect before editing. Preserve the existing stack, package manager, CI,

docs, naming, and architecture.

  • Add the smallest useful harness. Prefer updating existing files over adding

duplicate guidance.

  • Make important rules enforceable where practical through tests, linters,

type checks, CI, pre-commit hooks, or drift scripts.

  • Use manual review points only when automation would be brittle or misleading.
  • Record high-risk failures that should not recur, and name the check or review

point that catches recurrence.

  • Do not copy generic templates blindly. Adapt every artifact to real evidence

in the target repository.

Discovery

Before proposing or making harness changes, inspect the repository for existing

rules and evidence.

Read these files and folders when they exist:

  • README.md
  • AGENTS.md
  • .github/copilot-instructions.md
  • .github/instructions/
  • .github/workflows/
  • CONTRIBUTING.md
  • package manifests such as package.json, pyproject.toml, go.mod,

Cargo.toml, pom.xml, or build.gradle

  • existing docs under docs/
  • existing scripts under scripts/
  • existing tests and CI checks

Then summarize:

  • stack, package manager, and entry points
  • existing development and verification commands
  • current agent instructions or repository conventions
  • known failures, incidents, flaky paths, or repeated review comments
  • gaps where project rules are not enforced

Adoption Workflow

Follow this sequence:

  • Choose the harness surface that fits the target repository.
  • Write target-specific agent instructions.
  • Add enforceable checks for high-value rules.
  • Record failure memory for high-risk or recurring failures.
  • Add drift checks for guidance that can silently become stale.
  • Report the adoption with evidence, assumptions, and follow-up.

1. Choose the Harness Surface

Pick only the surfaces that fit the target repository:

| Need | Preferred artifact |

| --- | --- |

| Always-on agent behavior | AGENTS.md or .github/copilot-instructions.md |

| File-scoped guidance | .github/instructions/*.instructions.md |

| Recurring project checks | scripts/check_*.py, shell scripts, or package scripts |

| CI enforcement | existing workflow files or a small new workflow |

| Known failures | docs/failures/*.md |

| Architecture or process decisions | docs/decisions/*.md |

| Adoption evidence | docs/harness/adoption-report.md or similar |

If the repository already has an equivalent location, update it instead of

creating a parallel system.

2. Write Agent Instructions

Agent instructions should be concrete and operational. Include:

  • project purpose and major ownership boundaries
  • setup, test, lint, build, and verification commands
  • package manager and dependency rules
  • safe editing rules, generated file rules, and forbidden paths
  • testing expectations for changed code
  • PR and commit conventions if the repo has them
  • how to record new failures or decisions

Avoid broad personality guidance, generic best practices, and rules that cannot

be checked or reviewed.

3. Add Enforceable Checks

Convert high-value rules into checks. Good harness checks are:

  • narrow enough to avoid false positives
  • fast enough to run locally and in CI
  • named clearly so agents can run them before finishing
  • documented with the rule they protect

Examples:

Rule: Do not edit generated API clients.
Check: script scans diffs for generated paths and fails with a clear message.

Rule: Every failure memory note names a regression check.
Check: script validates docs/failures/*.md for a "Detection" section.

Rule: Profile docs and templates must stay aligned.
Check: test compares profile README files to expected template files.

4. Record Failure Memory

Record failures when they are user-visible, high-risk, or likely to recur.

Use a new file under docs/failures/ unless an existing note already covers

the same root cause.

Recommended structure:

# Short Failure Title

## Summary

What failed, who saw it, and why it matters.

## Root Cause

The technical or process cause. Avoid blame.

## Prevention

Instruction, test, drift check, CI gate, fixture, or manual review point that
prevents or detects recurrence.

## Evidence

Links to issue, PR, test, log, command output, or file paths.

If no automated check is practical, record the manual review point and why

automation would be unsafe or misleading.

5. Add Drift Checks

Use drift checks for guidance that can silently become stale. Common examples:

  • docs mention commands that no longer exist
  • profile snippets and generated examples diverge
  • failure notes omit regression checks
  • decision records are missing for structural changes
  • CI references stale scripts or package commands

Prefer small scripts using the repository's existing language. If the repo has

no scripting convention, Python with only the standard library is a portable

default.

6. Report the Adoption

Finish substantial harness work with an adoption report that includes:

  • files changed
  • rules added or updated
  • checks added or reused
  • commands run and results
  • assumptions and manual follow-up
  • failure memory created or intentionally skipped
  • how effectiveness will be measured

Review Workflow

When asked to review a harness change, take an opposing perspective. Look for:

  • generic rules copied without evidence from the target repository
  • duplicate or conflicting instruction files
  • broad checks that are likely to fail on valid changes
  • unenforced high-risk rules
  • missing failure memory for repeated mistakes or runtime failures
  • generated docs not refreshed after source changes
  • CI gates that do not run the relevant checks
  • target repository conventions being overwritten by harness defaults

Report findings first, ordered by severity, with file and line references when

available. Do not modify files during a review unless the user explicitly asks

for fixes.

Output Contract

Before finishing harness adoption work, verify:

  • the target repository was inspected before edits
  • new guidance is specific to the target repository
  • changed checks can be run locally or have a documented manual substitute
  • failure memory was recorded when required, or the final response explains why

it was skipped

  • generated docs or indexes are refreshed
  • the final report names every command run and its result

Optional Reference

The prompt-first workflow in

https://github.com/baskduf/harness-starter-kit is a reference implementation

of these ideas. Use it as reference material only when the user asks for it or

when the repository already includes it. The target repository remains the

source of truth.

Other skills for the same job

different authors, same section of the catalogue
Webapp Testing
by anthropics
vendor ×12

Toolkit for interacting with and testing local web applications using Playwright. Supports verifying frontend functionality, debugging UI behavior, capturing browser screenshots, and viewing browser logs.

6k tokens scripts
Finishing A Development Branch
by ZhanlinCui
×7

Use when implementation is complete, all tests pass, and you need to decide how to integrate the work - guides completion of development work by presenting structured options for merge, PR, or cleanup

1k tokens
Test Driven Development
by w95
×7

Use when implementing any feature or bugfix, before writing implementation code

2k tokens
Systematic Debugging
by ratacat
×7

Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes

10k tokens scripts
Verification Before Completion
by ZhanlinCui
×6

Use when about to claim work is complete, fixed, or passing, before committing or creating PRs - requires running verification commands and confirming output before making any success claims; evidence before assertions always

1k tokens
Backtest Expert
by BaggaT236
×3

Expert guidance for systematic backtesting of trading strategies. Use when developing, testing, stress-testing, or validating quantitative trading strategies. Covers "beating ideas to death" methodology, parameter robustness testing, slippage modeling, bias prevention, and interpreting backtest results. Applicable when user asks about backtesting, strategy validation, robustness testing, avoiding overfitting, or systematic trading development.

15k tokens scripts
Adaptyv
by christophacham
×3

Cloud laboratory platform for automated protein testing and validation. Use when designing proteins and needing experimental validation including binding assays, expression testing, thermostability measurements, enzyme activity assays, or protein sequence optimization. Also use for submitting experiments via API, tracking experiment status, downloading results, optimizing protein sequences for better expression using computational tools (NetSolP, SoluProt, SolubleMPNN, ESM), or managing protein design workflows with wet-lab validation.

16k tokens
Aeon
by christophacham
×3

This skill should be used for time series machine learning tasks including classification, regression, clustering, forecasting, anomaly detection, segmentation, and similarity search. Use when working with temporal data, sequential patterns, or time-indexed observations requiring specialized algorithms beyond standard ML approaches. Particularly suited for univariate and multivariate time series analysis with scikit-learn compatible APIs.

19k tokens

How to use it

Copy the folder

Take github/harness-engineering from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.