mcpbeat Sign in

Testing Quality Standards Skill for Claude

Defines testing quality metrics, coverage thresholds, and anti-patterns. Use when establishing test gates or validating a test suite's coverage targets.

3k tokens
context cost
the whole folder, loaded on every use
4
files
instructions only
0
copies elsewhere
how many repositories repackaged it
324
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/athola/claude-night-market --skill testing-quality-standards

The instruction itself

15 sections, as written by the author

Testing Quality Standards

Shared quality standards and metrics for testing across all plugins in the Claude Night Market ecosystem.

When To Use

  • Establishing test quality gates and coverage targets
  • Validating test suite against quality standards

When NOT To Use

  • Exploratory testing or spike work
  • Projects with established quality gates that meet requirements

Table of Contents

  • Coverage Thresholds
  • Quality Metrics
  • Detailed Topics

Coverage Thresholds

| Level | Coverage | Use Case |

|-------|----------|----------|

| Minimum | 60% | Legacy code |

| Standard | 80% | Normal development |

| High | 90% | Critical systems |

| detailed | 95%+ | Safety-critical |

Quality Metrics

Structure

  • [ ] Clear test organization
  • [ ] Meaningful test names
  • [ ] Proper setup/teardown
  • [ ] Isolated test cases

Coverage

  • [ ] Critical paths covered
  • [ ] Edge cases tested
  • [ ] Error conditions handled
  • [ ] Integration points verified

Maintainability

  • [ ] DRY test code
  • [ ] Reusable fixtures
  • [ ] Clear assertions
  • [ ] Minimal mocking

Reliability

  • [ ] No flaky tests
  • [ ] Deterministic execution
  • [ ] No order dependencies
  • [ ] Fast feedback loop

Detailed Topics

For implementation patterns and examples:

  • Anti-Patterns - Common testing mistakes with before/after examples
  • Best Practices - Core testing principles and exit criteria
  • Content Assertion Levels - L1/L2/L3 taxonomy for testing LLM-interpreted markdown files

Integration with Plugin Testing

This skill provides foundational standards referenced by:

  • pensive:test-review - Uses coverage thresholds and quality metrics
  • parseltongue:python-testing - Uses anti-patterns and best practices
  • sanctum:test-* - Uses quality checklist and content assertion levels for test validation
  • imbue:proof-of-work - Uses content assertion levels to enforce Iron Law on execution markdown

Reference in your skill's frontmatter:

dependencies: [leyline:testing-quality-standards]

Verification: Run pytest -v to verify tests pass.

Troubleshooting

Common Issues

Tests not discovered

Ensure test files match pattern test_*.py or *_test.py. Run pytest --collect-only to verify.

Import errors

Check that the module being tested is in PYTHONPATH or install with pip install -e .

Async tests failing

Install pytest-asyncio and decorate test functions with @pytest.mark.asyncio

Exit Criteria

  • [ ] Coverage threshold met for the project tier: 60% minimum for

legacy code, 80% for normal development, 90% for critical

systems, 95%+ for safety-critical; measured with

pytest --cov and threshold enforced in pyproject.toml

  • [ ] All four quality metric checklists pass: Structure (clear

organization, meaningful names, setup/teardown, isolation),

Coverage (critical paths, edge cases, error conditions,

integration points), Maintainability (DRY fixtures, clear

assertions, minimal mocking), Reliability (no flaky tests,

deterministic execution, no order dependencies)

  • [ ] Test files match discovery pattern test_*.py or *_test.py

confirmed by pytest --collect-only returning no errors

  • [ ] No snapshot tests present on non-deterministic output (hash

maps, timestamps, UUIDs); any found flagged as anti-patterns

per modules/anti-patterns.md

Other skills for the same job

different authors, same section of the catalogue
Webapp Testing
by anthropics
vendor ×12

Toolkit for interacting with and testing local web applications using Playwright. Supports verifying frontend functionality, debugging UI behavior, capturing browser screenshots, and viewing browser logs.

6k tokens scripts
Finishing A Development Branch
by ZhanlinCui
×7

Use when implementation is complete, all tests pass, and you need to decide how to integrate the work - guides completion of development work by presenting structured options for merge, PR, or cleanup

1k tokens
Test Driven Development
by w95
×7

Use when implementing any feature or bugfix, before writing implementation code

2k tokens
Systematic Debugging
by ratacat
×7

Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes

10k tokens scripts
Verification Before Completion
by ZhanlinCui
×6

Use when about to claim work is complete, fixed, or passing, before committing or creating PRs - requires running verification commands and confirming output before making any success claims; evidence before assertions always

1k tokens
Backtest Expert
by BaggaT236
×3

Expert guidance for systematic backtesting of trading strategies. Use when developing, testing, stress-testing, or validating quantitative trading strategies. Covers "beating ideas to death" methodology, parameter robustness testing, slippage modeling, bias prevention, and interpreting backtest results. Applicable when user asks about backtesting, strategy validation, robustness testing, avoiding overfitting, or systematic trading development.

15k tokens scripts
Adaptyv
by christophacham
×3

Cloud laboratory platform for automated protein testing and validation. Use when designing proteins and needing experimental validation including binding assays, expression testing, thermostability measurements, enzyme activity assays, or protein sequence optimization. Also use for submitting experiments via API, tracking experiment status, downloading results, optimizing protein sequences for better expression using computational tools (NetSolP, SoluProt, SolubleMPNN, ESM), or managing protein design workflows with wet-lab validation.

16k tokens
Aeon
by christophacham
×3

This skill should be used for time series machine learning tasks including classification, regression, clustering, forecasting, anomaly detection, segmentation, and similarity search. Use when working with temporal data, sequential patterns, or time-indexed observations requiring specialized algorithms beyond standard ML approaches. Particularly suited for univariate and multivariate time series analysis with scikit-learn compatible APIs.

19k tokens

How to use it

Copy the folder

Take athola/testing-quality-standards from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.

Install what it needs

The instructions reference pip. Without those the skill loads but fails at the first command.