mcpbeat

Testing Skills

3 538 testing skills from 425 authors. They run test suites, check accessibility and catch what broke. Half of them fit into 1 800 tokens or less — that is what one costs your context window when the agent loads it. 405 ship runnable scripts rather than instructions alone. 2 of them cannot work without an MCP server, most often rube. We also found 403 copies of these same skills sitting in other people's repositories — counted once here, not 403 times.

3 538 unique 425 authors 2 160 updated this month 320 from vendors

1 800
tokens, median
what a typical one costs in context
405
ship scripts
code that runs, not instructions alone
2
need a server
most often rube
403
copies elsewhere
counted once here, not once per repository

49–96 of 3 538

page 2 of 74
Ssh Penetration Testing ×2
ComeOnOliver

This skill should be used when the user asks to "pentest SSH services", "enumerate SSH configurations", "brute force SSH credentials", "exploit SSH vulnerabilities", "perform SSH tunneling", or "audit SSH security". It provides comprehensive SSH penetration testing methodologies and techniques.

8k tokens
Tdd Orchestrator ×2
ComeOnOliver

Master TDD orchestrator specializing in red-green-refactor discipline, multi-agent workflow coordination, and comprehensive test-driven development practices. Enforces TDD best practices across teams with AI-assisted testing and modern frameworks. Use PROACTIVELY for TDD implementation and governance.

5k tokens
Tdd Workflow ×2
ComeOnOliver

Test-Driven Development workflow principles. RED-GREEN-REFACTOR cycle.

3k tokens
Tdd Workflows Tdd Cycle ×2
ComeOnOliver

Use when working with tdd workflows tdd cycle

5k tokens
Tdd Workflows Tdd Green ×2
ComeOnOliver

Implement the minimal code needed to make failing tests pass in the TDD green phase.

9k tokens
Tdd Workflows Tdd Red ×2
ComeOnOliver

Generate failing tests for the TDD red phase to define expected behavior and edge cases.

3k tokens
Tdd Workflows Tdd Refactor ×2
ComeOnOliver

Use when working with tdd workflows tdd refactor

4k tokens
Temporal Python Testing ×2
ComeOnOliver

Test Temporal workflows with pytest, time-skipping, and mocking strategies. Covers unit testing, integration testing, replay testing, and local development setup. Use when implementing Temporal workflow tests or debugging test failures.

16k tokens
Test Automator ×2
ComeOnOliver

Master AI-powered test automation with modern frameworks, self-healing tests, and comprehensive quality engineering. Build scalable testing strategies with advanced CI/CD integration. Use PROACTIVELY for testing automation or quality assurance.

5k tokens
Test Fixing ×2
ComeOnOliver

Run tests and systematically fix all failing tests using smart error grouping. Use when user asks to fix failing tests, mentions test failures, runs test suite and failures occur, or requests to make tests pass.

3k tokens
UI Visual Validator ×2
ComeOnOliver

Rigorous visual validation expert specializing in UI testing, design system compliance, and accessibility verification. Masters screenshot analysis, visual regression testing, and component validation. Use PROACTIVELY to verify UI modifications have achieved their intended goals through comprehensive visual analysis.

5k tokens
Unit Testing Test Generate ×2
ComeOnOliver

Generate comprehensive, maintainable unit tests across languages with strong coverage and edge case focus.

5k tokens
Web3 Testing ×2
ComeOnOliver

Test smart contracts comprehensively using Hardhat and Foundry with unit tests, integration tests, and mainnet forking. Use when testing Solidity contracts, setting up blockchain test suites, or validating DeFi protocols.

6k tokens
Wordpress Penetration Testing ×2
ComeOnOliver

This skill should be used when the user asks to "pentest WordPress sites", "scan WordPress for vulnerabilities", "enumerate WordPress users, themes, or plugins", "exploit WordPress vulnerabilities", or "use WPScan". It provides comprehensive WordPress security assessment methodologies.

7k tokens
Workflow Patterns ×2
ComeOnOliver

Use this skill when implementing tasks according to Conductor's TDD workflow, handling phase checkpoints, managing git commits for tasks, or understanding the verification protocol.

6k tokens
Bats Testing Patterns ×2
ComeOnOliver

Master Bash Automated Testing System (Bats) for comprehensive shell script testing. Use when writing tests for shell scripts, CI/CD pipelines, or requiring test-driven development of shell utilities.

5k tokens
Debugging Strategies ×2
ComeOnOliver

Master systematic debugging techniques, profiling tools, and root cause analysis to efficiently track down bugs across any codebase or technology stack. Use when investigating bugs, performance issues, or unexpected behavior.

6k tokens
E2e Testing Patterns ×2
ComeOnOliver

Master end-to-end testing with Playwright and Cypress to build reliable test suites that catch bugs, improve confidence, and enable fast deployment. Use when implementing E2E tests, debugging flaky tests, or establishing testing standards.

6k tokens
Javascript Testing Patterns ×2
ComeOnOliver

Implement comprehensive testing strategies using Jest, Vitest, and Testing Library for unit tests, integration tests, and end-to-end testing with mocking, fixtures, and test-driven development. Use when writing JavaScript/TypeScript tests, setting up test infrastructure, or implementing TDD/BDD workflows.

11k tokens
LLM Evaluation ×2
ComeOnOliver

Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking. Use when testing LLM performance, measuring AI application quality, or establishing evaluation frameworks.

6k tokens
Sast Configuration ×2
ComeOnOliver

Configure Static Application Security Testing (SAST) tools for automated vulnerability detection in application code. Use when setting up security scanning, implementing DevSecOps practices, or automating code vulnerability detection.

4k tokens
Temporal Python Testing ×2
ComeOnOliver

Test Temporal workflows with pytest, time-skipping, and mocking strategies. Covers unit testing, integration testing, replay testing, and local development setup. Use when implementing Temporal workflow tests or debugging test failures.

20k tokens
Web3 Testing ×2
ComeOnOliver

Test smart contracts comprehensively using Hardhat and Foundry with unit tests, integration tests, and mainnet forking. Use when testing Solidity contracts, setting up blockchain test suites, or validating DeFi protocols.

5k tokens
Evaluating Code Models ×1
Orchestra-Research

Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics. Use when benchmarking code models, comparing coding abilities, testing multi-language support, or measuring code generation quality. Industry standard from BigCode Project used by HuggingFace leaderboards.

10k tokens
Dart Generate Test Mocks vendor ×1
flutter

Define and generate mock objects for external dependencies using `package:mockito` and `build_runner`. Use when unit testing classes that depend on complex external services like APIs or databases.

2k tokens
Dart Add Unit Test vendor ×1
flutter

Write and organize unit tests for functions, methods, and classes using `package:test`. Use when creating new logic or fixing bugs to ensure code remains correct and regression-free.

1k tokens
Durable Objects vendor ×1
openai

Create and review Cloudflare Durable Objects. Use when building stateful coordination (chat rooms, multiplayer games, booking systems), implementing RPC methods, SQLite storage, alarms, WebSockets, or reviewing DO code for best practices. Covers Workers integration, wrangler config, and testing with Vitest. Biases towards retrieval from Cloudflare docs over pre-trained knowledge.

7k tokens
Test Discipline vendor ×1
microsoft

Update tests when changing APIs — no exceptions

484 tokens
Systematic Debugging ×1
mrgoonie

Four-phase debugging framework that ensures root cause investigation before attempting fixes. Never jump to solutions.

5k tokens
Scale Game ×1
mrgoonie

Test at extremes (1000x bigger/smaller, instant/year-long) to expose fundamental truths hidden at normal scales

576 tokens
Genlayer Intelligent Contracts ×1
internet-court

Internet Court adapter for GenLayer Intelligent Contract supervision. Use to specify agent-performance rubrics, evidence schemas, decision outputs, and ERC-7710 connector expectations, while delegating actual GenLayer contract writing, linting, testing, deployment, and CLI interaction to the official GenLayer skills at https://skills.genlayer.com/.

2k tokens
Direct Tests ×1
internet-court

Write and run fast direct mode tests for GenLayer intelligent contracts.

2k tokens
Integration Tests ×1
internet-court

Write and run integration tests against a GenLayer environment.

2k tokens
Near Smart Contracts ×1
internet-court

NEAR Protocol smart contract development in Rust. Use when writing, reviewing, or deploying NEAR smart contracts. Covers contract structure, state management, cross-contract calls, testing, security, and optimization patterns. Based on near-sdk v5.x with modern macro syntax.

18k tokens
Android Testing ×1
new-silvermoon

Comprehensive testing strategy involving Unit, Integration, Hilt, and Screenshot tests.

731 tokens
Tdd Guide ×1
alirezarezvani

Comprehensive Test Driven Development guide for engineering subagents with multi-framework support, coverage analysis, and intelligent test generation

40k tokens scripts
Diagnose ×1
AvdLee

Disciplined diagnosis loop for hard bugs and performance regressions. Reproduce → minimise → hypothesise → instrument → fix → regression-test. Use when user says "diagnose this" / "debug this", reports a bug, says something is broken/throwing/failing, or describes a performance regression.

2k tokens scripts
Qa ×1
AvdLee

Interactive QA session where user reports bugs or issues conversationally, and the agent files GitHub issues. Explores the codebase in the background for context and domain language. Use when user wants to report bugs, do QA, file issues conversationally, or mentions "QA session".

1k tokens
Improve ×1
gadicc

Survey any codebase as a senior advisor and produce prioritized, self-contained implementation plans for OTHER models/agents to execute. Strictly read-only on source code — never implements, fixes, or refactors anything itself. Use when asked to audit a codebase, find improvement opportunities (bugs, security, performance, test coverage, tech debt, migrations, DX), suggest features or where to take the project next (roadmap, product direction), or generate handoff plans for another agent to implement.

11k tokens
Diagnosing Bugs ×1
mxyhi

Diagnosis loop for hard bugs and performance regressions. Use when the user says "diagnose"/"debug this", or reports something broken/throwing/failing/slow.

3k tokens scripts
Migrate To Shoehorn ×1
mxyhi

Migrate test files from `as` type assertions to @total-typescript/shoehorn. Use when user mentions shoehorn, wants to replace `as` in tests, or needs partial test data.

965 tokens
Swift Testing Expert ×1
AvdLee

Expert guidance for Swift Testing: test structure, #expect/#require macros, traits and tags, parameterized tests, test plans, parallel execution, async waiting patterns, and XCTest migration. Use when writing new Swift tests, modernizing XCTest suites, debugging flaky tests, or improving test quality and maintainability in Apple-platform or Swift server projects.

10k tokens
Pair Programming ×1
Microck

AI-assisted pair programming with multiple modes (driver/navigator/switch), real-time verification, quality monitoring, and comprehensive testing. Supports TDD, debugging, refactoring, and learning sessions. Features automatic role switching, continuous code review, security scanning, and performance optimization with truth-score verification.

6k tokens
E2e Testing ×1
loulanyue

Playwright E2E testing patterns, Page Object Model, configuration, CI/CD integration, artifact management, and flaky test strategies.

2k tokens
Ab Test Setup ×1
lingxling

Structured guide for setting up A/B tests with mandatory gates for hypothesis, metrics, and execution readiness.

2k tokens
Active Directory Attacks ×1
lingxling

Provide comprehensive techniques for attacking Microsoft Active Directory environments. Covers reconnaissance, credential harvesting, Kerberos attacks, lateral movement, privilege escalation, and domain dominance for red team operations and penetration testing.

5k tokens
Agent Evaluation ×1
lingxling

Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarks

9k tokens
Airflow Dag Patterns ×1
lingxling

Build production Apache Airflow DAGs with best practices for operators, sensors, testing, and deployment. Use when creating data pipelines, orchestrating workflows, or scheduling batch jobs.

4k tokens