mcpbeat

Testing Claude Skills

3 535 testing skills from 425 authors. They run test suites, check accessibility and catch what broke. Half of them fit into 1 800 tokens or less — that is what one costs your context window when the agent loads it. 404 ship runnable scripts rather than instructions alone. 2 of them cannot work without an MCP server, most often rube. We also found 403 copies of these same skills sitting in other people's repositories — counted once here, not 403 times.

3 535 unique 425 authors 2 151 updated this month 319 from vendors

1 800
tokens, median
what a typical one costs in context
404
ship scripts
code that runs, not instructions alone
2
need a server
most often rube
403
copies elsewhere
counted once here, not once per repository

913–960 of 3 535

page 20 of 74
Recipe Add Integration Tests
shinpr

Add integration/E2E tests to existing codebase using Design Docs

2k tokens
Recipe Prepare Implementation
shinpr

Verifies the work plan is implementable end-to-end and resolves verification-lane / fixture / E2E-environment gaps before the build phase begins. Use when "implement-ready/verification readiness/lane setup/E2E environment missing" is mentioned, or before any build phase begins on a work plan whose readiness has not been preflight-checked.

3k tokens
Test Fixing
mhattingpete

Run tests and systematically fix all failing tests using smart error grouping. Use when user asks to fix failing tests, mentions test failures, runs test suite and failures occur, or requests to make tests pass.

741 tokens
Code Auditor
mhattingpete

Performs comprehensive codebase analysis covering architecture, code quality, security, performance, testing, and maintainability. Use when user wants to audit code quality, identify technical debt, find security issues, assess test coverage, or get a codebase health check.

896 tokens
Build Test Verify
bitwarden

Build the project, run tests, lint, format, spell check, generate mocks, or verify the build passes for Bitwarden iOS. Use when asked to "build", "run tests", "lint", "format", "verify build", "check if it compiles", "run swiftlint", "run swiftformat", "generate mocks", or to execute any part of the build/test/verify pipeline.

1k tokens
Reviewing Changes
bitwarden

Performs comprehensive code reviews for Bitwarden iOS projects, verifying architecture compliance, style guidelines, compilation safety, test coverage, and security requirements. Use when reviewing pull requests, checking commits, analyzing code changes, verifying Bitwarden coding standards, evaluating unidirectional data flow pattern, checking services container dependency injection usage, reviewing security implementations, or assessing test coverage. Automatically invoked by CI pipeline or manually for interactive code reviews.

11k tokens
Converting Xctest To Swift Testing
bitwarden

Convert an XCTest file to Swift Testing framework. Use when asked to "convert to Swift Testing", "migrate XCTest", "convert test file", "xctest to swift testing", "migrate tests to Swift Testing", or when explicitly asked to convert existing XCTest-based tests.

6k tokens
Fixing Flaky Tests
bitwarden

> Diagnose and fix flaky (intermittently failing) tests in Bitwarden iOS. Finds the root cause of non-deterministic test failures (race conditions, timing issues, shared state, order dependence), applies a targeted fix, then stress-tests the fix by running the test 100 times to confirm stability. Use this skill whenever a test is failing intermittently, sometimes passes / sometimes fails, or someone describes a test as flaky, unstable, or non-deterministic — even if the exact "intermittent failure", "test is non-deterministic", "test fails sometimes", "test randomly fails".

2k tokens
Testing Ios Code
bitwarden

Write tests, add test coverage, unit test, or add missing tests for Bitwarden iOS. Use when asked to "write tests", "add test coverage", "test this", "unit test", "add tests for", "missing tests", or when creating test files for new implementations.

8k tokens
Adding Dbt Unit Test vendor
dbt-labs

Creates unit test YAML definitions that mock upstream model inputs and validate expected outputs. Use when adding unit tests for a dbt model or practicing test-driven development (TDD) in dbt.

9k tokens
Answering Natural Language Questions With Dbt vendor
dbt-labs

Writes and executes SQL queries against the data warehouse using dbt's Semantic Layer or ad-hoc SQL to answer business questions. Use when a user asks about analytics, metrics, KPIs, or data (e.g., "What were total sales last quarter?", "Show me top customers by revenue"). NOT for validating, testing, or building dbt models during development.

2k tokens
Using Dbt For Analytics Engineering vendor
dbt-labs

Builds and modifies dbt models, writes SQL transformations using ref() and source(), creates tests, and validates results with dbt show. Use when doing any dbt work - building or modifying models, debugging errors, exploring unfamiliar data sources, writing tests, or evaluating impact of changes.

10k tokens scripts
Apm Integrations
DataDog

| dd-trace-py integration development guide. Use when creating, modifying, or debugging contrib integrations in the Python tracer. Covers the patch module system, context_with_data, context_with_event (new), registration, testing with riot, and common anti-patterns. LLM/AI integrations should use this skill for APM-side workflow only; use llmobs-integrations for LLMObs-specific lifecycle, extraction, streaming, and VCR guidance. Pin is DEPRECATED. "trace_handlers", "PATCH_MODULES", "context_with_data", "context_with_event", "TracingEvent", "VCR", "cassette", "generative-ai", "LLM integration", "riot", "riotfile", "suitespec", "new integration", "wrap", "unwrap".

6k tokens
Llmobs Integrations
DataDog

| dd-trace-py LLMObs integration development guide. Use when creating, modifying, or debugging LLMObs integrations for LLM/AI libraries in the Python tracer. Covers BaseLLMIntegration, stream handling, message extraction, token counting, tool call parsing, and VCR-based testing patterns. "_llmobs_set_tags", "BaseStreamHandler", "submit_to_llmobs", "integration.trace", "LLM span", "VCR", "cassette", "anthropic", "openai", "google_genai", "claude_agent_sdk", "generative-ai", "LLM integration", "llmobs_enabled".

9k tokens
Run Tests
DataDog

> Validate code changes by intelligently selecting and running the appropriate test suites. Use this when editing code to verify changes work correctly, run tests, validate functionality, or check for regressions. Automatically discovers affected test suites, selects the minimal set of venvs needed for validation, and handles test execution with Docker services as needed.

4k tokens
Ambiguity Stress Test
lawve-ai

>- Adversarially stress-tests a legal text — a contract, statute, regulation, or people governed by it will later disagree about what it means and turns each into a concrete dispute scenario with both sides' arguments, the likely outcome, and a fix. Use it whenever someone wants to pressure-test, red-team, audit, or find weak spots, loopholes, gaps, ambiguities, or drafting problems in a legal document; whenever a drafter wants to tighten a contract, statute, regulation, or opinion before it issues; whenever a litigator wants to mine an opinion or contract for arguments; or whenever someone hands over a legal text and asks where it will be fought over or for issue-spotting. Trigger it for contract review, statutory-ambiguity analysis, judicial-opinion scope analysis, and drafting QA — even if the user never says "stress-test" or "ambiguity."

18k tokens
Legal Simulation Patrick Munro
lawve-ai

Framework for demonstrating AI capabilities in legal contexts. Provides detailed personas across tenant law, business contracts, startup disputes, employment claims, and consumer protection with progressive complexity scenarios. Use when: (1) Demonstrating AI-powered legal triage or intake systems, (2) Showcasing responsible AI-assisted client interactions, (3) Training staff on appropriate AI use in legal contexts, (4) Creating realistic scenarios for legal tech presentations, (5) Developing educational materials about AI in legal services, or (6) Testing AI-powered legal information systems in controlled environments.

14k tokens
Legal Test Builder Patrick Munro
lawve-ai

> Builds a high-fidelity interactive legal assessment as a single self-contained HTML artifact. Output includes a live countdown timer, contract review tasks with hover-annotated problem clauses, candidate answer textareas, model answers hidden behind reveal blocks, scenario-based legal memo tasks, strategy and function-building questions, and a pre-submission checklist that encodes the marking criteria. Use when the user needs to (1) assess a legal candidate with a realistic timed exercise, (2) train or onboard junior lawyers using problem sets rather than doctrine, (3) help a candidate prepare for a real take-home assessment they are facing, (4) build educational materials for law students, in-house teams, or compliance training, or (5) produce scenario-based training modules on specific legal topics. Triggers on "legal test", "take-home", "mock exam", "contract redline exercise", "candidate assessment", "legal training exercise", "practice test", or similar phrasing even when informal.

21k tokens
New Designation Screening Test
lawve-ai

Generate a spreadsheet of test entries — newly designated names from OFAC, OFSI, and EU sanctions lists plus deliberate variations of those names — to validate that a sanctions screening system catches fresh designations and is tuned to the right fuzziness threshold. Use this whenever the user asks for sanctions list update test data, screening regression test data, screening QA, fuzzy match calibration, or wants to verify their screening lists are current. Trigger even if the user doesn't say 'screening' explicitly — phrases like 'test my sanctions list', 'check our SDN coverage', 'is my list up to date', or 'build me a regression set from the latest designations' should also invoke this skill.

5k tokens
Opposing Counsel Review
lawve-ai

>- Act as experienced opposing counsel to attack, undermine, and expose weaknesses in a legal argument, submission, witness statement, or structured reasoning. 1. A core theory of attack identifying the single most effective way to defeat the argument; 2. A reconstructed version of the opposing argument stripped of rhetoric to expose its fragility; 3. Primary lines of attack grouped by category (legal misstatement, evidential gaps, causation failures, internal inconsistency, over-reliance on assertion, procedural weakness); 4. An "if I were the judge" section showing how a sceptical tribunal would dismantle the argument; 5. Surgical strikes - 3 to 5 high-impact points ready for oral submissions; and 6. An analysis of what the argument is trying to hide. Written in formal, adversarial British English for a legally trained audience.

2k tokens
Oral Argument
lawve-ai

| Prepare a lawyer for an adversarial proceeding the way Neal Katyal's "Harvey" prior opinions and questions, predict the specific questions you will face, map narrow "escape routes" each judge can walk through without abandoning prior commitments, and spar adversarially until only the strongest answers survive. The model is the sparring partner. The human still wins the case. evidentiary hearing, deposition (taking or defending), arbitration, mediation, or any proceeding where a known decision-maker will fire questions at you in real time. Also useful for stress-testing a brief before filing. named judges or arbitrators, drafts a brief and wants it pressure-tested, asks "what will Justice X / Judge Y ask me," or talks about preparing for a specific bench.

2k tokens
Settlement Pressure Tester Larissa Meredith Flister
lawve-ai

This skill stress-tests a proposed settlement position before an offer goes out or comes back: the assumptions it depends on, your leverage and the opponent’s, the evidential weaknesses, the likely opponent response, and the timing and costs pressures around it. It structures settlement judgment for a better-informed decision; it does not advise whether to settle.

3k tokens
Sustainable Opposing Counsel Review
lawve-ai

>- Produces an adversarial attack on a legal argument that survives reply. Runs the turned on that pass to cut every point that collapses under challenge - attacks on conceded facts, wrong-forum objections, speculation about documents and motives, gotchas with innocent explanations, overclaims, self-refuting assertions, and padding. The deliverable is one standalone attack containing only what can be defended, capped at five heads, not a discussion of the discarded draft. Use to attack, stress-test, red-team or rebut a submission, brief, motion, witness statement, letter or structured legal reasoning. Triggers on "attack this but only with points that hold", "no cheap shots", "what survives reply", "which points can I actually defend", "give me the version I can file", "double-pass adversarial review". Also use when an earlier adversarial review came back overlong or scattershot. Formal British English.

9k tokens
Bloc
evanca

Use when creating a Cubit or Bloc, modeling state with sealed classes or status enums, wiring BlocBuilder/BlocListener/BlocProvider, writing bloc tests, or choosing between Cubit and Bloc.

3k tokens
Firebase Cloud Functions
evanca

Use when calling callable functions (httpsCallable), passing data to server-side logic, handling function errors/timeouts, configuring regions, or testing with the Emulator Suite.

1k tokens
Firebase In App Messaging
evanca

Use when running in-app message campaigns, triggering/suppressing messages, configuring opt-in data collection, testing on specific devices, or handling message callbacks.

1k tokens
Firebase Remote Config
evanca

Use when implementing feature flags, running A/B tests, setting parameter defaults, fetching/activating config, or enabling real-time config updates.

1k tokens
Patrol E2e Testing
evanca

Use when writing E2E/integration tests, testing native interactions like permissions or system dialogs, capturing UI regressions, or validating cross-platform behavior (Patrol 4.x).

1k tokens
Revenuecat Testing
evanca

Use when testing RevenueCat purchases/subscriptions, setting up sandbox testing, debugging a purchase/restore/trial, verifying entitlements or events, or writing an IAP QA plan.

3k tokens
Riverpod
evanca

Use when setting up providers, combining requests, managing state disposal, passing arguments, performing side effects, or testing providers (Riverpod).

2k tokens
Testing
evanca

Use when writing or reviewing Flutter/Dart tests (unit, widget, golden), fixing flaky tests, adding coverage, or choosing between unit and widget tests.

1k tokens
Generate MCP App UI vendor
microsoft

Generate an MCP App widget (self-contained HTML) for an MCP tool. Describe the visual you want and paste your tool's test output. Use when user asks to create an MCP App, widget, or visual for a tool.

2k tokens
Add Sample Data vendor
microsoft

>- Populates Dataverse tables with sample records for testing and demoing a Power Pages site. Use when the user wants to add sample data, seed data, generate test records, or insert demo data into their tables.

5k tokens
Test Site vendor
microsoft

>- Tests a deployed, activated Power Pages site at runtime using browser-based navigation, page crawling, and API request verification via Playwright. Use when the user wants to test, verify, or smoke-test their deployed site.

11k tokens
Nw Abr Critique Dimensions
nWave-ai

Review dimensions for validating agent quality - template compliance, safety, testing, and priority validation

1k tokens
Nw Agent Testing
nWave-ai

5-layer testing approach for agent validation including adversarial testing, security validation, and prompt injection resistance

881 tokens
Nw Ab Critique Dimensions
nWave-ai

Review dimensions for validating agent quality - template compliance, safety, testing, and priority validation

1k tokens
Nw Ad Critique Dimensions
nWave-ai

Review dimensions for acceptance test quality - happy path bias, GWT compliance, business language purity, coverage completeness, walking skeleton user-centricity, priority validation, observable behavior assertions, traceability coverage, and walking skeleton boundary proof

2k tokens
Nw Bdd Methodology
nWave-ai

BDD patterns for acceptance test design - Given-When-Then structure, scenario writing rules, pytest-bdd implementation, anti-patterns, and living documentation

2k tokens
Nw Bugfix
nWave-ai

Bug fix workflow: root cause analysis → user review → regression test + fix via TDD

1k tokens
Nw Collaboration And Handoffs
nWave-ai

Cross-agent collaboration protocols, workflow handoff patterns, and commit message formats for TDD/Mikado/refactoring workflows

1k tokens
Nw Discover
nWave-ai

Conducts evidence-based product discovery through customer interviews and assumption testing. Use at project start to validate problem-solution fit.

2k tokens
Nw Distill
nWave-ai

Acceptance test creation methodology for the DISTILL wave. Domain knowledge for the acceptance designer agent: port-to-port principle, prior wave reading, wave-decision reconciliation, graceful degradation, and document back-propagation.

17k tokens
Nw Execute
nWave-ai

Dispatches one unit of DELIVER work to a specialized agent for TDD execution. Runs a single roadmap.json step through the TDD cycle.

3k tokens
Nw Finalize
nWave-ai

Archives a completed feature to docs/evolution/, migrates lasting artifacts to permanent directories, and cleans up the temporary workspace. Use after all implementation steps pass and mutation testing completes.

2k tokens
Nw Hexagonal Testing
nWave-ai

5-layer agent output validation, I/O contract specification, vertical slice development, and test doubles policy with per-layer examples

1k tokens
Nw Interviewing Techniques
nWave-ai

Mom Test questioning toolkit, JTBD analysis, interview conduct, assumption testing framework, and hypothesis design

1k tokens
Nw Mutation Test
nWave-ai

Runs feature-scoped mutation testing to validate test suite quality. Use after implementation to verify tests catch real bugs (kill rate >= 80%).

1k tokens