mcpbeat

Testing Claude Skills

3 535 testing skills from 425 authors. They run test suites, check accessibility and catch what broke. Half of them fit into 1 800 tokens or less — that is what one costs your context window when the agent loads it. 404 ship runnable scripts rather than instructions alone. 2 of them cannot work without an MCP server, most often rube. We also found 403 copies of these same skills sitting in other people's repositories — counted once here, not 403 times.

3 535 unique 425 authors 2 151 updated this month 319 from vendors

1 800
tokens, median
what a typical one costs in context
404
ship scripts
code that runs, not instructions alone
2
need a server
most often rube
403
copies elsewhere
counted once here, not once per repository

2 401–2 448 of 3 535

page 51 of 74
Test Driven Development
ravila4

(TDD) Use when implementing any feature or bugfix. Apply the Logic Gate to identify what needs tests, then use the Iron Rule for strict test-first development on logic.

1k tokens
Odoo 19
unclecatvn

>- Odoo 19 development knowledge base with 18 specialized guides covering Actions (ir.actions.*, cron jobs, server actions), Controllers (HTTP routing, endpoints, auth types), Data files (XML/CSV records, shortcuts, noupdate), API Decorators (@api.depends, @api.constrains, @api.ondelete, @api.onchange, @api.model, @api.private), SQL Constraints (models.Constraint replacing _sql_constraints), Database Indexes (models.Index), Module development (manifest, wizards, reports), Field types (Char, Text, Monetary, relational fields), Manifest configuration (__manifest__.py, dependencies, asset bundles), Mixins (mail.thread, mail.activity.mixin, mail.alias.mixin, utm.mixin), ORM Model methods (search, CRUD, domain filters, recordsets, CamelCase model naming), Migration scripts (pre/post/end hooks, data migration), OWL frontend components (hooks, services, lifecycle), Performance optimization (N+1 prevention, batch ops, _read_group), QWeb Reports (PDF/HTML, paper formats, barcodes, t-out), Security/ACL (record rules, field permissions, privilege-based groups, @api.private), Testing (TransactionCase, HttpCase, mocking, query count assertions), Transactions (savepoints, UniqueViolation, serialization failures), Translations (i18n, PO files, translatable fields), XML Views (list/form/search, kanban card templates, xpath inheritance, QWeb templates). Use when writing, reviewing, or debugging any Odoo 19 Python or XML code, creating or modifying modules, fixing performance issues, or looking up Odoo 19 API patterns and best practices.

47k tokens
Systematic Debugging
HezaoHezao

4-phase root cause debugging: understand bugs before fixing.

3k tokens
Test Driven Development
HezaoHezao

TDD: enforce RED-GREEN-REFACTOR, tests before code.

2k tokens
Running Bug Review Board
RayFernando1337

>- Runs real-user QA, manual test plans, UX bug hunts, build sign-off, bug filing, and bug triage for web or iOS/iPadOS apps. Use when asked "QA this", "is this ready to ship?", or similar. Produces P0/P1/P2 bug reports, YES/NO phase sign-off, tracker sync guidance, and an HTML QA dashboard; keeps Interactive BRB triage in a separate session.

215k tokens scripts
Bootstrap Ios
RayFernando1337

Bootstrap agents for iOS, iPadOS, macOS, Swift, SwiftUI, SwiftData/Core Data, Swift Testing, Xcode build/test/debug, Simulator, App Intents, or XcodeBuildMCP work. Use before building, fixing, refactoring, QAing, or setting up Apple-platform repos, and when asked to load/install Ray's iOS skills or bootstrap iOS.

6k tokens scripts
Testing Dbt Models
AltimateAI

| (1) Adding or modifying tests in schema.yml files (2) Task mentions "test", "validate", "data quality", "unique", "not_null", or "accepted_values" (3) Ensuring data integrity - primary keys, foreign keys, relationships (4) Debugging test failures or understanding why dbt test failed Matches existing project test patterns and YAML style before adding new tests.

1k tokens
Debugging Dbt Errors
AltimateAI

| (1) Task mentions "fix", "error", "broken", "failing", "debug", "wrong", or "not working" (2) Compilation Error, Database Error, or test failures occur (3) Model produces incorrect output or unexpected results (4) Need to troubleshoot why a dbt command failed Reads full error, checks upstream first, runs dbt build (not just compile) to verify fix.

1k tokens
Animation Jank QA
TheGoat395

Run animation jank QA. Use for GSAP, Motion, CSS transitions, scroll animation, pinned scenes, parallax, hover states, page transitions, mobile motion, reduced-motion fallbacks, performance-heavy animations, and final motion polish before delivery.

1k tokens
Cross Browser QA
TheGoat395

Run cross-browser QA for websites and apps. Use before production or after CSS, media, forms, animation, canvas, layout, or interaction changes to test Chromium, WebKit/Safari-like, and Firefox behavior, browser-specific CSS/media issues, focus/keyboard differences, and responsive rendering.

1k tokens
Image Crop Responsive QA
TheGoat395

Run responsive image crop QA for websites. Use after adding or changing images, hero media, galleries, product grids, portraits, venue photos, screenshots, background images, object-fit/object-position rules, art-directed picture sources, image dimensions, layout stability, and mobile crop polish.

2k tokens
Motion Performance QA
TheGoat395

Run motion performance QA for websites and apps. Use after adding Motion, GSAP, Lenis, scroll scenes, page transitions, hover states, carousels, WebGL sync, observers, timers, RAF loops, or any animation-heavy UI to catch jank, layout shift, leaks, cleanup bugs, reduced-motion gaps, and mobile issues.

1k tokens
Mobile Responsive QA
TheGoat395

Run mobile responsive QA for websites and apps. Use after layout, media, navigation, form, pricing, dashboard, gallery, checkout, or motion changes to inspect phone and tablet viewports, prevent horizontal overflow, fix crop issues, verify touch targets, and make mobile feel intentionally composed.

1k tokens
Premium Web Build Gate
TheGoat395

Premium website pre-build gate for cinematic, editorial, Squarespace-polished, Raycast/Linear/Vercel-precise, or motion-heavy web work. Use before implementing or redesigning a homepage, landing page, portfolio, product site, agency/studio site, or high-polish web app when the user wants non-generic visual quality, strong scroll/motion choreography, strict content scope, or an award-level result. Produces a build-ready UI spec, source-trust decisions, motion storyboard, and QA gates before code.

2k tokens
Responsive Visual Polish QA
TheGoat395

Run responsive visual QA and polish for websites and apps. Use after building or editing any frontend to inspect desktop, laptop, tablet, and mobile widths; fix typography, spacing, overlaps, media crops, interactions, motion, forms, canvas rendering, accessibility, and final build quality.

1k tokens
SEO Technical QA
TheGoat395

Run technical SEO QA for websites. Use for metadata, titles, descriptions, canonical URLs, robots, sitemap, Open Graph, structured data, heading structure, internal links, image alt, crawlability, noindex mistakes, redirects, JavaScript-rendered content, and launch readiness.

1k tokens
Visual Regression Lab
TheGoat395

Use after frontend changes or before delivery to run rendered visual QA with Playwright/browser screenshots, viewport checks, console/network checks, overflow detection, canvas/media verification, before/after comparisons, and responsive polish review.

366 tokens
Website Blueprint First
TheGoat395

Create a website blueprint before coding. Use for new sites, redesigns, image-to-site work, premium landing pages, portfolios, brand sites, product sites, and any frontend project where Codex should define pages, sections, assets, tokens, layout, motion, and QA criteria before implementation.

1k tokens
Bug Fix Protocol
CodeAlive-AI

8-step disciplined bug-fix protocol that treats every production bug as two failures — the code defect itself and the testing system that allowed it through. Use when fixing a production bug, investigating a regression, writing a post-mortem, or auditing a missed defect. Triggers on "fix this bug", "production bug", "regression test", "post-mortem", "test gap", "why did the tests miss this".

2k tokens
Windows QA Engineer
CodeAlive-AI

Use when testing Windows 11 desktop apps (WinForms/WPF/UWP) via UFO UIA/Win32 automation MCP. Triggers on "test this Windows app", "QA the app", "run smoke test", "click the button", "fill the form", "check the UI", "Windows automation", "UFO QA", "verify the dialog", or any Windows desktop UI testing task. Not for web/browser testing (use Playwright), mobile testing, or non-Windows platforms.

10k tokens scripts
Tdd
selmakcby

Test-Driven Development workflow. Use when writing new features, fixing bugs, or refactoring code. Enforces RED-GREEN-REFACTOR cycle with 80%+ coverage.

349 tokens
Fix
avibebuilder

Fix bugs and broken behavior when there is enough evidence to act on a repair path. Use for errors, crashes, incorrect results, API failures (500, 404, 403), CORS problems, database exceptions, broken rendering, duplicated or wrong data, off-by-one mistakes, timezone/date bugs, broken forms, config-caused runtime failures, and regressions. Trigger when the user wants the bug repaired and the conversation already contains a clear failing area, a reproducible failing test, a concrete error path, or a prior diagnosis to implement. Do NOT use for new features, pure explanation, architecture discussion, broad research, or bug reports where the main need is figuring out why the behavior happens — use diagnose for that.

1k tokens
Cook
avibebuilder

Implement, build, create, or add any feature, endpoint, page, component, or functionality. Use this skill whenever the user asks you to write new code or make code changes — whether it's adding an API endpoint, building a UI page, creating an export feature, wiring up a webhook, implementing a search/filter, or any other hands-on coding task. This is the default skill for all 'build this', 'add this', 'create this', 'wire up', 'implement' requests. Covers the full cycle: clarify requirements, plan if needed, write code, verify, and review. Do NOT use for pure research, debugging, documentation, or explanation — only when the user wants working code delivered.

973 tokens
Skill Creator
avibebuilder

Use when the user wants to work on a Claude Code skill file (SKILL.md): writing one from scratch, testing whether an existing one works well, running evals or benchmarks, improving its instructions, or fixing why it isn't triggering. Triggers on: 'make a skill for X', 'test this skill', 'run evals on my SKILL.md', 'touch-skill', sharing a SKILL.md and asking if it's ready to ship. The key signal is intent to create, validate, or improve a skill — not just mention one. Do NOT trigger for general Claude Code questions, hook debugging, or CLAUDE.md configuration.

87k tokens scripts
Testing Patterns vendor
langchain-ai

Unit testing and integration testing best practices

446 tokens
Pi En
share-skills

PI Cognitive AI. Trigger: coding/development/fleet/architecture/API/debugging/bug/error/testing/compile/test/git/make/release/verify/product/requirements/ops/growth/creative/design/collaboration/team/communication/interaction/support, or 2+ failures/looping/giving-up/retry/nevermind

17k tokens
Pi
share-skills

PI 智行合一。触发:编程/开发/fleet/代码/架构/API/调试/修复/优化/bug/报错/测试/编译/compile/test/git/make/发布/验证/产品/需求/运营/增长/创意/设计/协作/团队/沟通/交互/陪伴/情感,或失败2+次/打转/言退/再试试/换个参数/算了

13k tokens zh
Eval Debate
YIKUAIBANZI

测试 use-self 替身会议的辩论质量。给定 persona + 3 个决策场景,运行完整三阶段辩论并按 5 个维度评分,输出质量报告。

1k tokens zh
Eval Consistency
YIKUAIBANZI

测试 use-persona 的角色扮演一致性。给定 persona + 10 个对话场景,生成回复并按 5 个维度评分,输出一致性报告。

941 tokens zh
Minecraft Testing
Jahrome907

Write automated tests for Minecraft mods and plugins for 1.21.x. Covers NeoForge GameTests (@GameTest annotation, GameTestHelper assertions, test structure placement), Fabric game tests (fabric-gametest-api-v1), unit testing non-Minecraft logic with JUnit 5, MockBukkit for Paper/Bukkit plugin testing (mock server, mock player, event dispatching, inventory checking), integration testing with a test server via Gradle, and GitHub Actions CI workflows that run GameTests headlessly. Includes patterns for mocking registries, testing event handlers, testing commands, and test-driven development for Minecraft projects. Use when the user asks about testing Minecraft mods or plugins, writing GameTests, setting up MockBukkit, or configuring CI for Minecraft projects.

6k tokens scripts
Flutter Testing
MADTeacher

>- Write, fix, review, debug, and validate Flutter tests for apps, packages, and plugins. Use when adding unit tests, widget tests, integration tests, MethodChannel or plugin mocks, Mockito or mocktail test doubles, golden or accessibility checks, CI test commands, test failures, MissingPluginException, pump or pumpAndSettle problems, finder errors, build_runner mock generation, device integration testing, web integration testing, or flaky Flutter tests.

20k tokens scripts
Anchor Repro
lynxlangya

Reproduce a behavioral bug before fixing it, record the failing probe, and verify the fix with the same probe. Use for bug reports, failing tests, crashes, hangs, regressions, wrong output, and observable behavior that should change.

10k tokens scripts
Langgraph Testing Evaluation
Lubu-Labs

Use this skill when you need to test or evaluate LangGraph/LangChain agents: writing unit or integration tests, generating test scaffolds, mocking LLM/tool behavior, running trajectory evaluation (match or LLM-as-judge), running LangSmith dataset evaluations, and comparing two agent versions with A/B-style offline analysis. Use it for Python and JavaScript/TypeScript workflows, evaluator design, experiment setup, regression gates, and debugging flaky/incorrect evaluation results.

39k tokens scripts
Systematic Debugging
itseffi

Four-phase debugging process - root cause first, then fix. Use when encountering bugs or unexpected behavior.

846 tokens
Tdd
itseffi

Test-driven development - write failing test first, then minimal code. Use before implementing any feature or bugfix.

778 tokens
Ab Test Design
Mehdibargach

Design a complete A/B test or experiment from a hypothesis. Takes what you want to test and outputs a full experiment design with hypothesis, metrics, sample size, duration, and success criteria.

952 tokens
Ios Launch Performance
Livsy90

Use this skill when diagnosing iOS app launch performance, startup regressions, first-frame readiness, or early responsiveness. Covers pre-main/dyld work, AppDelegate/SceneDelegate, SwiftUI App startup, launch orchestration, SDK initialization, and launch measurement. Do not use for general performance unless the code runs on the launch path.

59k tokens
A B Test Design
Infrasity-Labs

Design rigorous A/B tests with hypotheses, variants, metrics, and sample size calculations.

445 tokens
Ab Testing
Infrasity-Labs

When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program. Also use when the user mentions "A/B test," "split test," "experiment," "test this change," "variant copy," "multivariate test," "hypothesis," "should I test this," "which version is better," "test two versions," "statistical significance," "how long should I run this test," "growth experiments," "experiment velocity," "experiment backlog," "ICE score," "experimentation program," or "experiment playbook." Use this whenever someone is comparing two approaches and wants to measure which performs better, or when they want to build a systematic experimentation practice. For tracking implementation, see analytics. For page-level conversion optimization, see cro.

6k tokens
Click Test Plan
Infrasity-Labs

Design click/first-click tests to evaluate navigation and information findability.

402 tokens
Design QA Checklist
Infrasity-Labs

Create QA checklists for verifying design implementation accuracy.

512 tokens
Design Sprint Plan
Infrasity-Labs

Plan and facilitate design sprints from challenge framing through prototype testing.

449 tokens
Test Scenario
Infrasity-Labs

Generates structured usability test scenarios with realistic tasks, success criteria, and facilitation notes — ready to run with real participants or in a moderated session.

424 tokens
Usability Test Plan
Infrasity-Labs

Design a usability test plan with tasks, success metrics, participant criteria, and facilitation guide. Use when planning moderated or unmoderated usability testing sessions.

394 tokens
Ux Researcher Designer
Infrasity-Labs

UX research and design toolkit for Senior UX Designer/Researcher including data-driven persona generation, journey mapping, usability testing frameworks, and research synthesis. Use when conducting user research, creating personas, mapping user journeys, planning usability tests, or validating designs.

24k tokens scripts
Agentforce Observe
SalesforceAIResearch

Analyze production Agentforce agent behavior using session traces and Data Cloud. TRIGGER when: user queries STDM session data or Data Cloud trace records; investigates production agent failures, regressions, or performance issues; asks about session traces, conversation logs, or agent metrics; wants to reproduce a reported production issue in preview; runs findSessions or trace analysis queries. DO NOT TRIGGER when: user creates, modifies, or debugs .agent files during development (use agentforce-generate); writes or runs test specs (use agentforce-test); uses sf agent preview for local development iteration; deploys or publishes agents.

37k tokens
Agentforce Test
SalesforceAIResearch

Write, run, and analyze structured test suites for Agentforce agents — functional AND security. TRIGGER when: user writes or modifies test spec YAML (AiEvaluationDefinition); runs sf agent test create, run, run-eval, or results commands; asks about test coverage strategy, metric selection, or custom evaluations; interprets test results or diagnoses test failures; asks about batch testing, regression suites, or CI/CD test integration; requests security testing, OWASP LLM Top 10, red-teaming, penetration testing, prompt-injection tests, a security grade, or a vulnerability assessment of an agent. DO NOT TRIGGER when: user creates, modifies, previews, or debugs .agent files (use agentforce-generate); deploys or publishes agents; writes Agent Script code; uses sf agent preview for development iteration; analyzes production session traces (use agentforce-observe); performs a static safety review of .agent file content (use agentforce-generate Section 15).

49k tokens
Android Reverse Engineering
incogbyte

Decompile Android APK, XAPK, AAB, DEX, JAR, and AAR files using jadx or Fernflower/Vineflower. Reverse engineer Android apps, extract HTTP API endpoints (Retrofit, OkHttp, Volley, GraphQL, WebSocket), trace call flows from UI to network layer, analyze security patterns (cert pinning, exposed secrets, Android Fragment Injection via exported PreferenceActivity), perform dynamic analysis with Frida (adaptive bypass generation, crash analysis, runtime hooking), and — only when the decompiled app contains Google API keys or Firebase configuration — run a conditional Firebase & Google API testing phase (Auth, Realtime DB, Firestore, Remote Config, Storage, Dynamic Links, FCM, Gemini, Maps). Use when the user wants to decompile, analyze, or reverse engineer Android packages, find API endpoints, follow call flows, audit app security, bypass runtime protections, test exposed Google/Firebase credentials, or check for Fragment Injection exposure.

56k tokens scripts