mcpbeat Sign in

Ab Test Setup Skill for Claude

> significance for conversion experiments. Use when setting up an A/B test, calculating sample size, designing an experiment, or analyzing results.

13k tokens
context cost
the whole folder, loaded on every use
6
files
ships runnable scripts
0
copies elsewhere
how many repositories repackaged it
447
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/borghei/Claude-Skills --skill ab-test-setup

What comes with it

45 514 bytes besides the instruction
examples/test_results.csv
references/ab-testing-guide.md
scripts/results_analyzer.py
scripts/sample_size_calculator.py
scripts/test_designer.py

The instruction itself

14 sections, as written by the author

A/B Test Setup Skill

Overview

Production-ready A/B testing toolkit for calculating sample sizes, designing rigorous test plans, and analyzing results with statistical significance testing. Designed for growth teams, product managers, and marketers who need to make data-driven decisions from controlled experiments.

Clarify First

Before designing the test, confirm these inputs. If any is unknown or vague, ASK — do not assume:

  • [ ] Hypothesis + primary metric — what change you expect and the single metric that judges it (drives test plan + analysis)
  • [ ] Baseline conversion rate — the current rate the metric sits at today (drives sample size calculation)
  • [ ] Minimum detectable effect (MDE) — smallest lift worth detecting (drives required samples + duration)
  • [ ] Daily traffic available — eligible visitors per day per variant (determines how long the test must run)

Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions at the top of the artifact.

Quick Start

# Calculate required sample sizes for a test
python scripts/sample_size_calculator.py --baseline 0.05 --mde 0.10 --power 0.80

# Design a complete A/B test plan
python scripts/test_designer.py test_config.json

# Analyze A/B test results
python scripts/results_analyzer.py results.json

Tools Overview

| Tool | Purpose | Input | Output |

|------|---------|-------|--------|

| sample_size_calculator.py | Sample size calculation | Baseline rate, MDE, power | Required samples + duration |

| test_designer.py | Test plan design | JSON test config | Complete test plan document |

| results_analyzer.py | Results analysis | JSON with test results | Statistical analysis + recommendation |

Workflows

Workflow 1: New A/B Test Setup

  • Define hypothesis and success metric
  • Run sample_size_calculator.py with baseline conversion and minimum detectable effect
  • Create test configuration JSON (see Common Patterns)
  • Run test_designer.py to generate complete test plan
  • Share plan with stakeholders for alignment before launch

Workflow 2: Test Results Analysis

  • Collect test results into JSON format
  • Run results_analyzer.py to get statistical significance
  • Review confidence interval, p-value, and effect size
  • Check for segment-level effects if overall result is inconclusive
  • Make ship/no-ship decision based on analysis

Workflow 3: Experimentation Program Review

  • Compile results from multiple past tests
  • Run results_analyzer.py --batch on all results
  • Review win rate, average effect size, and velocity
  • Identify patterns in winning vs losing tests
  • Optimize test pipeline based on learnings

Reference Documentation

See references/ab-testing-guide.md for comprehensive methodology covering:

  • Statistical foundations (z-tests, confidence intervals)
  • Sample size theory and trade-offs
  • Common experimentation pitfalls
  • Multi-variant and sequential testing
  • Bayesian vs frequentist approaches

Common Patterns

Pattern: Test Configuration JSON

{
  "test_name": "Homepage CTA Button Color",
  "hypothesis": "Changing the CTA button from blue to green will increase click-through rate",
  "metric_primary": "cta_click_rate",
  "metric_secondary": ["signup_rate", "bounce_rate"],
  "baseline_rate": 0.045,
  "minimum_detectable_effect": 0.10,
  "significance_level": 0.05,
  "power": 0.80,
  "variants": [
    {"name": "control", "description": "Current blue CTA button"},
    {"name": "treatment", "description": "Green CTA button"}
  ],
  "daily_traffic": 5000,
  "allocation": {"control": 0.50, "treatment": 0.50}
}

Pattern: Test Results JSON

{
  "test_name": "Homepage CTA Button Color",
  "variants": {
    "control": {"visitors": 12500, "conversions": 563},
    "treatment": {"visitors": 12500, "conversions": 625}
  },
  "metric": "cta_click_rate",
  "significance_level": 0.05
}

Quick Reference: Common Effect Sizes

| Context | Small Effect | Medium Effect | Large Effect |

|---------|-------------|---------------|--------------|

| Conversion Rate | 2-5% relative | 5-15% relative | > 15% relative |

| Revenue per User | 1-3% | 3-8% | > 8% |

| Engagement Rate | 3-5% | 5-10% | > 10% |

Other skills for the same job

different authors, same section of the catalogue
Viral Instagram Reels
by vyralcontent
×1

Plan, write, and diagnose Instagram Reels that earn cold-audience reach. Use whenever someone wants a reels script or reels hook for a specific Reel, is debugging why a Reel flopped, wants to know if a draft is worth testing with Trial Reels before going public, or needs a reels caption tuned for the post-hashtag instagram algorithm. Built around what Mosseri has publicly named as the signal hierarchy (watch time, sends per reach, likes per reach), the Trial Reels test-then-publish loop, the Original Content Guidelines and 30-day recovery window, the Edits app, and Reels Insights metrics (skip rate, share rate, followers from this post). Covers a Reels-specific reels strategy: send-driving CTAs, originality without watermarks, audio licensing by account type, captions as the primary SEO signal, and the anti-patterns that quietly cap distribution. Pattern-based guidance, not a virality promise.

18k tokens
Bond Relative Value
by anthropics
vendor

Perform relative value analysis on bonds by combining pricing, yield curve context, credit spreads, and scenario stress testing. Use when analyzing bond richness/cheapness, computing spread decomposition, comparing bonds, assessing bond value vs curves, or running rate shock scenarios.

910 tokens
Returns Analysis
by anthropics
vendor

Build quick IRR/MOIC sensitivity tables for PE deal evaluation. Models returns across entry multiple, leverage, exit multiple, growth, and hold period scenarios. Use when sizing up a deal, stress-testing assumptions, or preparing IC returns exhibits. Triggers on "returns analysis", "IRR sensitivity", "MOIC table", "what's the return at", "model the returns", or "back of the envelope".

837 tokens
Brainstorm Experiments New
by phuryn

Design lean startup experiments (pretotypes) for a new product. Creates XYZ hypotheses and suggests low-effort validation methods like landing pages, explainer videos, and pre-orders. Use when validating a new product idea, creating pretotypes, or testing market demand.

634 tokens
Amazon Alexa QA
by browser-act

Amazon Alexa for Shopping Q&A automation: submits questions to Amazon's Alexa/Rufus AI shopping assistant and collects response text; supports optional keyword search context (navigate to search results page before asking for category-specific answers). Use when user mentions Amazon Alexa, Rufus, Amazon shopping assistant, Amazon AI chat, ask Amazon, Amazon Q&A, automate Alexa questions, Rufus chatbot, Amazon assistant automation, collect Alexa responses, bulk question submission to Amazon, keyword search context, category research. Also applies to extracting Amazon product recommendations from conversational AI, automating repeated queries to Amazon's AI shopping feature, collecting Alexa shopping responses at scale, or market research within a specific product category.

4k tokens scripts
AI Ugc Ads
by tech-leads-club

When the user wants to create UGC ad campaigns, recruit UGC creators, generate AI UGC content, or scale with user-generated content. Also use when the user mentions 'UGC,' 'user-generated content,' 'creator ads,' 'Spark Ads,' 'whitelisting,' 'AI UGC,' 'Arcads,' 'Creatify,' 'creator brief,' or 'UGC testing.' This skill covers the UGC growth framework from creator recruitment through AI-powered scaling. Do NOT use for technical implementation, code review, or software architecture.

5k tokens
Input File Skill
by NVIDIA

Parse, modify, validate, and patch simulator input files. Use when working with reservoir simulation input files, testing scenarios, or validating simulation configurations. This implementation supports reference format (.DATA); other simulators use different extensions (e.g., .afi, .DAT). Supports natural language modifications, keyword patching, and syntax validation.

11k tokens scripts
Recon Scope Triage
by elementalsouls

Triage ASM/recon output for ownership before testing — separate the target's real assets from namespace-collision noise. Automated recon keyword-matches on the brand name, so for any target whose name is a common/dictionary word, the output is dominated by assets belonging to UNRELATED same-named companies (repos, cloud buckets, mobile apps, breach corpora, typosquats). Built from an authorized engagement where an ASM report's "Criticals" were overwhelmingly false positives and the combo/repos/mobile/bucket lists were polluted with unrelated same-named orgs. Use at the START of any engagement, immediately on receiving any ASM/recon/OSINT dataset, BEFORE testing anything.

2k tokens

How to use it

Copy the folder

Take borghei/ab-test-setup from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.