Khi nguoi dung can thiet lap A/B test cho ads, landing page, email, hoac product. Cung dung khi nguoi dung nhac 'A/B test', 'split test', 'test creative', 'test copy', 'test landing', 'test gia', 'kiem dinh thong ke', 'sample size'. Skill nay giup chon bien can test, tinh sample size, setup tracking, va phan tich ket qua thong ke (significance) — phu hop cho team khong phai data scientist.
npx skills add https://github.com/minhnv0807/ai-business-skills --skill 19-ab-test-setup
Doc .agents/product-marketing-context.md neu co.
Quy tac cot loi. Neu test 2 bien cung luc → khong biet bien nao tac dong.
Format: "Neu lam [X], metric [Y] se tang [Z%]"
Khong duoc dung test som.
Test tung tuan tron (khong 3 ngay, khong 10 ngay) — khach mua hanh vi khac nhau cac ngay.
Khong check ket qua moi gio. Dung ket qua cho den khi het thoi han + du sample.
Chi tin khi p-value < 0.05. Neu khong: ngau nhien, khong phai do bien cua ban.
Ghi lai:
→ De sau nay ban khong test lai cai da test.
Sample size/variant =
16 × p × (1-p) / MDE²
Trong do:
p = Baseline conversion rate hien tai (vi du 0.02 = 2%)
MDE = Minimum Detectable Effect (muc do thay doi toi thieu ban muon phat hien, vi du 0.2 = 20% tang)
Landing page hien co conv rate 3%, muon phat hien thay doi 20%+:
=> Can 25,866 visit total, neu 500 visit/ngay → 52 ngay.
Landing page hien conv rate 10%, muon phat hien 20%+:
=> 7,200 visit total, 500/ngay → 14 ngay (kha thi).
| Traffic/ngay | Conv rate | Thoi gian can | Co nen test? |
|-------------|-----------|--------------|-------------|
| < 100 visit | Bat ky | 2+ thang | Khong — traffic thap |
| 100-500 | 2-5% | 3-6 tuan | Co, nhung kien nhan |
| 500-2K | 2-5% | 2-3 tuan | Co — ly tuong |
| 2K+ | Bat ky | 1-2 tuan | Co — nhanh |
Neu traffic thap (< 100/ngay) → khong test, thay vao do tap trung tang traffic truoc.
Tac dong cao nhat. 80% nguoi doc headline, 20% doc body → headline thay doi toi uu = bien nho nhat.
Idea:
Do thay doi duoc de test. Thuong tang 5-25%.
Test:
Test:
Test:
Test:
Test:
Test:
Test:
| Tool | Phu hop | Chi phi |
|------|---------|---------|
| Meta Ads A/B test built-in | Test ads (creative/audience/placement) | Free |
| TikTok Ads Split test | Test TikTok ads | Free |
| Google Optimize (deprecated) | Da ngung — dung thay the: VWO, Convert.com | Tu $199/thang |
| PostHog | Test product features, landing | Free tier OK |
| Unbounce / Instapage | Landing page built-in A/B | $90+/thang |
| Manual (split URL) | Don gian, 50/50 URL redirect | Free |
1. Tao 2 phien ban landing page: /page-a va /page-b
2. Split traffic 50/50:
- Qua Meta Ads: 2 ad set cung audience, khac link dich
- Qua Google Ads: 2 ad cung traffic, khac link
- Qua organic: dung URL rewrite (rare)
3. Tracking conversion rieng tung page:
- Meta Pixel: event "Lead" voi custom parameter "page_version" = A/B
- GA4: UTM campaign=test-header, utm_content=a/b
4. Sau 14 ngay + du sample, export data
Dung calculator online: https://www.evanmiller.org/ab-testing/chi-squared.html
Input:
Output:
| p-value | Uplift | Ket luan |
|---------|--------|---------|
| < 0.05 | > 5% | B thang — implement B |
| < 0.05 | < 5% | Significant nhung nho — can nhac cost |
| 0.05-0.10 | > 10% | Ket qua borderline — keo dai test |
| > 0.10 | Bat ky | Khong significant — giu A hoac test tiep |
# A/B Test: [Ten test]
Ngay tao: [YYYY-MM-DD]
Chu so huu: [Ten nguoi]
## 1. Gia thuyet
"Neu [thay doi X], [metric Y] se tang [Z%]"
## 2. Bien test
- Variant A (Control): [Mo ta hien tai]
- Variant B (Challenger): [Mo ta thay doi]
- Chi thay doi: [1 thing]
## 3. Metric
- Primary: [VD: Conversion rate]
- Secondary: [VD: Bounce rate, Time on page]
## 4. Sample size & Duration
- Baseline: [p = %]
- MDE: [% change muon phat hien]
- Sample/variant: [so]
- Traffic/ngay: [so]
- Thoi gian can: [so ngay]
## 5. Setup
- Tool: [Ten]
- URL A: [...]
- URL B: [...]
- Tracking: [Meta Pixel / GA4 / PostHog events]
- Split: 50/50 (hoac khac ghi ro)
## 6. Timeline
- Start: [ngay]
- End: [ngay]
- Review: [ngay]
## 7. Ket qua (dien sau khi ket thuc)
| Variant | Visitors | Conversions | Rate | Uplift vs A |
|---------|----------|-------------|------|-------------|
| A | | | | — |
| B | | | | +X% |
p-value: [x]
Significant (p<0.05): [Yes/No]
## 8. Ket luan
[B thang / A thang / Khong significant]
## 9. Hanh dong
[Implement / Test tiep / Rollback]
## 10. Bai hoc
[Dieu hoc duoc → ap dung cho test sau]
Plan, write, and diagnose Instagram Reels that earn cold-audience reach. Use whenever someone wants a reels script or reels hook for a specific Reel, is debugging why a Reel flopped, wants to know if a draft is worth testing with Trial Reels before going public, or needs a reels caption tuned for the post-hashtag instagram algorithm. Built around what Mosseri has publicly named as the signal hierarchy (watch time, sends per reach, likes per reach), the Trial Reels test-then-publish loop, the Original Content Guidelines and 30-day recovery window, the Edits app, and Reels Insights metrics (skip rate, share rate, followers from this post). Covers a Reels-specific reels strategy: send-driving CTAs, originality without watermarks, audio licensing by account type, captions as the primary SEO signal, and the anti-patterns that quietly cap distribution. Pattern-based guidance, not a virality promise.
Perform relative value analysis on bonds by combining pricing, yield curve context, credit spreads, and scenario stress testing. Use when analyzing bond richness/cheapness, computing spread decomposition, comparing bonds, assessing bond value vs curves, or running rate shock scenarios.
Build quick IRR/MOIC sensitivity tables for PE deal evaluation. Models returns across entry multiple, leverage, exit multiple, growth, and hold period scenarios. Use when sizing up a deal, stress-testing assumptions, or preparing IC returns exhibits. Triggers on "returns analysis", "IRR sensitivity", "MOIC table", "what's the return at", "model the returns", or "back of the envelope".
Design lean startup experiments (pretotypes) for a new product. Creates XYZ hypotheses and suggests low-effort validation methods like landing pages, explainer videos, and pre-orders. Use when validating a new product idea, creating pretotypes, or testing market demand.
Amazon Alexa for Shopping Q&A automation: submits questions to Amazon's Alexa/Rufus AI shopping assistant and collects response text; supports optional keyword search context (navigate to search results page before asking for category-specific answers). Use when user mentions Amazon Alexa, Rufus, Amazon shopping assistant, Amazon AI chat, ask Amazon, Amazon Q&A, automate Alexa questions, Rufus chatbot, Amazon assistant automation, collect Alexa responses, bulk question submission to Amazon, keyword search context, category research. Also applies to extracting Amazon product recommendations from conversational AI, automating repeated queries to Amazon's AI shopping feature, collecting Alexa shopping responses at scale, or market research within a specific product category.
When the user wants to create UGC ad campaigns, recruit UGC creators, generate AI UGC content, or scale with user-generated content. Also use when the user mentions 'UGC,' 'user-generated content,' 'creator ads,' 'Spark Ads,' 'whitelisting,' 'AI UGC,' 'Arcads,' 'Creatify,' 'creator brief,' or 'UGC testing.' This skill covers the UGC growth framework from creator recruitment through AI-powered scaling. Do NOT use for technical implementation, code review, or software architecture.
Parse, modify, validate, and patch simulator input files. Use when working with reservoir simulation input files, testing scenarios, or validating simulation configurations. This implementation supports reference format (.DATA); other simulators use different extensions (e.g., .afi, .DAT). Supports natural language modifications, keyword patching, and syntax validation.
Triage ASM/recon output for ownership before testing — separate the target's real assets from namespace-collision noise. Automated recon keyword-matches on the brand name, so for any target whose name is a common/dictionary word, the output is dominated by assets belonging to UNRELATED same-named companies (repos, cloud buckets, mobile apps, breach corpora, typosquats). Built from an authorized engagement where an ASM report's "Criticals" were overwhelmingly false positives and the combo/repos/mobile/bucket lists were polluted with unrelated same-named orgs. Use at the START of any engagement, immediately on receiving any ASM/recon/OSINT dataset, BEFORE testing anything.
Take minhnv0807/19-ab-test-setup from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.