홈 / Skills / 브랜드·마케팅 / A/B 테스트와 실험 (A/B Testing)
브랜드·마케팅

A/B 테스트와 실험 (A/B Testing)

Corey HainesMITSkill 언어: 영어
라이선스 확인됨스크립트 없음낮은 위험
데모실행 기록 · claude-sonnet-5-5 · 2026년 10월
Using the ab-testing skill, plan a test for the signup page headline of a fictional app. Baseline signup conversion is 4%, the page gets about 8,000 visitors a week, and we want to detect a 15% relative lift. Tell me whether we have enough traffic, how long to run it, and what to write down before we start.
Skill: ab-testing

Signup headline test plan

Do we have enough traffic?

Yes, but only just. At 8,000 visitors a week, a 15% relative lift is detectable in about 5 weeks.

Input Value
Baseline conversion 4.0%
Target (15% relative lift) 4.6% (+0.6 points absolute)
Significance / power 95% (two-sided) / 80%
Sample size about 18,000 visitors per variant, about 36,000 total

This uses the standard two-proportion formula. Check it against Evan Miller's calculator before launch. Expect small differences of a few hundred visitors.

How long to run it

  • 36,000 visitors at 8,000 a week is 4.5 weeks. Run it for 5 full weeks with a 50/50 split, which is about 40,000 visitors.
  • Use whole weeks so every weekday and weekend is represented equally.
  • Don't stop early, even if it looks significant at week 2. Peeking inflates false positives.
  • Don't extend it either if it's flat at week 5. Call it inconclusive.

Sensitivity to the lift you target:

Detectable lift Per variant Duration
20% relative (4.8%) about 10,300 about 3 weeks
15% relative (4.6%) about 18,000 about 5 weeks
10% relative (4.4%) about 39,500 about 10 weeks

Headline tests often produce lifts under 15%. If the true effect is 10%, this test will probably miss it, so you'd see "no significant difference" even though the headline helps. Only run it if a 15% lift is the smallest result worth acting on. If you want to detect smaller effects, test a bolder change than a wording tweak, such as a different value proposition rather than a rephrase.

Constraints:

  • Two variants only. An A/B/C test would need about 7 weeks.
  • Counting visitors. The 8,000 must be unique visitors who see the headline, not sessions or page views. If it's sessions, the test will take longer.

What to write down before starting

  1. Hypothesis, in the skill's format: Because [evidence, e.g. session recordings, survey responses, or ad-to-page message mismatch], we believe [new headline] will cause [higher signup conversion] for [new visitors to the signup page]. We'll know this is true when conversion rises from 4.0% to at least 4.6% at 95% confidence. Without real evidence in the "because" clause, you're guessing.
  2. Variants. Screenshots of control and variant, with the exact copy. Change only the headline.
  3. Primary metric. Unique visitor to completed signup, with a precise definition of "completed" (submitted form or verified email). Decide this now.
  4. Secondary metrics. Signup-form start rate, scroll depth, and time on page. These help explain why it worked or didn't.
  5. Guardrail metrics. Activation or week-1 retention of signups, and bounce rate. A headline that over-promises can lift signups and attract low-quality users. Define the stop rule: halt if a guardrail is significantly negative.
  6. Test parameters.
    • Sample size per variant (about 18,000), power, alpha, and the MDE (15% relative).
    • Fixed end date, about 5 weeks out.
    • 50/50 split with persistent assignment, so returning visitors see the same variant.
    • Population, including whether to exclude internal and bot traffic.
  7. Decision rules.
    • Ship if the variant is significant and positive with no guardrail harm.
    • Keep control if it's significant and negative.
    • If it's flat, treat it as inconclusive, not as evidence of no effect.
    • Note which segments you'll look at (device, new vs. returning, traffic source). Fix these now to avoid cherry-picking later.
  8. Implementation notes.
    • Server-side or client-side. Client-side can flicker on a headline, so prefer server-side or an anti-flicker setup.
    • Tracking verified on both variants.
    • A sample ratio mismatch (SRM) check planned. The split should land close to 50/50.
  9. Risks and external factors. Planned campaigns, launches, seasonality, or traffic-source changes during the 5 weeks. Log anything that happens. Avoid adding new traffic sources mid-test.
  10. Pre-launch QA. Check both variants on mobile and desktop, and confirm events fire.

When it finishes, record the outcome in the playbook format: result with CI and p-value, guardrails, segment deltas, the reusable pattern, and status.

If you give me the actual control headline and candidate variant, I can draft the hypothesis and a few evidence-based headline options.

Real run in an isolated folder with only this skill installed. Only the Skill and Read tools were enabled, so nothing was fetched from the web and no file was written. The example is fictional. The traffic numbers are invented.

하는 일

믿을 수 있는 실험을 설계하도록 돕습니다. 가설 템플릿("X 때문에 Y가 대상 W에게 Z를 일으킬 것이라 믿으며, 지표 M이 움직이면 맞다고 본다"), 테스트 유형(A/B, A/B/n, 다변량, URL 분할), 기준선과 기대 개선폭별 표본 크기 빠른 표, 주요·보조·가드레일 지표, 무엇을 바꿀지, 트래픽 배분, 클라이언트 측과 서버 측 구현, 출시 전 체크리스트, '중간에 들여다보기' 문제, 결과별 해석이 있는 분석 체크리스트를 다룹니다. 후반부는 단발성 테스트를 운영 체계로 만듭니다. 실험 루프, ICE 우선순위, 실험 속도 목표, 플레이북 항목 템플릿, 주간·격주·월간·분기 운영 리듬입니다. 참고 문서는 테스트 템플릿과 표본 크기 가이드입니다.

동작 방식

  1. 기준 전환율, 트래픽, 바꿀 내용, 감지하고 싶은 최소 개선폭을 묻습니다.
  2. 트래픽이 충분한지 확인하고 필요한 기간을 추정합니다.
  3. 가설, 지표, 시작 전에 기록할 내용을 작성합니다.

이런 때 좋습니다

변경이 정말 더 나은지 판단하려는 사람, 꾸준히 실험하는 습관을 만들려는 팀.

참고 및 위험

낮은 위험:스크립트가 없는 지침 패키지로, 네트워크 접속이나 파일 쓰기가 없습니다. 표본 크기 수치와 운영 목표(월 4~8개 실험, 승률 20~30% 등)는 경험칙이므로 실제 결정에는 제대로 된 계산기를 쓰세요(스킬에 공개 계산기 링크 두 개가 있습니다). 테스트를 실행하거나 분석 데이터를 읽지 않으므로 실제 테스트의 결과는 알려 줄 수 없습니다. 일부 링크는 이 패키지에 없는 자매 스킬과 도구 가이드를 가리킵니다. 가상의 트래픽 수치로 한 번 시험 실행했습니다.