Signup headline test plan
Do we have enough traffic?
Yes, but only just. At 8,000 visitors a week, a 15% relative lift is detectable in about 5 weeks.
| Input | Value |
|---|---|
| Baseline conversion | 4.0% |
| Target (15% relative lift) | 4.6% (+0.6 points absolute) |
| Significance / power | 95% (two-sided) / 80% |
| Sample size | about 18,000 visitors per variant, about 36,000 total |
This uses the standard two-proportion formula. Check it against Evan Miller's calculator before launch. Expect small differences of a few hundred visitors.
How long to run it
- 36,000 visitors at 8,000 a week is 4.5 weeks. Run it for 5 full weeks with a 50/50 split, which is about 40,000 visitors.
- Use whole weeks so every weekday and weekend is represented equally.
- Don't stop early, even if it looks significant at week 2. Peeking inflates false positives.
- Don't extend it either if it's flat at week 5. Call it inconclusive.
Sensitivity to the lift you target:
| Detectable lift | Per variant | Duration |
|---|---|---|
| 20% relative (4.8%) | about 10,300 | about 3 weeks |
| 15% relative (4.6%) | about 18,000 | about 5 weeks |
| 10% relative (4.4%) | about 39,500 | about 10 weeks |
Headline tests often produce lifts under 15%. If the true effect is 10%, this test will probably miss it, so you'd see "no significant difference" even though the headline helps. Only run it if a 15% lift is the smallest result worth acting on. If you want to detect smaller effects, test a bolder change than a wording tweak, such as a different value proposition rather than a rephrase.
Constraints:
- Two variants only. An A/B/C test would need about 7 weeks.
- Counting visitors. The 8,000 must be unique visitors who see the headline, not sessions or page views. If it's sessions, the test will take longer.
What to write down before starting
- Hypothesis, in the skill's format: Because [evidence, e.g. session recordings, survey responses, or ad-to-page message mismatch], we believe [new headline] will cause [higher signup conversion] for [new visitors to the signup page]. We'll know this is true when conversion rises from 4.0% to at least 4.6% at 95% confidence. Without real evidence in the "because" clause, you're guessing.
- Variants. Screenshots of control and variant, with the exact copy. Change only the headline.
- Primary metric. Unique visitor to completed signup, with a precise definition of "completed" (submitted form or verified email). Decide this now.
- Secondary metrics. Signup-form start rate, scroll depth, and time on page. These help explain why it worked or didn't.
- Guardrail metrics. Activation or week-1 retention of signups, and bounce rate. A headline that over-promises can lift signups and attract low-quality users. Define the stop rule: halt if a guardrail is significantly negative.
- Test parameters.
- Sample size per variant (about 18,000), power, alpha, and the MDE (15% relative).
- Fixed end date, about 5 weeks out.
- 50/50 split with persistent assignment, so returning visitors see the same variant.
- Population, including whether to exclude internal and bot traffic.
- Decision rules.
- Ship if the variant is significant and positive with no guardrail harm.
- Keep control if it's significant and negative.
- If it's flat, treat it as inconclusive, not as evidence of no effect.
- Note which segments you'll look at (device, new vs. returning, traffic source). Fix these now to avoid cherry-picking later.
- Implementation notes.
- Server-side or client-side. Client-side can flicker on a headline, so prefer server-side or an anti-flicker setup.
- Tracking verified on both variants.
- A sample ratio mismatch (SRM) check planned. The split should land close to 50/50.
- Risks and external factors. Planned campaigns, launches, seasonality, or traffic-source changes during the 5 weeks. Log anything that happens. Avoid adding new traffic sources mid-test.
- Pre-launch QA. Check both variants on mobile and desktop, and confirm events fire.
When it finishes, record the outcome in the playbook format: result with CI and p-value, guardrails, segment deltas, the reusable pattern, and status.
If you give me the actual control headline and candidate variant, I can draft the hypothesis and a few evidence-based headline options.
Real run in an isolated folder with only this skill installed. Only the Skill and Read tools were enabled, so nothing was fetched from the web and no file was written. The example is fictional. The traffic numbers are invented.
Was es macht
Hilft beim Entwurf vertrauenswürdiger Experimente. Enthalten sind eine Hypothesenvorlage („Weil X, glauben wir, dass Y bei Zielgruppe W Z bewirkt, und wir erkennen das, wenn sich Metrik M bewegt“), Testarten (A/B, A/B/n, multivariat, Split-URL), eine Schnelltabelle zur Stichprobengröße nach Ausgangswert und erwarteter Steigerung, primäre, sekundäre und Guardrail-Metriken, was variiert wird, Traffic-Verteilung, Client- und Server-seitige Umsetzung, eine Checkliste vor dem Start, das Peeking-Problem und eine Auswertungs-Checkliste mit Deutung jedes Ergebnisses. Ein zweiter Teil macht aus Einzeltests ein Programm: Experimentschleife, ICE-Priorisierung, Zielwerte für Tempo, eine Playbook-Vorlage und ein Rhythmus aus wöchentlich, zweiwöchentlich, monatlich und quartalsweise. Referenzen enthalten Testvorlagen und einen Leitfaden zur Stichprobengröße.
So funktioniert es
- Es fragt nach Ausgangsrate, Traffic, der Änderung und der kleinsten sinnvoll nachweisbaren Steigerung.
- Es prüft, ob der Traffic reicht, und schätzt die Dauer.
- Es schreibt Hypothese, Metriken und das, was vor dem Start zu dokumentieren ist.
Geeignet für
Alle, die entscheiden wollen, ob eine Änderung besser ist, und Teams, die regelmäßig experimentieren wollen.
Geringes Risiko: Reine Anweisungen ohne Skripte, ohne Netzwerkzugriff und ohne Dateischreiben. Die Zahlen zur Stichprobengröße und die Programmziele (etwa 4-8 Experimente pro Monat, 20-30 % Erfolgsquote) sind Faustregeln; für echte Entscheidungen nutzen Sie einen richtigen Rechner (der Skill verlinkt zwei öffentliche). Er führt keine Tests aus und liest Ihre Analysedaten nicht, kann also nichts über das Ergebnis eines echten Tests sagen. Einige Links zeigen auf Schwester-Skills und Werkzeug-Leitfäden, die nicht in diesem Paket sind. Einmal mit erfundenen Traffic-Zahlen ausprobiert.