Izveidot kontuCreate account
‹ All playbooks
Run a growth experiment with a stop rule

growth.run-experiment·version 1.0.0·draft1 to verify

Run a growth experiment with a stop rule

A change to the site or funnel is tested against a written hypothesis and a stop rule, and the result is recorded either way.

MartaGrowth Leadruns itProfile ›
Whenon request — owner or employee wants to test a change — headline, price display, form, ad, onboarding step
Who actsthe agent acts after approval
Time45 min active; 1–4 weeks running depending on volume
Countryany country
Sign in to run thisThis playbook opens inside Brain Club. Sign in to read and run it.

When to use

The company wants to know whether a specific change improves a number, and is willing to accept "no" as an answer. Not for a one-off campaign or content push — use growth.social-post-weekly or growth.newsletter. Not for an analytics setup itself — use tech.install-analytics first; an experiment without working tracking is not run at all. If the idea came from a competitor scan, start from growth.competitor-scan output, but this playbook still applies.

Before you start

  • The primary metric is named and its current baseline is read from site-analytics (not from memory).
  • Tracking for that metric already works — an event or pageview exists and fires today.
  • Someone can change the site or channel (the agent prepares, a person or agent with deploy access ships).
  • If the experiment touches prices, terms or personal data handling, flag it in S3 — that may need a

What a run requires3

  • Approval · S3 · ownerthe run stops until a named person records the decision
  • Approval · S6 · ownerthe run stops until a named person records the decision
  • Approval · S9 · ownerthe run stops until a named person records the decision

Any step can wait until a date and reopens by itself; every closed step leaves evidence (a note, a link, a number).

The trail9 steps

  1. Frame the problemagent

    Write one sentence: what is observed, where, and what the change would be. Note anything that could confound it (a campaign, a season, a price change in the same period).

    Done when the sentence and the confound list exist in the experiment record.

  2. Write the hypothesis and the stop ruleagent

    Format: "If we <change>, then <metric> moves from <baseline> to <target>, measured over <at least N weeks or M observations>." Write the stop rule: the date or volume at which the test ends no matter what the interim numbers look like, and the rule for stopping early only if the variant is clearly hurting (state the threshold now).

    Done when hypothesis, metric, baseline, end date and stop rule are all written down.

  3. Approve the hypothesisownerneeds approval · owner

    Approval · S3 · owner — the run stops until a named person records the decision

    Show hypothesis, metric, baseline, cost of running it (time, any spend), and what happens if it loses.

    Done when the owner has approved or amended the hypothesis, and the end date is fixed.

  4. Check the test can actually concludeagent

    From site-analytics, estimate how many observations per week the page or channel gets. If the volume cannot plausibly reach a readable difference within 4–6 weeks, say so and propose a bigger change or a different page — do not launch a test that cannot end in a decision.

    Done when the record states the expected observations and the agent's verdict: runnable / not runnable.

  5. Instrument and verify tracking BEFORE launchagent

    Add or confirm the events in site-analytics for both control and variant. Trigger each event yourself and read it back in the reports.

    Done when every event the hypothesis needs fires and appears in the analytics within minutes.

    ⛔ A test launched on unverified tracking is void — restart the clock after the fix.

  6. Launchagentneeds approval · owner

    Approval · S6 · owner — the run stops until a named person records the decision

    Ship the variant (site change, ad variant, copy) with the split as even as the tooling allows. Record the exact launch time — it starts the clock on the stop rule.

    Done when the variant is live and the launch time is in the record.

  7. Monitor, do not peek-decideagent

    Check weekly that both variants receive traffic and events and that tracking is still firing. Do not call a winner mid-run; the stop rule from S2 governs.

    Done when each weekly check is logged with traffic counts per variant and a one-line "tracking OK / broken".

  8. Close and analyseagent

    At the end date, pull the numbers for both variants: primary metric, plus any secondary metric already named in S2 (never invented after seeing the data). State the difference plainly and whether it meets the target from S2.

    Done when the result — including "no difference" — is written in the experiment record.

  9. Decide, then actownerneeds approval · owner

    Approval · S9 · owner — the run stops until a named person records the decision

    Present the result with a recommendation: adopt the variant, reject it, or rerun with a change (and why). The decision is executed the same week — winner shipped, loser removed — and linked via bc tasks add or bc goals add to the goal it serves.

    Done when the site or channel matches the decision, and the record shows the decision, who made it, and when.

Checks — how we know it worked

  • The record contains a baseline read before launch, not after.
  • The end date in S2 and the launch time in S6 give a test length that matches the S4 estimate.
  • Both variants received traffic every week of the run (no variant silently at zero).
  • The result names the primary metric and the end date — a result without an end date is not a result.
  • The site after S9 matches the decision, whichever way it went.

If it goes wrong

SymptomResponse
Tracking breaks mid-runLog the gap with dates; if the gap covers more than a third of the run, extend the end date once, say so in the record, and never quietly.
Interim numbers look like a big winDo not stop. The stop rule exists for this; early winners are the most common false positive.
Not enough traffic to concludeClose it as "inconclusive — volume too low" in S8; that is a valid result, not a failure. Propose a bigger change or a higher-traffic page.
Variant hurts clearly before the end dateApply only the pre-written early-stop threshold from S2; if none was written, ride to the end date or shorten it once with the owner's ok.
Result contradicts last experiment on the same pageCheck whether the audience or season differed; note both records as related before deciding.

What each step leaves behind

  1. S1the sentence and the confound list exist in the experiment record.
  2. S2hypothesis, metric, baseline, end date and stop rule are all written down.
  3. S3the owner has approved or amended the hypothesis, and the end date is fixed.
  4. S4the record states the expected observations and the agent's verdict: runnable / not runnable.
  5. S5every event the hypothesis needs fires and appears in the analytics within minutes.
  6. S6the variant is live and the launch time is in the record.
  7. S7each weekly check is logged with traffic counts per variant and a one-line "tracking OK / broken".
  8. S8the result — including "no difference" — is written in the experiment record.
  9. S9the site or channel matches the decision, and the record shows the decision, who made it, and when.

Evidence to keep

The experiment record: hypothesis, baseline, stop rule, launch time · weekly monitoring log · final numbers per variant with the pull date · the S9 decision and who made it · links to the task/goal and, if adopted, to the change that shipped.

How this playbook improves

After every 5 finished experiments ask: how many ended in a written decision vs quietly abandoned? How many results were "no difference", and did those pages get bigger changes proposed? Did any test run past its end date, and which step let it slip? Did any launch ship with unverified tracking? A new version changes the step that let the slip happen, and says so in its change note.