
Use this when
- Next year’s plan is a stack of untested ICP, message, and channel bets.
- Campaigns launch at full budget because “we aligned in Q4.”
- Sales, marketing, and product each have a different story about what worked last quarter, and none of it is written down.
- Leadership wants more tests, and the calendar is already a pile of one-off ads with no decision attached.
Do not use this when
- There is no ICP and no primary motion. Stay in ICP and channel strategy.
- You need a page a champion can forward. Write content strategy first; then test distribution.
- Legal or privacy forbids the treatment. This page will not bless dark patterns or unconsented lists.
- The request is “prove marketing with a multi-touch model.” That is a different job.
A few useful terms
Keep this in mind
One assumption per test, and a decision the result is allowed to change. “We think this might work” is not a hypothesis. A test that cannot alter the plan is a campaign with extra slides. Do not A/B-test a strategy question (which ICP, which motion) as if it were a headline.How to do it
Step 1: Write an assumption you can test
Pull the hidden bets out of the annual plan: who the buyer is, which pain they will move for, which words they use, which channel they will answer. Isolate one. Shapes that work:- Segment: [role] in [industry] will [book / click / reply] more on [this offer] than [that adjacent role].
- Message: language about [problem A] will beat [problem B] on [metric] in [this list].
- Channel: [channel] will beat [channel] for [this segment] on [conversion to a sales-usable next step]—not on vanity reach.
Step 2: Design a small, useful test
Answer, on one card:- Decision this informs (ICP list, campaign narrative, spend, enablement).
- Success metric that a seller would recognize (meeting, qualified conversation)—not “engagement” unless that is the actual decision.
- Who you are testing (a named slice, not “the market”).
- How small you can go and still believe the comparison (two variants, one audience, a clock).
- When you will stop (a date, not “until it works”).
Step 3: Record the result and its meaning
Collect the number and what sales heard. A higher CTR that produces worse meetings is a failed hypothesis if the decision was “which message we put in the pitch.” When it ends, write four lines: what we tested, what happened, what we now believe, which artifact changes (ICP tier, messaging hierarchy, channel mix, lead scoring inputs). Tag the learning: segment / message / channel / motion. Do not let the result live in a Slack screenshot.Step 4: Expand only after reviewing the evidence
Scaling is not “do more of everything.” Promote a winning message into the core narrative. Promote a winning slice into tier criteria. Move spend toward a channel that produced the sales-usable step cheaper. Feed a behavioral signal into scoring only if you will inspect it—see lead scoring. A failed test that kills a bad ICP story is a win. Celebrate learning velocity in the review, not only the variant that “won.”Step 5: Keep a shared experiment log
Marketing, sales, and product put hypotheses in the same ledger. Airtable vs Notion vs a Sheet is a tooling choice; the requirement is one list of in-flight and closed tests. If MarTech governance later buys a testing SKU, it still writes into this log. Hold a weekly experiment review that only asks: what closed, what did we believe, which artifact changed. That is how a growth mindset shows up. A slide of “tests launched” is not the review. Name each bet as engine (a loop that can compound), lubricant (makes the engine cheaper), or turbo (a one-off). A calendar of only turbos is a launch habit, not a system. Channel strategy still names the primary motion; this page refuses to optimize a motion you have not chosen.Worked example (illustrative)
Sales-assist cybersecurity. Not your vertical.Copy: experiment card (fill)
- Assumption we are willing to be wrong about:
- If true we will ______; if false we will stop ______:
- Audience slice:
- Treatment vs contrast:
- Sales-usable success metric:
- Start / stop dates:
- Owner:
- Result (number + what sales heard):
- Artifact we updated (ICP / message / channel / scoring / none):
- Tag (segment / message / channel / motion):
Before you start
- The hypothesis names one assumption and a falsifier.
- A GTM decision is written before launch.
- The metric is a next step sales would use, or an explicit diagnostic.
- Scope and stop date exist.
- The log is shared; this test has a row before it spends.
- We will not scale “everything that moved a vanity chart.”
- Privacy and list source are allowed for this treatment.
Metrics
Do not count experiments launched, or a decorated “innovation” slide, as a testing system.
Common mistakes
- Vague hunches dressed as hypotheses.
- Tests that cannot change the plan.
- Six-month campaigns called experiments.
- Optimizing CTR while the meeting quality collapses.
- Scaling volume because one cell was green.
- No shared log, so Q3 repeats Q1.
- Copying another consultancy’s “2–3 tests per quarter” or vertical examples as your calendar.
- Using an LLM citation score as the success metric for a message test—see SEO and AEO for that job.
What to read next
Whether the year can even close is GTM planning. Which AI problem deserves a bet before you open a test row is AI use-case selection. Which pages exist before you test distribution is content strategy. Which motion is primary is channel strategy. How buyers find you in search and answers is SEO and AEO. Routing a validated signal is lead scoring.Sources and evidence boundary
This is an owner-maintained operating synthesis. It is not a causal-inference textbook and not a promise that small tests replace a strategy. Hypothesis → small informative test → classify the learning → feed ICP, message, channel, and scoring is distilled from a public GTM testing essay (Heinz Marketing, Win Dean-Salyards, undated 2026 planning post). That essay is a method prompt, not a source to copy. Vertical examples, suggested quarterly test counts, and vendor scoring tools named there are not this library’s calendar or a requirement to buy a platform. The weekly review as the growth-mindset ritual, the warning not to A/B-test strategy questions, and the engine / lubricant / turbo split (Racecar) draw on Elena Verna’s framework essays (9 favorite growth frameworks, 2024-10-25; weekly experiment review (subscription may be required)). Racecar as originally written by Lenny Rachitsky and Dan Hockenmaier stays with those authors. None of those pieces are a command to hire a growth squad or to copy Dropbox’s loop.Copyright © 2026 Ivan Xu. All rights reserved. See the copyright and reuse terms. Canonical source: github.com/weilun88313/B2B-Playbook