By following this guide, you will set up a repeatable email A/B testing workflow that generates statistically reliable insights, not just directional guesses. The full process takes roughly two to three hours to configure the first time, then under thirty minutes per test after that.
What You'll Build
- A documented hypothesis template that forces clear thinking before each test
- A segmentation and sample-size calculation method tied to your real list size
- A consistent naming and tagging convention so results stay organised across campaigns
- A results log in Notion or Google Sheets that accumulates learnings over time
- A decision framework for when to ship a winner and when to retest
Prerequisites
- An email platform that supports A/B testing natively (Klaviyo, Loops, Mailchimp, HubSpot, or ActiveCampaign all work)
- A subscriber list of at least 1,000 contacts (smaller lists require adjusted expectations, covered in Step 3)
- Basic familiarity with your platform's campaign builder
- A spreadsheet tool or Notion for logging results
Step 1: Write a Testable Hypothesis Before You Touch the Platform
Most email A/B tests fail to teach you anything because they start with a vague idea rather than a structured hypothesis. A hypothesis forces you to commit to what you expect to happen and why.
Use this sentence structure for every test:
If we change [element] from [control version] to [variant version],
then [metric] will [increase/decrease] by [estimated amount]
because [reason based on customer behaviour or prior data].
A real example for a SaaS company targeting small businesses in Australia might read: "If we change the subject line from 'Your free trial ends soon' to 'Your account closes in 48 hours', then open rate will increase by at least 4 percentage points because urgency framing with a specific time window outperforms generic trial language in our segment."
The reason matters. If you cannot write a reason, you are guessing. Guessing produces data you cannot act on.
Common pitfall: Testing two completely different subject lines that change multiple variables at once. If the winner is longer, uses a number, and drops emoji while the loser is short, casual, and uses emoji, you have no idea what drove the result.
Step 2: Choose One Variable and Lock Everything Else
Your test should isolate a single element. Here are the most reliable elements to test in order of typical impact:
- Subject line (highest leverage, easiest to test, fastest results)
- Preview text (often overlooked, significant on mobile)
- Send time or day (requires at least three weeks of data to be meaningful)
- CTA button copy (affects click-through rate, not open rate)
- Email length or content structure (slowest to produce clean data)
Pick one. Write your control version (A) and your variant version (B). Document both in your hypothesis template before creating anything in your email platform.
Pro tip: Subject line and preview text work as a pair. If you change the subject line, freeze the preview text and vice versa. Many marketers accidentally change both and contaminate their results.
Step 3: Calculate Your Sample Size
Why does sample size matter so much?
Sending to too small a sample produces results that look decisive but are actually noise. A difference of 2 percentage points in open rate on a 200-person test is statistically meaningless. The same difference on a 5,000-person test might be real.
Use a free sample size calculator. Evan Miller's calculator at evanmiller.org is the standard reference. For most email marketers, these inputs are reasonable starting points:
- Baseline conversion rate: your current average open rate (or click rate for CTA tests)
- Minimum detectable effect: 3 to 5 percentage points for open rate, 1 to 2 percentage points for click rate
- Statistical significance: 95%
- Statistical power: 80%
For a list with a 30% average open rate where you want to detect a 4-point improvement, you typically need around 1,400 subscribers per variant. That means 2,800 total for a two-variant test, which is well within reach for most businesses with lists above 5,000.
What if your list is under 1,000 contacts?
Treat results as directional only. Run the test, record the outcome, but do not declare a winner with confidence. Accumulate three to four tests on the same variable before making a permanent change to your template or strategy.
Step 4: Set Up the Test in Your Email Platform
The mechanics differ slightly by platform, but the principles are consistent. The following uses Klaviyo as a reference since it is widely used by e-commerce and SaaS teams in Canada, Singapore, and Australia as of 2026.
- Create a new campaign and select A/B test mode.
- Set variant A to your control. Set variant B to your single changed element.
- Set the split to 50/50. Do not use a 70/30 or 80/20 split unless you have a specific reason tied to list size constraints.
- Set the winner metric to match your hypothesis. Open rate for subject line tests. Click rate for CTA tests.
- Set the evaluation window to at least 4 hours for open rate tests. Some platforms default to 1 hour, which is too short for audiences spread across time zones.
- Disable automatic winner sending for your first five tests. Review results manually so you build the habit of reading the data rather than outsourcing the decision.
Common pitfall: Letting the platform auto-send the winner to the remaining list before the evaluation window closes. Set a calendar reminder to review results at least 30 minutes before the window ends.
Step 5: Name and Tag Every Test Consistently
After ten campaigns, an inconsistent naming convention makes it nearly impossible to find historical results. Use a convention like this:
[DATE]_[CAMPAIGN TYPE]_[VARIABLE TESTED]_[TEST ID]
Example:
2026-09_PROMO_SUBJECT-LINE_T012
In Klaviyo and Mailchimp, you can add custom tags or labels. Tag every A/B test with "ab-test" plus the variable name (e.g. "ab-test-subject"). This makes filtering your campaign history trivial when you want to review patterns across six months.
Store the same naming convention in your results log so each row ties back directly to the campaign in your platform.
Step 6: Record Results in a Shared Log
Create a results log in Notion or Google Sheets with these columns:
| Test ID | Date | Variable | Control | Variant | Sample Size | Control Result | Variant Result | Winner | Confidence | Hypothesis Confirmed? | Notes |
|---|---|---|---|---|---|---|---|---|---|---|---|
| T012 | 2026-09-15 | Subject line | Free trial ends soon | Account closes in 48h | 3,200 | 28.4% | 33.1% | Variant | 97% | Yes | Specific time window outperformed generic urgency |
The "Hypothesis Confirmed?" column is the most important field in the log. Over time, a pattern of confirmed hypotheses tells you that your mental model of your audience is accurate. A pattern of disconfirmed hypotheses tells you that your assumptions need revisiting.
If you are managing email alongside a wider content strategy, the same systematic mindset applies. Lenka Studio's content and email clients often pair this log with a broader editorial calendar to track which messages perform best by audience segment and channel.
Step 7: Apply the Decision Framework Before Shipping
When should you declare a winner?
Declare a winner when all three of these conditions are true:
- Statistical significance is at or above 95% (most platforms show this directly)
- The evaluation window has fully closed
- The absolute difference is large enough to matter operationally (a 0.3% open rate lift on a 500-person list is not worth changing your default template)
When should you retest instead of shipping?
Retest when significance is below 90%, when the test ran during an unusual period (a public holiday in Australia, a major sale event), or when the sample was smaller than your calculated minimum. Mark the test in your log as "inconclusive" with a note, then schedule a retest with a larger window.
Pro tip: Apply winners to your templates and ongoing flows, not just the one-off campaign you tested. The compounding effect of improving your base templates is where the real long-term gains come from.
Step 8: Build a Testing Cadence
Running one test every two months produces six data points per year. Running one test per send (where your send volume allows) can produce 50 or more. The right cadence depends on your list size and send frequency.
A realistic starting cadence for a business sending two campaigns per week:
- Weeks 1 and 2: Test subject line framing (urgency vs. benefit)
- Weeks 3 and 4: Test preview text (question vs. statement)
- Weeks 5 and 6: Apply winners, send without testing to establish a new baseline
- Week 7 onwards: Begin testing the next variable
If you are also running social media campaigns and want to coordinate your messaging across channels, the free Lenka Studio social media toolkit includes a content calendar template that pairs well with this email testing cadence.
Frequently Asked Questions
How long should I run an email A/B test?
For subject line tests, a minimum of 4 hours and a maximum of 24 hours works for most lists. Send time tests need at least two to three weeks of data across multiple sends to account for day-of-week variation.
Does a 95% confidence level mean there's a 5% chance the result is wrong?
Yes. It means that if you ran the same test 100 times under identical conditions, roughly 5 of those tests would show a difference this large by chance alone. For most marketing decisions, 95% is a reasonable threshold. For high-stakes decisions affecting your entire automated flow, aim for 99%.
Can I test more than two variants at once?
Yes, most platforms support multivariate tests. However, each additional variant requires a proportionally larger sample size to reach significance. For lists under 10,000 contacts, stick to two variants at a time.
What if my test results conflict with what I expected?
That is useful data. Record the disconfirmed hypothesis in your log with a note on why you think the result went the opposite direction. Revisit it after three to five more tests on the same variable. Patterns across multiple disconfirmed hypotheses often reveal something real about how your audience reads email.
Does this workflow apply to transactional emails too?
Yes, but with care. Transactional emails (receipts, password resets, shipping notifications) have much higher open rates and different goals than promotional campaigns. Test them separately and track click rate or downstream action rather than open rate as your primary metric.
Next Steps
Start with your highest-volume campaign send from the last 90 days. Write one hypothesis using the template in Step 1. Calculate your sample size. Run the test with manual winner review turned on.
The first test will take longer than expected. The fifth test will feel routine. By the tenth, you will have a results log that shapes every email decision your team makes.
If you want help building email workflows that feed into a broader conversion strategy, the team at Lenka Studio works with SMBs across Australia, Singapore, Canada, and the US to design and automate marketing systems that compound over time. Get in touch to talk through what a structured testing programme could look like for your business.




