How to Test Webinar Reminder Calls: The Holdout Design
A holdout test on a webinar list means splitting your registrants into two groups at random, running your reminder calls or texts on one group only, leaving the other group entirely alone, and comparing the two show up rates. It is the only design that tells you whether the calls caused anything, because everything else about the two groups - the ads, the offer, the week, the weather - is identical by construction. This page gives you the design, and then does the thing nobody else publishes: it works through the sample-size arithmetic so you can see, before you spend anything, whether your own webinar is big enough for the answer to mean anything.
The finding this page exists to publish
At a 30% baseline show up rate, detecting a realistic reminder effect - the kind the transferable research actually supports, somewhere around 30% to 34-38% - needs roughly 1,100 to 4,300 registrants split evenly across the two arms, at 80% power and a 5% significance level. A single webinar with 400-600 registrants split 50/50 is almost certainly underpowered. With 600 registrants you can only reliably detect a jump from 30% to about 41%, which is a bigger effect than the published evidence supports for any reminder. The arithmetic is below and you can check every step of it. This page tells you how to test our own product, including why a single test of it probably cannot settle the question either way.
Terms used on this page
- Holdout
- A randomly selected group of registrants who are deliberately excluded from the treatment, so their outcome can be compared against the group that received it. Source
- Arm
- One of the two groups in the test: the treated arm receives the calls, the holdout arm does not. Source
- Statistical power
- The probability that a test detects an effect of a given size when that effect is genuinely there. 80% is the conventional target, meaning a one-in-five chance of missing a real effect. Source
- Minimum detectable effect
- The smallest difference between the two arms that a test of a given size can reliably distinguish from noise. Smaller samples can only detect larger effects. Source
- Underpowered
- A test too small to detect the effect it is looking for. It will usually return "no significant difference" whether or not the effect is real, so a null result carries no information. Source
What is a holdout test on a webinar list?
Take the registrant list for one webinar. Before any reminder goes out, assign each registrant at random to one of two groups. Group A gets the full treatment - the calls, the texts, whatever you are testing. Group B gets nothing beyond whatever your baseline already was. When the session runs, you count how many people from each group attended, and the difference between the two rates is your estimate of what the treatment did.
The reason this design is worth the discomfort is that it removes every alternative explanation at once. If you instead compare this month's webinar (with calls) against last month's (without), the difference could be the calls, or it could be that you changed the ad creative, or that it is now July, or that the topic was better. A random split within one event holds all of that constant by construction, which is exactly why it is the test a serious buyer asks for.
The buyer's version of this, in their own words
An operator we spoke to during the research for this cluster put the requirement more plainly than we could: "Split my registrant list down the middle. Same ads, same week, same offer. That's real, everything else is a story." That is the correct instinct, and it is the standard we think any vendor in this category should be held to - including us.
One clarification before the arithmetic, because it changes what "the treatment" means. A holdout measures the whole package you applied to the treated arm, not any single component inside it. If the treated arm gets a text, a call, and a different email sequence, you have measured all three together. That is fine if what you want to know is "does buying this thing help me". It is useless if what you want to know is "which of the three did it", which needs a much larger test with more arms and is out of reach for almost every operator running this.
How many registrants do you need to detect a real lift?
This is the section that decides whether the rest of the page is worth acting on, and it is arithmetic rather than a claim. We are going to work it through in full so you can redo it with your own numbers and disagree with us if the disagreement is real.
The formula, and where it comes from
For a superiority test comparing two independent proportions with equal group sizes, the standard sample size per arm is:
n per arm = f(α, β) × [ p₁(1 − p₁) + p₂(1 − p₂) ] / (p₂ − p₁)²
where p₁ is the control rate, p₂ is the rate you want to be able to detect, and f(α, β) is the squared sum of two standard normal deviates. Sealed Envelope publishes this formula verbatim for binary superiority trials. For the deviates we use the conventional values given by Suresh and Chandrashekara (2012) in the Journal of Human Reproductive Sciences: 1.96 for a 5% two-sided significance level and 0.84 for 80% power. So f = (1.96 + 0.84)² = 7.84.
Worked example, step by step
Take the base case: a 30% baseline lifted to 38%. Substitute and the whole thing is four lines of school arithmetic.
- p₁(1 − p₁) = 0.30 × 0.70 = 0.21
- p₂(1 − p₂) = 0.38 × 0.62 = 0.2356
- Sum = 0.4456. Difference squared = (0.38 − 0.30)² = 0.0064
- n per arm = 7.84 × 0.4456 / 0.0064 = 7.84 × 69.625 = 545.9, round up to 546
So 1,092 registrants in total, 546 in each arm, to have a four-in-five chance of detecting an eight-point lift if that lift is genuinely there. Here is the same calculation across the range of effects worth caring about.
| Effect you want to detect | Registrants per arm | Total registrants (50/50) | Webinars at 500 registrants each | Is this effect size plausible? |
|---|---|---|---|---|
| 30% to 34% (+4pp) | 2,129 | 4,258 | about 9 | Yes - this is roughly what the Cochrane SMS reminder review implies when its RR 1.14 is applied to a 30% baseline |
| 30% to 36% (+6pp) | 960 | 1,920 | about 4 | Optimistic but inside the Cochrane confidence interval |
| 30% to 38% (+8pp) | 546 | 1,092 | about 2 or 3 | A relative lift of 1.27x - at the upper edge of the published reminder evidence |
| 30% to 40% (+10pp) | 353 | 706 | about 1.5 | Beyond anything the transferable literature supports for a reminder |
| 30% to 45% (+15pp) | 160 | 320 | under 1 | Not supported by any published reminder study we could verify |
Read that table backwards
The tests that fit inside one normal webinar are exactly the tests looking for effects nobody has ever measured. The effect sizes that the evidence actually supports are the ones that need thousands of registrants. That inversion is the whole problem with testing reminder calls, it is not specific to us, and any vendor who tells you a single event will settle it has either not done this arithmetic or is hoping you have not.
Where the "realistic effect" number comes from
We did not pick 34% to make our own product look hard to prove. It is what happens when you take the strongest transferable reminder evidence and apply it to a 30% baseline.
Apply RR 1.14 to a 30% baseline and you get 34.2%. That is a four-point lift, and four points is what the 4,258-registrant row in the table above is priced for. The Cochrane review is Gurol-Urganci and colleagues, 2013, and it covers healthcare appointment attendance in Australia, China, Kenya, Malaysia and the UK - not webinars. We use it because it is the best-designed reminder evidence that exists, and we label the transfer as an assumption rather than a finding.
The other anchor is Nickerson and Rogers (2010), whose abstract states that facilitating a voting plan "can increase turnout by 4.1 percentage points among those contacted, but a standard encouragement call and self-prediction have no significant impact". Two things travel with that number every time we use it. The 4.1 points is the effect among people the callers actually reached, not among everyone assigned to be called. And the setting is the 2008 US presidential election, with live human callers - whether the same effect survives when the caller is a disclosed AI is, as far as we can establish, unpublished. Nobody has tested it.
The result you should be prepared for is zero
The single strongest contradicting finding in this literature is Ødegård and colleagues (2022) in PLOS One: two-way text messages did not improve appointment attendance compared with standard care, RR 1.03, 95% CI 0.95 to 1.12, I² = 53%. Two caveats are mandatory. The review states that the 31 included trials were "all set in Sub-Saharan Africa", which is a serious external-validity limit. And the source contradicts itself on the sample: its abstract gives "5 trials, 4374 participants" for the attendance analysis while its body text says "Five trials (6627 participants) were included in our primary analysis on appointment attendance". We quote both rather than quietly picking one. A well-run holdout that returns nothing is not a failed test - it is a result that already has precedent.
What can a single webinar actually detect?
Run the formula in the other direction. Fix the number of registrants you actually have, and solve for the smallest effect the test could reliably find. This is the table that should decide whether you run the test at all.
| Registrants on the list | Per arm at 50/50 | Smallest lift detectable at 80% power | Verdict |
|---|---|---|---|
| 400 | 200 | 30% to about 43% (+13pp) | Underpowered for any realistic effect. A null result here means nothing |
| 500 | 250 | 30% to about 42% (+12pp) | Underpowered. This is the size most operators imagine is enough |
| 600 | 300 | 30% to about 41% (+11pp) | Still underpowered for anything the evidence supports |
| 1,000 | 500 | 30% to about 38% (+8pp) | Can detect a large effect only. Directional at best |
| 2,000 | 1,000 | 30% to about 36% (+6pp) | Now inside the optimistic end of the plausible range |
| 4,258 | 2,129 | 30% to about 34% (+4pp) | The realistic effect finally becomes detectable |
So the plain answer to "can I settle this with my next webinar" is: almost certainly not. A 400 to 600 registrant event split down the middle can only see effects roughly two to three times larger than the ones the published reminder research supports. If you run it anyway and get a null, you have learned nothing about the treatment - you have learned that your test was too small, which you already knew from this table.
What to do instead of giving up
Pool across events. Keep the same random-split procedure running across four, six, nine consecutive webinars, holding the design constant, and add the arms together. Nine events of 500 registrants gets you to the 4,258 that the realistic effect needs. Or accept a smaller question. Run the split anyway, report the direction and the confidence interval rather than a yes/no verdict, and treat it as one input rather than a proof. Both of those are honest. What is not honest is running an underpowered test, getting a positive result by luck, and putting it in a deck.
Should you split 50/50 or 90/10?
Most operators want a 90/10 split. The instinct is understandable: treat almost everyone, hold back a token slice, lose almost nothing. The arithmetic says this choice is expensive, and the reason is worth understanding in plain language rather than taking on trust.
The smaller arm limits you. The precision of a comparison is governed by whichever side has fewer people, because a rate measured on 60 people is a wobbly number no matter how precisely you have measured the other 540. Adding more people to the already-large arm buys you very little; the uncertainty is coming from the small side. A 50/50 split is the allocation that makes both arms as precise as they can be for a fixed total.
The published correction for this is simple. Suresh and Chandrashekara (2012) state it directly: "If r = n1/n2 is the ratio of sample size in 2 groups, then the required sample size is N1 = N(1+r)²/4r", where N is the total you calculated assuming equal groups.
| Split | Ratio r | Multiplier (1+r)²/4r | Total registrants for the 30% to 38% test | Size of the holdout arm |
|---|---|---|---|---|
| 50 / 50 | 1 | 1.00 | 1,092 | 546 |
| 70 / 30 | 2.33 | 1.19 | about 1,300 | about 390 |
| 80 / 20 | 4 | 1.56 | about 1,706 | about 341 |
| 90 / 10 | 9 | 2.78 | about 3,033 | about 303 |
Read the last two columns together, because the trade is not the one people expect. Going lopsided shrinks your holdout arm only modestly - from 546 at an even split down to about 303 at 90/10 - while the total list you must accumulate nearly triples. You are buying a slightly smaller holdout at the cost of almost three times the registrants. And there is a floor: no matter how enormous you make the treated arm, the holdout arm for this effect cannot go below about 257, because at that point the uncertainty is coming entirely from the small side and adding to the other side buys nothing at all.
A note on how we computed that
The 2.78 multiplier is the published formula, which assumes the same variance in both arms. If you instead substitute the two different variances directly into the two-proportion formula for a 9:1 allocation, you get a total of about 2,893 rather than 3,033 - a penalty of 2.65x instead of 2.78x. Both are checkable, both point the same way, and neither changes the recommendation. We show you the difference rather than presenting one figure as exact.
What should you measure, and where is the trap?
Show up rate is the obvious metric and it is the wrong one to measure alone. The objection that gets raised in every serious conversation about this is: show rate is a vanity metric, where are the sales? That objection is correct, and it has a specific mechanism behind it rather than just cynicism.
The marginal attendee assumption - ours, not a finding
The people your calls bring into the room are, by definition, the ones who would not have come on their own. It is reasonable to expect them to convert worse than a self-motivated attendee who cleared their calendar unprompted. This is our assumption and we cannot cite it - we found no published study measuring the conversion rate of reminder-induced attendees against organic ones. State it as an assumption in your own analysis too. Its practical consequence is that a lift in show up rate does not translate one-for-one into a lift in revenue, and a test that only counts bodies in the room will systematically overstate what the treatment is worth to you.
Measure the whole chain in both arms, and keep the arm label attached to every downstream record so you can trace a sale back to the group it came from. That last part is the piece people forget, and without it the test is unrecoverable after the fact.
| Metric | What it tells you | Why it is not enough on its own |
|---|---|---|
| Reached rate | What share of the treated arm you actually got hold of | Without it you cannot tell a weak treatment from a treatment that never landed. It also defines what an "among those reached" estimate would even mean |
| Show up rate (live) | The headline number, the one the treatment is aimed at | It is the input to the thing you care about, not the thing itself |
| Stay rate / time in session | Whether the extra attendees stayed or bounced in four minutes | A body that leaves before the offer is worth roughly nothing, and reminder-induced attendees are the ones most likely to do it |
| Booked calls | Whether attendance converted into a next step | The first metric with a plausible link to money. Also the first place the marginal-attendee effect shows up |
| Attended calls | Whether the booking held | Booked-and-missed is a well-known way for a funnel metric to look healthy while producing nothing |
| Sales and revenue | The only number that pays for the system | Needs the largest sample of all - see the note below on why you will rarely power this arm properly |
| Complaints, opt-outs, unsubscribes | The cost side | A treatment that lifts attendance and burns your list is a bad trade you will only notice if you counted |
Why you almost certainly cannot power the sales arm
Run the same formula on a conversion step rather than an attendance step. If 4% of attendees buy and you hope the treatment moves that to 5%, then p₁ = 0.04 and p₂ = 0.05, so the numerator is 0.0384 + 0.0475 = 0.0859 and the denominator is 0.0001. That gives 7.84 × 859 = 6,735 attendees per arm. Not registrants - attendees. For most operators that is several years of webinars. The honest consequence: measure sales, report them, and treat them as directional evidence rather than as a test you have powered. Anyone claiming a statistically established revenue lift from a reminder system should be asked how many attendees were in each arm.
How do you actually run the test?
The design is not complicated. The discipline is where these fall apart, and every step below exists because it is a way we have seen a test become unreadable.
Write the plan down before the list exists
Record the primary metric, the arm sizes, the split ratio, how many events you will pool, and what result would make you stop. Save it with a timestamp somewhere you cannot quietly edit it. Everything after this step is easier to fudge than you think.
Randomise at the person level, not by any list attribute
Assign each registrant to an arm with a random number. Do not split by registration date, alphabet, source, time zone or "the ones who look engaged". Every one of those correlates with attendance, and any of them turns your treatment effect into a measurement of the sorting rule.
Freeze the assignment before the first touch goes out
Assignment happens once, before anything is sent. Nobody moves between arms afterwards, including the person who complained, the person who looks like a big buyer, and the person your setter thinks is worth a call.
Keep everything else identical across both arms
Same registration page, same confirmation email, same standard reminder emails, same landing experience, same session. The only difference between the arms is the thing you are testing. If the treated arm also gets a different email sequence, you have measured the package, not the calls.
Log the arm label on every downstream record
Attendance, stay time, bookings, attended calls, sales, refunds, opt-outs. If the arm label is not attached at the moment the record is created, you will not be able to reconstruct it later, and the test dies quietly at the analysis stage.
Run it across several events without changing the design
One event is a data point. Pool the arms across four to nine consecutive webinars with the same traffic source, offer and slot. Changing the script, the timing or the channel mid-way restarts the count - you now have two smaller tests instead of one adequate one.
Analyse once, at the pre-declared stopping point
Checking after every event and stopping when the numbers look good is the single most common way to manufacture a false positive. If you genuinely need interim looks, decide the rule for them in step one rather than inventing it when the graph looks promising.
Report the interval, not just the point estimate
Publish or record the difference between the arms with its confidence interval and both arm sizes. "Plus six points, interval minus two to plus fourteen" is an honest sentence. "It lifted attendance by six points" from the same data is not.
What contaminates a webinar holdout?
Randomisation protects you from confounding, not from leakage. These are the failure modes specific to running this on a registrant list rather than in a lab, and most of them do not announce themselves.
| Contaminant | What it does to your result | What to do about it |
|---|---|---|
| People on the list talk to each other | A holdout registrant hears about the call from a treated colleague, or gets forwarded the link. The two arms stop being independent and the measured difference shrinks toward zero | Randomise by company or household rather than by person where the list is clustered. Expect any measured effect to be a floor, not a ceiling |
| Traffic source changed between events | The single biggest confounder in this subject. Attendance moves enormously with who supplied the registrant, so a mid-test change in ad mix can swamp the treatment entirely | Hold the source and the mix constant across all pooled events, and record the mix per event so you can check afterwards |
| Seasonality and calendar effects | A December webinar and a September one are different populations behaving differently. Pooling across a season is fine; comparing across one is not | Because the split is within-event, seasonality hits both arms equally. This is the main reason within-event beats event-to-event |
| The offer changed | A new promise, price or bonus changes who registers and who shows. You are now measuring the offer | One variable at a time. If the offer must change, close the test, record it, and start a fresh count |
| You behaved differently toward one arm | The most common contamination and the hardest to see. Extra effort on treated registrants, a warmer email, a personal nudge from a setter - all of it lands on the treatment side | Make the operational process identical apart from the tested touch, and ideally keep the arm assignment out of view of whoever is doing the manual work |
| Attendance attribution is broken | Duplicate joins, shared links, and people joining on a second device can inflate or split attendance records unevenly | Deduplicate on the platform’s own registrant identifier before you analyse, not on email address |
| Small-arm dropout | Refunds, bounces, invalid numbers and unreachable registrants thin the treated arm and change what you are comparing | Analyse everyone as assigned, then report the reached rate separately. Do not silently delete the unreachable from the treated arm |
Which of these actually matters most
If you can only control one thing, control the traffic source. Platform data shows attendance moving by more than twenty percentage points depending on whose list the registrant came from - Banzai and Demio's 2024 report on 800,000-plus webinars puts host companies under $1M in revenue at 22% attendance against 46-48% for companies over $10M, on the same platform in the same year. An effect that large will bury a four-point treatment effect without difficulty.
What if you will not withhold it from half your list?
Some operators refuse to run a holdout because they believe the treatment helps, and deliberately withholding help from half their audience feels wrong. That is a legitimate ethical position and not an objection to argue past. If you genuinely believe the calls help the registrant - not just you - then randomising who gets helped is a real cost, and the fact that it produces a cleaner number does not automatically settle the question.
There is a weaker design available that most people find ethically comfortable: alternate the policy across consecutive events. Run the system on event one, not on event two, on event three, not on event four, and compare the pooled treated events against the pooled untreated ones. Everyone at a given event is treated the same, which is what the objection was actually about.
| Within-event holdout (50/50) | Alternating across events | |
|---|---|---|
| Controls for seasonality, ads, offer, week | Yes, by construction | No. Time is confounded with the treatment |
| Registrants needed for a realistic effect | About 4,258 across pooled events | More, because event-to-event variance adds noise on top |
| Two people at the same event treated differently | Yes - this is the objection | No |
| Vulnerable to a mid-test change in traffic source | Only if it happens between pooled events | Severely - a source change can land entirely on one arm |
| How many events before it says anything | Several | More than several, and honestly more than most operators will run |
| Verdict | Cleaner, faster, ethically harder | Noisier, slower, ethically comfortable. Still far better than no test |
If you alternate, alternate rather than switching once. Treated, untreated, treated, untreated, so that a drift in your ad costs or your list fatigue lands roughly equally on both arms instead of entirely on the second half of the experiment. And write down the traffic mix per event, because if it moves, you will want to know whether it moved before or after you saw the result.
How long before you know anything?
Several events, not one. That is the whole answer, and the tables above are why. If your webinars run 500 registrants and you are looking for the realistic four-point effect, you need roughly nine events before the arithmetic gives you the power you assumed you had from the beginning. At a monthly cadence that is most of a year.
There is a shorter honest version. If you are prepared to settle for "is this pointing the right way and is the interval consistent with the published evidence" rather than "is this proven", three or four events will tell you something worth knowing, and you will have a confidence interval you can report without embarrassment. What three or four events will not give you is a number to put in a case study, and if a vendor produces one from a sample that size, the arithmetic on this page is how you check it.
Write the rules down before you start
Pre-registration is a formal habit borrowed from clinical research and it takes fifteen minutes. Write down, before the first registrant is assigned: the primary metric, the secondary metrics, the arm sizes, the number of events you will pool, the analysis you will run, and the stopping rule. Then do not change any of it while the test is running.
Why this is the step that decides whether the test was worth running
Without a pre-declared metric and stopping rule, you have several plausible outcome measures, several plausible cut-off points, and a strong preference about the answer. That combination reliably produces a significant-looking result whether or not an effect exists, and it will produce it for the person who genuinely believes they are being fair. The discipline is not about honesty; it is about the fact that honesty is not sufficient here. If you take one thing from this page other than the arithmetic, take this.
When should you not bother testing at all?
Testing is not free. It costs list size, operational discipline, and the option of just getting on with it. There are cases where the correct decision is to skip the test and decide on judgement instead, and it is worth naming them rather than implying everyone should run an experiment.
| Situation | Test? | Why |
|---|---|---|
| Your webinars have fewer than about 200 registrants | No | You cannot detect anything smaller than a 20-point swing, and no reminder produces one. Decide on judgement, or pool over so many events that the world will have changed underneath you |
| Traffic source is unstable month to month | No, not yet | The confounder is larger than the effect. Stabilise the source first, then test. A test run on shifting traffic produces a number that means nothing and gets quoted anyway |
| You are changing three things at once | No | You will learn that the package did something, and nothing about which part. If that is genuinely what you want to know, fine - but do not later claim you tested the calls |
| The offer or the funnel is being rebuilt | No | Wait. Anything measured across a rebuild measures the rebuild |
| You already know you will buy it regardless | No | Then the test is theatre. Spend the effort on the script and the disclosure instead, and be honest that this was a decision rather than a finding |
| The cost is small and the downside is a slightly worse show rate | Probably not | A test that costs more to run properly than the decision is worth is a bad use of a year. Just try it and watch your own numbers over time |
| You are being asked to sign a large or long contract | Yes | This is exactly the case the holdout exists for. Pool across events, pre-register, and make the pilot the thing you are buying first |
| A vendor is quoting you a specific lift | Yes | Ask which sample it came from and run the arithmetic on this page against it. Most published lifts in this category do not survive that question |
What this means if you are thinking of buying this from us
We sell a system that calls and texts webinar registrants, and this page hands you the tool to check whether it works, including the arithmetic showing that most single tests cannot prove it either way. That is deliberate. We would rather you arrive at a conversation knowing what a real answer would cost than sign on the strength of a number we made comfortable.
What we can and cannot show you
We have zero closed clients on this offer. There is no case study, no logo wall, no result to point at, and we are not going to manufacture one. That is precisely why the holdout matters here: with no client evidence to lean on, the only honest thing we can offer is a design for measuring us. We also cannot tell you whether a disclosed AI caller produces the same effect as a human one - the plan-making research used live human callers, and we could find no published study on the disclosed-AI version. That is unknown, not established, and we are not going to write it either way.
One consequence follows for anyone shopping this category. If a vendor tries to talk you out of a holdout, you have learned what you needed to know. The design is not exotic, it costs the vendor nothing but the risk of a null result, and the only reason to resist it is not wanting the number. The same test applies to us - and if we ever ask you to skip it, hold us to this paragraph.
For what the numbers underneath a test like this actually look like, our teardown of every published webinar attendance study gives the sources and the samples. The two questions most operators should settle before spending money on reminders are covered in does charging for a webinar increase attendance - the largest reported lever in the whole corpus, and one nobody has controlled either - and should you offer the webinar replay, which is the most-stated reason people skip. Both of those pages end by telling you to test it on your own list; this is the page that tells you how. The operational side of what we build sits on AI webinar attendance, the evidence review on webinar reminder calls, and the after-the-event work on post-webinar follow-up calls. If you are in the EU or the UK, check what your registration form has to say before anyone may call those registrants on webinar registrant consent, and if you would rather work through the design with us, our contact page books a scoped conversation.
Frequently Asked Questions
Frequently Asked Questions
You split your registrant list at random into two groups before any reminder goes out. One group gets the calls and texts, the other gets nothing beyond your existing baseline. When the session runs you compare the two show up rates. Because the split is random and within a single event, the ads, the offer, the week and the topic are identical across both groups, so a difference in attendance is attributable to the treatment rather than to anything else that changed.
At a 30% baseline, 80% power and a 5% two-sided significance level, you need 546 registrants per arm (1,092 total) to detect a lift to 38%, and 2,129 per arm (4,258 total) to detect a lift to 34%. The formula is n per arm = 7.84 x [p1(1-p1) + p2(1-p2)] / (p2-p1) squared, where 7.84 is (1.96 + 0.84) squared. A lift to 34% is roughly what the Cochrane reminder review implies when its risk ratio of 1.14 is applied to a 30% baseline, so the realistic case is the expensive one.
Almost certainly not. A 400-registrant webinar split 50/50 can only reliably detect a jump from 30% to about 43%. A 600-registrant one gets you to about 41%. Both of those are far larger than any effect the published reminder literature supports, so a null result from a single event tells you your test was too small rather than telling you anything about the treatment. Pool the same design across several events, or report direction and a confidence interval rather than a verdict.
Substantially, because the smaller arm limits the precision of the whole comparison. Suresh and Chandrashekara (2012) give the penalty as N(1+r) squared divided by 4r, so at a 9:1 allocation you need 2.78 times as many registrants in total to answer the same question - about 3,033 instead of 1,092 for a 30% to 38% test. Computing it exactly with the two arms' own variances gives about 2,893, a penalty of 2.65 times. What the lopsided split buys you is a holdout arm that shrinks only modestly, from 546 down to about 303, with a hard floor around 257 no matter how large the treated arm gets. You pay nearly three times the total list for that. A 50/50 split is the efficient choice.
Reached rate, stay rate, booked calls, attended calls, sales and revenue, and complaints or opt-outs - with the arm label attached to every downstream record at the moment it is created. Show up rate alone will overstate the value of the treatment, because the attendees your calls bring in are by definition the ones who would not have come unprompted and may convert worse than self-motivated ones. That last point is our assumption, not a published finding; we could not locate any study measuring it.
Very probably not, and the arithmetic says why. If 4% of attendees buy and you hope for 5%, the same formula gives 7.84 x [0.04(0.96) + 0.05(0.95)] / 0.0001 = about 6,735 attendees per arm. Not registrants - attendees. For most operators that is several years of webinars. Measure sales, report them, and treat them as directional. Anyone claiming a statistically established revenue lift from a reminder system should be asked how many attendees were in each arm.
That is a legitimate position, not an obstacle. The alternative is to alternate the policy across consecutive events - system on, system off, system on, system off - and pool the treated events against the untreated ones. Everyone at a given event is treated the same. It is weaker, because event-to-event variance in traffic source, season and list fatigue is real and now sits inside your comparison, so it needs more events to say the same thing. It is still far better than no test.
Registrants on the same list talking to each other or forwarding the link, so the arms stop being independent. The traffic source changing between pooled events, which is the largest confounder in this subject. Seasonality, if you compare across events instead of splitting within one. Changing the offer mid-test. And the most common one: behaving differently toward the treated arm - extra effort, a warmer email, a setter making an unscheduled call - which quietly adds an untested treatment to one side.
If your webinars have fewer than about 200 registrants, if your traffic source is unstable month to month, if you are changing three things at once, or if you have already decided to buy regardless. In those cases the test either cannot detect anything meaningful or measures something other than what you think. Run one when the contract is large or long, or when a vendor quotes you a specific lift - then check their claimed sample against the arithmetic on this page.
No. We have zero closed clients on this offer, so there is no case study, no logo wall and no result to point at, and we would rather say so than manufacture one. That is exactly why we publish the holdout design and the sample-size arithmetic, including the part showing that a single test of our own product probably cannot settle it either way. If a vendor in this category tries to talk you out of a holdout, you have learned what you needed to know.
Unknown. The strongest evidence for asking someone to form a concrete plan comes from Nickerson and Rogers (2010), a field experiment with N=287,228 that used live human callers and reported a 4.1 percentage point turnout lift among those contacted, with a standard encouragement call having no significant impact. We could find no published study testing whether the same effect survives when the caller identifies itself as an AI assistant. It is untested rather than established, and we are not going to assert it in either direction.
Founder & CEO, AInora
Building AI digital administrators that replace front-desk overhead for service businesses across Europe. Previously built voice AI systems for dental clinics, hotels, and restaurants.
View all articlesReady to try AI for your business?
Hear how AInora sounds handling a real business call. Try the live voice demo or book a consultation.
Related Articles
Does Charging for a Webinar Increase Attendance? The Honest Answer
The largest reported lever in the corpus, from one organiser over six years - and why nobody has ever controlled it.
Should You Offer the Webinar Replay? Nobody Has Tested It
The most-stated reason people skip a live session, and the least-studied thing in the subject. The trade-off in both directions.
Average Webinar Attendance Rate: Every Study (2026)
Every benchmark with its real sample, source market and figure - plus the widely-quoted numbers that do not survive a click.
Why Webinar Registrants Do Not Show Up (Causes, Ranked)
The ranked reasons people register and then skip, in operators’ own words - starting with the replay.