AInora
Webinar AttendanceHoldout TestShow Up RateExperiment DesignStatistical PowerEvidence Review

How to Test Webinar Reminder Calls: The Holdout Design

JB
Justas ButkusFounder, Ainora
··15 min read

A holdout test on a webinar list means splitting your registrants into two groups at random, running your reminder calls or texts on one group only, leaving the other group entirely alone, and comparing the two show up rates. It is the only design that tells you whether the calls caused anything, because everything else about the two groups - the ads, the offer, the week, the weather - is identical by construction. This page gives you the design, and then does the thing nobody else publishes: it works through the sample-size arithmetic so you can see, before you spend anything, whether your own webinar is big enough for the answer to mean anything.

The finding this page exists to publish

At a 30% baseline show up rate, detecting a realistic reminder effect - the kind the transferable research actually supports, somewhere around 30% to 34-38% - needs roughly 1,100 to 4,300 registrants split evenly across the two arms, at 80% power and a 5% significance level. A single webinar with 400-600 registrants split 50/50 is almost certainly underpowered. With 600 registrants you can only reliably detect a jump from 30% to about 41%, which is a bigger effect than the published evidence supports for any reminder. The arithmetic is below and you can check every step of it. This page tells you how to test our own product, including why a single test of it probably cannot settle the question either way.

1,092
Registrants needed in total (546 per arm) to detect 30% to 38% at 80% power, alpha 0.05
Source: Two-proportion sample-size formula, Sealed Envelope
4,258
Registrants needed in total (2,129 per arm) to detect 30% to 34% - the effect size implied by the Cochrane reminder review
Source: Cochrane 2013, RR 1.14 applied to a 30% baseline
2.78x
More total registrants needed at a 90/10 split than at 50/50, from the published unequal-allocation multiplier
Source: Suresh & Chandrashekara 2012, J Hum Reprod Sci

Terms used on this page

Holdout
A randomly selected group of registrants who are deliberately excluded from the treatment, so their outcome can be compared against the group that received it. Source
Arm
One of the two groups in the test: the treated arm receives the calls, the holdout arm does not. Source
Statistical power
The probability that a test detects an effect of a given size when that effect is genuinely there. 80% is the conventional target, meaning a one-in-five chance of missing a real effect. Source
Minimum detectable effect
The smallest difference between the two arms that a test of a given size can reliably distinguish from noise. Smaller samples can only detect larger effects. Source
Underpowered
A test too small to detect the effect it is looking for. It will usually return "no significant difference" whether or not the effect is real, so a null result carries no information. Source

What is a holdout test on a webinar list?

Take the registrant list for one webinar. Before any reminder goes out, assign each registrant at random to one of two groups. Group A gets the full treatment - the calls, the texts, whatever you are testing. Group B gets nothing beyond whatever your baseline already was. When the session runs, you count how many people from each group attended, and the difference between the two rates is your estimate of what the treatment did.

The reason this design is worth the discomfort is that it removes every alternative explanation at once. If you instead compare this month's webinar (with calls) against last month's (without), the difference could be the calls, or it could be that you changed the ad creative, or that it is now July, or that the topic was better. A random split within one event holds all of that constant by construction, which is exactly why it is the test a serious buyer asks for.

The buyer's version of this, in their own words

An operator we spoke to during the research for this cluster put the requirement more plainly than we could: "Split my registrant list down the middle. Same ads, same week, same offer. That's real, everything else is a story." That is the correct instinct, and it is the standard we think any vendor in this category should be held to - including us.

One clarification before the arithmetic, because it changes what "the treatment" means. A holdout measures the whole package you applied to the treated arm, not any single component inside it. If the treated arm gets a text, a call, and a different email sequence, you have measured all three together. That is fine if what you want to know is "does buying this thing help me". It is useless if what you want to know is "which of the three did it", which needs a much larger test with more arms and is out of reach for almost every operator running this.

How many registrants do you need to detect a real lift?

This is the section that decides whether the rest of the page is worth acting on, and it is arithmetic rather than a claim. We are going to work it through in full so you can redo it with your own numbers and disagree with us if the disagreement is real.

The formula, and where it comes from

For a superiority test comparing two independent proportions with equal group sizes, the standard sample size per arm is:

n per arm = f(α, β) × [ p₁(1 − p₁) + p₂(1 − p₂) ] / (p₂ − p₁)²

where p₁ is the control rate, p₂ is the rate you want to be able to detect, and f(α, β) is the squared sum of two standard normal deviates. Sealed Envelope publishes this formula verbatim for binary superiority trials. For the deviates we use the conventional values given by Suresh and Chandrashekara (2012) in the Journal of Human Reproductive Sciences: 1.96 for a 5% two-sided significance level and 0.84 for 80% power. So f = (1.96 + 0.84)² = 7.84.

Worked example, step by step

Take the base case: a 30% baseline lifted to 38%. Substitute and the whole thing is four lines of school arithmetic.

  • p₁(1 − p₁) = 0.30 × 0.70 = 0.21
  • p₂(1 − p₂) = 0.38 × 0.62 = 0.2356
  • Sum = 0.4456. Difference squared = (0.38 − 0.30)² = 0.0064
  • n per arm = 7.84 × 0.4456 / 0.0064 = 7.84 × 69.625 = 545.9, round up to 546

So 1,092 registrants in total, 546 in each arm, to have a four-in-five chance of detecting an eight-point lift if that lift is genuinely there. Here is the same calculation across the range of effects worth caring about.

Effect you want to detectRegistrants per armTotal registrants (50/50)Webinars at 500 registrants eachIs this effect size plausible?
30% to 34% (+4pp)2,1294,258about 9Yes - this is roughly what the Cochrane SMS reminder review implies when its RR 1.14 is applied to a 30% baseline
30% to 36% (+6pp)9601,920about 4Optimistic but inside the Cochrane confidence interval
30% to 38% (+8pp)5461,092about 2 or 3A relative lift of 1.27x - at the upper edge of the published reminder evidence
30% to 40% (+10pp)353706about 1.5Beyond anything the transferable literature supports for a reminder
30% to 45% (+15pp)160320under 1Not supported by any published reminder study we could verify

Read that table backwards

The tests that fit inside one normal webinar are exactly the tests looking for effects nobody has ever measured. The effect sizes that the evidence actually supports are the ones that need thousands of registrants. That inversion is the whole problem with testing reminder calls, it is not specific to us, and any vendor who tells you a single event will settle it has either not done this arithmetic or is hoping you have not.

Where the "realistic effect" number comes from

We did not pick 34% to make our own product look hard to prove. It is what happens when you take the strongest transferable reminder evidence and apply it to a 30% baseline.

RR 1.14
SMS reminders vs no reminders for healthcare appointments - 7 studies, 5,841 participants, moderate quality evidence (95% CI 1.03 to 1.26)
Source: Cochrane 2013, Gurol-Urganci et al.
RR 0.99
SMS reminders vs phone call reminders - 3 studies, 2,509 participants. No detectable difference (95% CI 0.95 to 1.02)
Source: Cochrane 2013, Gurol-Urganci et al.
+4.1pp
Turnout lift among those contacted from a plan-making script, N=287,228 - a standard encouragement call had no significant impact
Source: Nickerson & Rogers 2010, Psychological Science 21(2) 194-199

Apply RR 1.14 to a 30% baseline and you get 34.2%. That is a four-point lift, and four points is what the 4,258-registrant row in the table above is priced for. The Cochrane review is Gurol-Urganci and colleagues, 2013, and it covers healthcare appointment attendance in Australia, China, Kenya, Malaysia and the UK - not webinars. We use it because it is the best-designed reminder evidence that exists, and we label the transfer as an assumption rather than a finding.

The other anchor is Nickerson and Rogers (2010), whose abstract states that facilitating a voting plan "can increase turnout by 4.1 percentage points among those contacted, but a standard encouragement call and self-prediction have no significant impact". Two things travel with that number every time we use it. The 4.1 points is the effect among people the callers actually reached, not among everyone assigned to be called. And the setting is the 2008 US presidential election, with live human callers - whether the same effect survives when the caller is a disclosed AI is, as far as we can establish, unpublished. Nobody has tested it.

The result you should be prepared for is zero

The single strongest contradicting finding in this literature is Ødegård and colleagues (2022) in PLOS One: two-way text messages did not improve appointment attendance compared with standard care, RR 1.03, 95% CI 0.95 to 1.12, I² = 53%. Two caveats are mandatory. The review states that the 31 included trials were "all set in Sub-Saharan Africa", which is a serious external-validity limit. And the source contradicts itself on the sample: its abstract gives "5 trials, 4374 participants" for the attendance analysis while its body text says "Five trials (6627 participants) were included in our primary analysis on appointment attendance". We quote both rather than quietly picking one. A well-run holdout that returns nothing is not a failed test - it is a result that already has precedent.

What can a single webinar actually detect?

Run the formula in the other direction. Fix the number of registrants you actually have, and solve for the smallest effect the test could reliably find. This is the table that should decide whether you run the test at all.

Registrants on the listPer arm at 50/50Smallest lift detectable at 80% powerVerdict
40020030% to about 43% (+13pp)Underpowered for any realistic effect. A null result here means nothing
50025030% to about 42% (+12pp)Underpowered. This is the size most operators imagine is enough
60030030% to about 41% (+11pp)Still underpowered for anything the evidence supports
1,00050030% to about 38% (+8pp)Can detect a large effect only. Directional at best
2,0001,00030% to about 36% (+6pp)Now inside the optimistic end of the plausible range
4,2582,12930% to about 34% (+4pp)The realistic effect finally becomes detectable

So the plain answer to "can I settle this with my next webinar" is: almost certainly not. A 400 to 600 registrant event split down the middle can only see effects roughly two to three times larger than the ones the published reminder research supports. If you run it anyway and get a null, you have learned nothing about the treatment - you have learned that your test was too small, which you already knew from this table.

What to do instead of giving up

Pool across events. Keep the same random-split procedure running across four, six, nine consecutive webinars, holding the design constant, and add the arms together. Nine events of 500 registrants gets you to the 4,258 that the realistic effect needs. Or accept a smaller question. Run the split anyway, report the direction and the confidence interval rather than a yes/no verdict, and treat it as one input rather than a proof. Both of those are honest. What is not honest is running an underpowered test, getting a positive result by luck, and putting it in a deck.

Should you split 50/50 or 90/10?

Most operators want a 90/10 split. The instinct is understandable: treat almost everyone, hold back a token slice, lose almost nothing. The arithmetic says this choice is expensive, and the reason is worth understanding in plain language rather than taking on trust.

The smaller arm limits you. The precision of a comparison is governed by whichever side has fewer people, because a rate measured on 60 people is a wobbly number no matter how precisely you have measured the other 540. Adding more people to the already-large arm buys you very little; the uncertainty is coming from the small side. A 50/50 split is the allocation that makes both arms as precise as they can be for a fixed total.

The published correction for this is simple. Suresh and Chandrashekara (2012) state it directly: "If r = n1/n2 is the ratio of sample size in 2 groups, then the required sample size is N1 = N(1+r)²/4r", where N is the total you calculated assuming equal groups.

SplitRatio rMultiplier (1+r)²/4rTotal registrants for the 30% to 38% testSize of the holdout arm
50 / 5011.001,092546
70 / 302.331.19about 1,300about 390
80 / 2041.56about 1,706about 341
90 / 1092.78about 3,033about 303

Read the last two columns together, because the trade is not the one people expect. Going lopsided shrinks your holdout arm only modestly - from 546 at an even split down to about 303 at 90/10 - while the total list you must accumulate nearly triples. You are buying a slightly smaller holdout at the cost of almost three times the registrants. And there is a floor: no matter how enormous you make the treated arm, the holdout arm for this effect cannot go below about 257, because at that point the uncertainty is coming entirely from the small side and adding to the other side buys nothing at all.

A note on how we computed that

The 2.78 multiplier is the published formula, which assumes the same variance in both arms. If you instead substitute the two different variances directly into the two-proportion formula for a 9:1 allocation, you get a total of about 2,893 rather than 3,033 - a penalty of 2.65x instead of 2.78x. Both are checkable, both point the same way, and neither changes the recommendation. We show you the difference rather than presenting one figure as exact.

What should you measure, and where is the trap?

Show up rate is the obvious metric and it is the wrong one to measure alone. The objection that gets raised in every serious conversation about this is: show rate is a vanity metric, where are the sales? That objection is correct, and it has a specific mechanism behind it rather than just cynicism.

The marginal attendee assumption - ours, not a finding

The people your calls bring into the room are, by definition, the ones who would not have come on their own. It is reasonable to expect them to convert worse than a self-motivated attendee who cleared their calendar unprompted. This is our assumption and we cannot cite it - we found no published study measuring the conversion rate of reminder-induced attendees against organic ones. State it as an assumption in your own analysis too. Its practical consequence is that a lift in show up rate does not translate one-for-one into a lift in revenue, and a test that only counts bodies in the room will systematically overstate what the treatment is worth to you.

Measure the whole chain in both arms, and keep the arm label attached to every downstream record so you can trace a sale back to the group it came from. That last part is the piece people forget, and without it the test is unrecoverable after the fact.

MetricWhat it tells youWhy it is not enough on its own
Reached rateWhat share of the treated arm you actually got hold ofWithout it you cannot tell a weak treatment from a treatment that never landed. It also defines what an "among those reached" estimate would even mean
Show up rate (live)The headline number, the one the treatment is aimed atIt is the input to the thing you care about, not the thing itself
Stay rate / time in sessionWhether the extra attendees stayed or bounced in four minutesA body that leaves before the offer is worth roughly nothing, and reminder-induced attendees are the ones most likely to do it
Booked callsWhether attendance converted into a next stepThe first metric with a plausible link to money. Also the first place the marginal-attendee effect shows up
Attended callsWhether the booking heldBooked-and-missed is a well-known way for a funnel metric to look healthy while producing nothing
Sales and revenueThe only number that pays for the systemNeeds the largest sample of all - see the note below on why you will rarely power this arm properly
Complaints, opt-outs, unsubscribesThe cost sideA treatment that lifts attendance and burns your list is a bad trade you will only notice if you counted

Why you almost certainly cannot power the sales arm

Run the same formula on a conversion step rather than an attendance step. If 4% of attendees buy and you hope the treatment moves that to 5%, then p₁ = 0.04 and p₂ = 0.05, so the numerator is 0.0384 + 0.0475 = 0.0859 and the denominator is 0.0001. That gives 7.84 × 859 = 6,735 attendees per arm. Not registrants - attendees. For most operators that is several years of webinars. The honest consequence: measure sales, report them, and treat them as directional evidence rather than as a test you have powered. Anyone claiming a statistically established revenue lift from a reminder system should be asked how many attendees were in each arm.

How do you actually run the test?

The design is not complicated. The discipline is where these fall apart, and every step below exists because it is a way we have seen a test become unreadable.

1

Write the plan down before the list exists

Record the primary metric, the arm sizes, the split ratio, how many events you will pool, and what result would make you stop. Save it with a timestamp somewhere you cannot quietly edit it. Everything after this step is easier to fudge than you think.

2

Randomise at the person level, not by any list attribute

Assign each registrant to an arm with a random number. Do not split by registration date, alphabet, source, time zone or "the ones who look engaged". Every one of those correlates with attendance, and any of them turns your treatment effect into a measurement of the sorting rule.

3

Freeze the assignment before the first touch goes out

Assignment happens once, before anything is sent. Nobody moves between arms afterwards, including the person who complained, the person who looks like a big buyer, and the person your setter thinks is worth a call.

4

Keep everything else identical across both arms

Same registration page, same confirmation email, same standard reminder emails, same landing experience, same session. The only difference between the arms is the thing you are testing. If the treated arm also gets a different email sequence, you have measured the package, not the calls.

5

Log the arm label on every downstream record

Attendance, stay time, bookings, attended calls, sales, refunds, opt-outs. If the arm label is not attached at the moment the record is created, you will not be able to reconstruct it later, and the test dies quietly at the analysis stage.

6

Run it across several events without changing the design

One event is a data point. Pool the arms across four to nine consecutive webinars with the same traffic source, offer and slot. Changing the script, the timing or the channel mid-way restarts the count - you now have two smaller tests instead of one adequate one.

7

Analyse once, at the pre-declared stopping point

Checking after every event and stopping when the numbers look good is the single most common way to manufacture a false positive. If you genuinely need interim looks, decide the rule for them in step one rather than inventing it when the graph looks promising.

8

Report the interval, not just the point estimate

Publish or record the difference between the arms with its confidence interval and both arm sizes. "Plus six points, interval minus two to plus fourteen" is an honest sentence. "It lifted attendance by six points" from the same data is not.

What contaminates a webinar holdout?

Randomisation protects you from confounding, not from leakage. These are the failure modes specific to running this on a registrant list rather than in a lab, and most of them do not announce themselves.

ContaminantWhat it does to your resultWhat to do about it
People on the list talk to each otherA holdout registrant hears about the call from a treated colleague, or gets forwarded the link. The two arms stop being independent and the measured difference shrinks toward zeroRandomise by company or household rather than by person where the list is clustered. Expect any measured effect to be a floor, not a ceiling
Traffic source changed between eventsThe single biggest confounder in this subject. Attendance moves enormously with who supplied the registrant, so a mid-test change in ad mix can swamp the treatment entirelyHold the source and the mix constant across all pooled events, and record the mix per event so you can check afterwards
Seasonality and calendar effectsA December webinar and a September one are different populations behaving differently. Pooling across a season is fine; comparing across one is notBecause the split is within-event, seasonality hits both arms equally. This is the main reason within-event beats event-to-event
The offer changedA new promise, price or bonus changes who registers and who shows. You are now measuring the offerOne variable at a time. If the offer must change, close the test, record it, and start a fresh count
You behaved differently toward one armThe most common contamination and the hardest to see. Extra effort on treated registrants, a warmer email, a personal nudge from a setter - all of it lands on the treatment sideMake the operational process identical apart from the tested touch, and ideally keep the arm assignment out of view of whoever is doing the manual work
Attendance attribution is brokenDuplicate joins, shared links, and people joining on a second device can inflate or split attendance records unevenlyDeduplicate on the platform’s own registrant identifier before you analyse, not on email address
Small-arm dropoutRefunds, bounces, invalid numbers and unreachable registrants thin the treated arm and change what you are comparingAnalyse everyone as assigned, then report the reached rate separately. Do not silently delete the unreachable from the treated arm

Which of these actually matters most

If you can only control one thing, control the traffic source. Platform data shows attendance moving by more than twenty percentage points depending on whose list the registrant came from - Banzai and Demio's 2024 report on 800,000-plus webinars puts host companies under $1M in revenue at 22% attendance against 46-48% for companies over $10M, on the same platform in the same year. An effect that large will bury a four-point treatment effect without difficulty.

What if you will not withhold it from half your list?

Some operators refuse to run a holdout because they believe the treatment helps, and deliberately withholding help from half their audience feels wrong. That is a legitimate ethical position and not an objection to argue past. If you genuinely believe the calls help the registrant - not just you - then randomising who gets helped is a real cost, and the fact that it produces a cleaner number does not automatically settle the question.

There is a weaker design available that most people find ethically comfortable: alternate the policy across consecutive events. Run the system on event one, not on event two, on event three, not on event four, and compare the pooled treated events against the pooled untreated ones. Everyone at a given event is treated the same, which is what the objection was actually about.

Within-event holdout (50/50)Alternating across events
Controls for seasonality, ads, offer, weekYes, by constructionNo. Time is confounded with the treatment
Registrants needed for a realistic effectAbout 4,258 across pooled eventsMore, because event-to-event variance adds noise on top
Two people at the same event treated differentlyYes - this is the objectionNo
Vulnerable to a mid-test change in traffic sourceOnly if it happens between pooled eventsSeverely - a source change can land entirely on one arm
How many events before it says anythingSeveralMore than several, and honestly more than most operators will run
VerdictCleaner, faster, ethically harderNoisier, slower, ethically comfortable. Still far better than no test

If you alternate, alternate rather than switching once. Treated, untreated, treated, untreated, so that a drift in your ad costs or your list fatigue lands roughly equally on both arms instead of entirely on the second half of the experiment. And write down the traffic mix per event, because if it moves, you will want to know whether it moved before or after you saw the result.

How long before you know anything?

Several events, not one. That is the whole answer, and the tables above are why. If your webinars run 500 registrants and you are looking for the realistic four-point effect, you need roughly nine events before the arithmetic gives you the power you assumed you had from the beginning. At a monthly cadence that is most of a year.

There is a shorter honest version. If you are prepared to settle for "is this pointing the right way and is the interval consistent with the published evidence" rather than "is this proven", three or four events will tell you something worth knowing, and you will have a confidence interval you can report without embarrassment. What three or four events will not give you is a number to put in a case study, and if a vendor produces one from a sample that size, the arithmetic on this page is how you check it.

Write the rules down before you start

Pre-registration is a formal habit borrowed from clinical research and it takes fifteen minutes. Write down, before the first registrant is assigned: the primary metric, the secondary metrics, the arm sizes, the number of events you will pool, the analysis you will run, and the stopping rule. Then do not change any of it while the test is running.

Why this is the step that decides whether the test was worth running

Without a pre-declared metric and stopping rule, you have several plausible outcome measures, several plausible cut-off points, and a strong preference about the answer. That combination reliably produces a significant-looking result whether or not an effect exists, and it will produce it for the person who genuinely believes they are being fair. The discipline is not about honesty; it is about the fact that honesty is not sufficient here. If you take one thing from this page other than the arithmetic, take this.

When should you not bother testing at all?

Testing is not free. It costs list size, operational discipline, and the option of just getting on with it. There are cases where the correct decision is to skip the test and decide on judgement instead, and it is worth naming them rather than implying everyone should run an experiment.

SituationTest?Why
Your webinars have fewer than about 200 registrantsNoYou cannot detect anything smaller than a 20-point swing, and no reminder produces one. Decide on judgement, or pool over so many events that the world will have changed underneath you
Traffic source is unstable month to monthNo, not yetThe confounder is larger than the effect. Stabilise the source first, then test. A test run on shifting traffic produces a number that means nothing and gets quoted anyway
You are changing three things at onceNoYou will learn that the package did something, and nothing about which part. If that is genuinely what you want to know, fine - but do not later claim you tested the calls
The offer or the funnel is being rebuiltNoWait. Anything measured across a rebuild measures the rebuild
You already know you will buy it regardlessNoThen the test is theatre. Spend the effort on the script and the disclosure instead, and be honest that this was a decision rather than a finding
The cost is small and the downside is a slightly worse show rateProbably notA test that costs more to run properly than the decision is worth is a bad use of a year. Just try it and watch your own numbers over time
You are being asked to sign a large or long contractYesThis is exactly the case the holdout exists for. Pool across events, pre-register, and make the pilot the thing you are buying first
A vendor is quoting you a specific liftYesAsk which sample it came from and run the arithmetic on this page against it. Most published lifts in this category do not survive that question

What this means if you are thinking of buying this from us

We sell a system that calls and texts webinar registrants, and this page hands you the tool to check whether it works, including the arithmetic showing that most single tests cannot prove it either way. That is deliberate. We would rather you arrive at a conversation knowing what a real answer would cost than sign on the strength of a number we made comfortable.

What we can and cannot show you

We have zero closed clients on this offer. There is no case study, no logo wall, no result to point at, and we are not going to manufacture one. That is precisely why the holdout matters here: with no client evidence to lean on, the only honest thing we can offer is a design for measuring us. We also cannot tell you whether a disclosed AI caller produces the same effect as a human one - the plan-making research used live human callers, and we could find no published study on the disclosed-AI version. That is unknown, not established, and we are not going to write it either way.

One consequence follows for anyone shopping this category. If a vendor tries to talk you out of a holdout, you have learned what you needed to know. The design is not exotic, it costs the vendor nothing but the risk of a null result, and the only reason to resist it is not wanting the number. The same test applies to us - and if we ever ask you to skip it, hold us to this paragraph.

For what the numbers underneath a test like this actually look like, our teardown of every published webinar attendance study gives the sources and the samples. The two questions most operators should settle before spending money on reminders are covered in does charging for a webinar increase attendance - the largest reported lever in the whole corpus, and one nobody has controlled either - and should you offer the webinar replay, which is the most-stated reason people skip. Both of those pages end by telling you to test it on your own list; this is the page that tells you how. The operational side of what we build sits on AI webinar attendance, the evidence review on webinar reminder calls, and the after-the-event work on post-webinar follow-up calls. If you are in the EU or the UK, check what your registration form has to say before anyone may call those registrants on webinar registrant consent, and if you would rather work through the design with us, our contact page books a scoped conversation.

Frequently Asked Questions

Frequently Asked Questions

You split your registrant list at random into two groups before any reminder goes out. One group gets the calls and texts, the other gets nothing beyond your existing baseline. When the session runs you compare the two show up rates. Because the split is random and within a single event, the ads, the offer, the week and the topic are identical across both groups, so a difference in attendance is attributable to the treatment rather than to anything else that changed.

At a 30% baseline, 80% power and a 5% two-sided significance level, you need 546 registrants per arm (1,092 total) to detect a lift to 38%, and 2,129 per arm (4,258 total) to detect a lift to 34%. The formula is n per arm = 7.84 x [p1(1-p1) + p2(1-p2)] / (p2-p1) squared, where 7.84 is (1.96 + 0.84) squared. A lift to 34% is roughly what the Cochrane reminder review implies when its risk ratio of 1.14 is applied to a 30% baseline, so the realistic case is the expensive one.

Almost certainly not. A 400-registrant webinar split 50/50 can only reliably detect a jump from 30% to about 43%. A 600-registrant one gets you to about 41%. Both of those are far larger than any effect the published reminder literature supports, so a null result from a single event tells you your test was too small rather than telling you anything about the treatment. Pool the same design across several events, or report direction and a confidence interval rather than a verdict.

Substantially, because the smaller arm limits the precision of the whole comparison. Suresh and Chandrashekara (2012) give the penalty as N(1+r) squared divided by 4r, so at a 9:1 allocation you need 2.78 times as many registrants in total to answer the same question - about 3,033 instead of 1,092 for a 30% to 38% test. Computing it exactly with the two arms' own variances gives about 2,893, a penalty of 2.65 times. What the lopsided split buys you is a holdout arm that shrinks only modestly, from 546 down to about 303, with a hard floor around 257 no matter how large the treated arm gets. You pay nearly three times the total list for that. A 50/50 split is the efficient choice.

Reached rate, stay rate, booked calls, attended calls, sales and revenue, and complaints or opt-outs - with the arm label attached to every downstream record at the moment it is created. Show up rate alone will overstate the value of the treatment, because the attendees your calls bring in are by definition the ones who would not have come unprompted and may convert worse than self-motivated ones. That last point is our assumption, not a published finding; we could not locate any study measuring it.

Very probably not, and the arithmetic says why. If 4% of attendees buy and you hope for 5%, the same formula gives 7.84 x [0.04(0.96) + 0.05(0.95)] / 0.0001 = about 6,735 attendees per arm. Not registrants - attendees. For most operators that is several years of webinars. Measure sales, report them, and treat them as directional. Anyone claiming a statistically established revenue lift from a reminder system should be asked how many attendees were in each arm.

That is a legitimate position, not an obstacle. The alternative is to alternate the policy across consecutive events - system on, system off, system on, system off - and pool the treated events against the untreated ones. Everyone at a given event is treated the same. It is weaker, because event-to-event variance in traffic source, season and list fatigue is real and now sits inside your comparison, so it needs more events to say the same thing. It is still far better than no test.

Registrants on the same list talking to each other or forwarding the link, so the arms stop being independent. The traffic source changing between pooled events, which is the largest confounder in this subject. Seasonality, if you compare across events instead of splitting within one. Changing the offer mid-test. And the most common one: behaving differently toward the treated arm - extra effort, a warmer email, a setter making an unscheduled call - which quietly adds an untested treatment to one side.

If your webinars have fewer than about 200 registrants, if your traffic source is unstable month to month, if you are changing three things at once, or if you have already decided to buy regardless. In those cases the test either cannot detect anything meaningful or measures something other than what you think. Run one when the contract is large or long, or when a vendor quotes you a specific lift - then check their claimed sample against the arithmetic on this page.

No. We have zero closed clients on this offer, so there is no case study, no logo wall and no result to point at, and we would rather say so than manufacture one. That is exactly why we publish the holdout design and the sample-size arithmetic, including the part showing that a single test of our own product probably cannot settle it either way. If a vendor in this category tries to talk you out of a holdout, you have learned what you needed to know.

Unknown. The strongest evidence for asking someone to form a concrete plan comes from Nickerson and Rogers (2010), a field experiment with N=287,228 that used live human callers and reported a 4.1 percentage point turnout lift among those contacted, with a standard encouragement call having no significant impact. We could find no published study testing whether the same effect survives when the caller identifies itself as an AI assistant. It is untested rather than established, and we are not going to assert it in either direction.

JB
Justas Butkus

Founder & CEO, AInora

Building AI digital administrators that replace front-desk overhead for service businesses across Europe. Previously built voice AI systems for dental clinics, hotels, and restaurants.

View all articles

Ready to try AI for your business?

Hear how AInora sounds handling a real business call. Try the live voice demo or book a consultation.