Sample Size Calculator

Set the precision you want and read back the sample it takes to get there. Survey margins, A/B test lifts, mean estimates, response rates, dropout and cluster design effects all sit in one panel, and the trade-off table underneath prices every extra point of precision before you commit to fieldwork.

Sample size worksheet

I am sizing
Start from
± %
How far the result is allowed to sit from the truth.
How often the interval covers the true value on repeat sampling.
%
Leave at 50 when you have no prior estimate. It costs the most and never underestimates.
Fill this in for a school, a payroll or a customer list. Blank means anything above roughly 100,000.
Fieldwork reality: response rate, clustering
%
Drops the invite count out the other side.
1.0 for simple random sampling. Higher when you sample whole villages, schools or offices.
Build the design effect from clusters
Design effect equals 1 + (cluster size - 1) × ICC. Health surveys often land between 1.5 and 3.
Start from
From a pilot, past data or a published study. Everything below hangs off this one number.
±
Same units as the standard deviation, not a percentage.
The t distribution is used, so small samples pay a penalty the z formula hides.
Optional. Applies the finite population correction.
Fieldwork reality: response rate, clustering
%
Turns completed measurements into invitations sent.
Raise it above 1 when the sample arrives in clusters rather than one person at a time.
Start from
%
The control arm, measured over a full weekly cycle rather than a good day.
%
A 10% relative lift moves 5.00% to 5.50%.
Your odds of catching the lift when it is genuinely there.
Your tolerance for calling a flat result a winner.
Traffic split, dropout, test duration, tails
×
1 is an even split. 0.25 sends a quarter of the control traffic to the variant.
Adds a run length in days and whole weeks.
%
Bounced sessions, uninstalls, people who never reach the step you measure.
Matches the chi-square test with Yates correction. Raises the count, most noticeably on small samples.
Start from
In real units. Minutes saved, points scored, grams lost.
Pooled across both arms. The ratio of these two fields is Cohen's d.
Allocation, dropout, tails
×
Uneven arms always need more people in total than an even split.
%
Recruit above the analysed count so the final numbers still hold.
%
%

What tightening the number costs

Why 385 turns up in every survey you have ever read

Ask for a national poll and the quote comes back at 1,000 interviews. Ask a statistician for the bare minimum and you get 385. The second number is not folklore. It falls out of one calculation with three assumptions baked in: 95 percent confidence, a margin of error of 5 points, and a result sitting at 50 percent.

n = 1.962 × 0.5 × 0.5 ÷ 0.052 = 384.16, rounded up to 385.

The 50 percent assumption does the heaviest lifting. Variance for a percentage is p(1-p), which peaks at exactly 0.5 and falls away fast on either side. A result near 90 percent carries variance of 0.09 rather than 0.25, so the same 5 point margin needs 139 responses instead of 385. Nobody knows the answer before fielding, so 50 percent stays as the safe default. It never asks for too few.

What 385 does not survive is a second question. Split those responses by four regions and each region rests on 96 replies, worth a margin near 10 points. Break them by age band as well and the cells fall into double digits. Sizing on the headline number and then reporting the crosstabs is the single most common way a survey budget gets wasted.

Precision is priced on a square root, and it is brutal

Sample size sits under a square root, so halving the margin of error quadruples the sample. Not double. Four times. This is the fact worth carrying into a budget meeting, because the gap between a 5 point margin and a 1 point margin is not five times the money, it is twenty five times.

Responses needed at 95 percent confidence, worst-case 50 percent split, unlimited population.
Margin of errorResponsesAgainst the row above
±10 points97Baseline
±5 points3854× more
±2.5 points1,5374× more
±1 point9,6046.2× more
±0.5 points38,4154× more

Confidence level moves the number too, though less violently. Going from 95 to 99 percent lifts 385 to 664, a 72 percent surcharge for four extra points of certainty. Dropping to 90 percent saves you 30 percent and takes you to 271. Most commercial research settles at 95 because the cost curve turns steep right after it, not because the number carries any deep meaning.

The practical read. If a stakeholder asks for a 2 point margin on a 5 point budget, the answer is not a negotiation. It is a factor of 6.25 in fieldwork cost. Show the trade-off table above the panel and let the number make the argument.

Your population size matters far less than you think

The intuition says a country of 60 million needs a bigger sample than a town of 60,000. The arithmetic disagrees. Both need about 385. Once the population passes roughly 20,000, the finite population correction stops moving the answer in any way you would notice.

Responses for a ±5 point margin at 95 percent confidence, corrected for population.
PopulationResponses neededShare of the population
50021843.6%
1,00027827.8%
10,0003703.7%
100,0003830.4%
60,000,0003850.0006%

The correction only bites when your sample is a large slice of the whole. A 480 person company auditing its own staff genuinely needs fewer replies than an open web poll, and the calculator applies n = n₀N ÷ (n₀ + N - 1) whenever you fill the population field. Leave it blank and nothing is deducted.

One caveat worth stating plainly. The correction assumes you drew names at random from a complete list. If you emailed everyone and analysed whoever replied, you have a self-selected sample, and no correction repairs that. The people who answer staff surveys are not a random draw from the payroll.

Measuring one number and comparing two numbers are different sports

A survey estimates a single percentage. An A/B test asks whether two percentages differ. The second job needs an order of magnitude more data, and teams who carry the 385 figure across from surveys run tests doomed before the first visitor lands.

Take a checkout converting at 5 percent. You want to catch a 10 percent relative improvement, which moves it to 5.5 percent. At 80 percent power and the usual 0.05 threshold, you need 31,234 users per arm, 62,468 in total. That is 162 times the survey figure, for a change most product teams would call modest.

Three levers move that number, and only one of them is free:

  • The size of the lift. Doubling the lift you are willing to detect cuts the sample to roughly a quarter. Chasing 20 percent instead of 10 percent takes you from 31,234 to 8,158 per arm. The word "roughly" is doing real work here, because the variance of the variant shifts with its own rate.
  • Power. Moving from 80 to 90 percent adds about a third to the count. Most teams keep 80 percent, which quietly means one real winner in five slips past unnoticed.
  • Baseline rate. Rare events are expensive. The lower the conversion rate, the more traffic each percentage point of relative change demands, which is why testing on a 0.4 percent enterprise demo form usually fails before it starts.

Relative and absolute lifts are not interchangeable. On a 5 percent baseline, a 10 percent relative lift is 0.5 percentage points. A 10 percentage point lift means tripling the rate. Enter the first as a relative lift, the second as points. Confusing the two changes the required sample by a factor of 400 in this example, and the toggle above the input exists to stop exactly that mix-up.

Sample size is not the number of people you contact

Everything above returns the count that ends up in the analysis. Four things sit between that number and your recruitment plan, and each one inflates it.

Response rate
Email surveys land between 10 and 25 percent in most sectors. At 20 percent, 385 completed responses means 1,925 invitations. Fill the response rate field and the invite count appears in the panel alongside the sample.
Dropout
Clinical trials, app onboarding funnels and multi-week studies all lose people after enrollment. Sizing at 200 per arm with 20 percent attrition means recruiting 250. The analysis needs the survivors, not the sign-ups.
Design effect
Sampling whole schools, villages or offices instead of individuals means neighbours answer alike, and each response carries less independent information. A design effect of 2 doubles the requirement. Feed in the intracluster correlation and the cluster size and the panel works the multiplier out for you.
Subgroup reporting
Any group you plan to report separately needs its own sample. Five customer segments at ±5 points each is five separate calculations, not one. Size for the smallest cell you intend to quote.

Stack all four and the gap grows fast. A cluster survey with a design effect of 1.5, a 30 percent response rate and three reported regions runs to 577 per region, 1,731 completed responses, and 5,770 households approached. Sizing that project off the 385 headline would have missed by a factor of fifteen.

Where this calculator stops

Every formula here rests on the normal approximation, on independent observations, and on a single analysis run at the end of data collection. Break any of those and the number stops applying.

  • Peeking at results as they arrive. A fixed sample size assumes one test at the end. Watching a dashboard and stopping the moment it crosses the threshold pushes the real false positive rate past 5 percent, often above 20 percent. Sequential and group sequential designs handle that properly, and they need different boundaries than anything on this page.
  • Very small or very large proportions. Below about 1 percent, the normal approximation drifts from the exact binomial. The panel flags it when the expected count of events drops under five. Clopper-Pearson intervals give the exact answer, and this tool does not compute them.
  • Time-to-event outcomes. Survival studies are powered on the number of events, not the number of participants. Log-rank sample sizes need the hazard ratio, the accrual period and the follow-up window. None of those inputs exist here.
  • Repeated measures and crossover designs. Measuring the same person twice changes the variance structure and usually reduces the count needed. Paired designs, mixed models and cluster randomised trials with baseline covariates all need their own treatment.
  • Equivalence and non-inferiority. Proving two treatments are close enough inverts the hypothesis and needs a margin you set in advance. Running an ordinary two-sided calculation for that question gives the wrong answer in a direction that flatters the study.
  • Two-group mean calculations round up. The iteration uses central t quantiles on both the alpha and beta side, which is the conventional approach and slightly conservative against the exact noncentral t. Expect it to sit one or two participants per arm above a specialist package like G*Power at small sample sizes.

For hypothesis testing after the data lands, the t-test calculator and the p-value calculator pick up where this one stops, and the confidence interval calculator turns your collected sample back into a range. Everything on this page runs in the browser, so nothing you type is uploaded or stored anywhere.

Sample size questions worth answering before you field

Common sticking points on margins, power, response rates and the limits of the formulas.

What sample size do I need for a 95 percent confidence level?

Confidence level alone does not fix a sample size. Pair it with a margin of error. At 95 percent confidence and a ±5 point margin you need 385 responses, at ±3 points you need 1,068, and at ±1 point you need 9,604. All three assume a worst-case 50 percent result and a large population.

Is 100 responses enough for a survey?

It buys a margin of error near ±9.8 points at 95 percent confidence. A result of 60 percent on 100 replies means the truth sits somewhere between 50 and 70 percent, so 100 works for a rough temperature check and fails for anything that has to distinguish a 55 percent result from a 45 percent one.

Why does my A/B test need tens of thousands of users when a poll needs 385?

A poll estimates one number. A test compares two, so it carries the uncertainty of both arms and has to separate them from noise. On a 5 percent baseline chasing a 10 percent relative lift, that works out to 31,234 users per arm at 80 percent power. Larger lifts and higher baseline rates bring it down quickly.

What is statistical power and why is 80 percent the default?

Power is your chance of detecting a real effect. At 80 percent, one genuine winner in five goes undetected. The convention traces back to Jacob Cohen treating a false negative as roughly four times less costly than a false positive. It is a budget compromise, not a law, and 90 percent is worth the extra third of traffic on decisions you cannot revisit.

Do I need a bigger sample for a bigger population?

Barely. A town of 60,000 and a country of 60 million both need about 385 responses for a ±5 point margin. The finite population correction only matters when your sample covers a large share of the whole, such as 218 replies from a 500 person company.

Should I use relative lift or percentage points for my A/B test?

Use relative lift when the goal reads as "improve conversion by 10 percent". Use percentage points when it reads as "move conversion from 5 percent to 6 percent". On a 5 percent baseline, a 10 percent relative lift is 0.5 points, so picking the wrong one changes the required sample enormously. The toggle beside the field switches between them and restates the target rate underneath.

How do I set the design effect for a cluster survey?

Compute 1 + (cluster size - 1) × ICC. Sampling 20 households per village with an intracluster correlation of 0.02 gives 1.38, so multiply your sample by 1.38. Household health surveys often land between 1.5 and 3. Enter the ICC and cluster size in the fieldwork section and the panel applies the multiplier.

What happens if I stop my test early once the p-value drops below 0.05?

Your false positive rate climbs well past the 5 percent you set. Checking daily on a four week test pushes it above 20 percent. Fixed sample sizes assume one analysis at the end, so either wait for the full count or switch to a sequential design with adjusted boundaries.

Why does the mean calculator return a larger number than the z formula?

It iterates with t distribution quantiles rather than a fixed z value. For a standard deviation of 10, a ±2 margin and 95 percent confidence, the z formula returns 97 while the t based iteration settles on 99. The gap widens as the sample shrinks and is the honest figure at small n.

Does the response rate change the statistical result?

No. It changes how many people you contact, not how many you need in the analysis. At a 20 percent response rate, 385 completed replies means 1,925 invitations. A low response rate does raise a separate concern, since the people who reply differ from the people who ignore you, and no sample size fixes that bias.