Detectable difference and statistical power
Could your study detect the difference that matters?
Compare the sample you can recruit with the difference you need to detect. Plan separate audiences or test cells using percentages or average scores.
Start with the comparison you intend to make.
Use assumptions about the true population difference before collecting data. The opening example uses 400 usable people per group and a 50% baseline. Add your own meaningful difference to assess whether that sample supports your decision.
Start with your own inputs, or try a worked example.
Model-based planning estimate
Smallest increase at 80% power
Group A baseline: 50%. Two-sided test; 5% significance threshold per comparison; 80% target power.
Enter a meaningful difference to compare this sample with your business requirement and calculate the sample needed.
Two-proportion normal approximation, with pooled null and separate alternative variances; no continuity correction. Method and scope.
This estimates detection of an audience difference. It does not establish what caused the difference. Selection, coverage, and measurement bias are not quantified.
Compare detection targets
Detectable increase at your entered sample sizes, in percentage points.
80% and 90% are alternatives, not universal requirements. Higher power usually needs more usable data.
Detection becomes more likely as the true difference grows.
Power for a true increase, using 400 people in A and 400 in B. The curve uses the same method as your result.
The slider explores assumptions; it does not change your meaningful difference. An 80% result means about 80 detections in 100 comparable studies under this model.
View curve values as a table
| Difference (percentage points) | Power |
|---|---|
| 0 | 5% |
| 1.845 | 8.2% |
| 3.691 | 18.1% |
| 5.536 | 34.8% |
| 7.382 | 55.3% |
| 9.227 | 74.7% |
| 11.073 | 88.5% |
| 12.918 | 95.9% |
| 14.764 | 98.9% |
Where would 100 more respondents help?
| Allocation | A + B | Difference |
|---|---|---|
| Current | 400 + 400 | 9.84 |
| Add 100 to A | 500 + 400 | 9.35 |
| Add 100 to B | 400 + 500 | 9.34 |
These scenarios change available bases only. Recruitment costs and group availability also matter.
Methods and your plan
The download records your assumptions, rounded usable bases, achieved power, and scope. These are planning calculations, not observed results.
Interactive calculator loading.
Keep three different questions separate.
What difference matters?
A five-point improvement might change an investment decision even when the planned study has little chance of detecting it. Choose the business threshold from the decision and evidence, then assess the sample against it.
What can this sample detect?
A minimum detectable difference is tied to a stated probability of detection. At 80% power, the model still allows about one in five comparable studies to miss that true difference. It is not a hard significance boundary.
What did the study find?
Observed results need an appropriate test, estimates, and uncertainty intervals. Power calculated from the same observed difference does not resolve an inconclusive result. This tool plans future comparisons.
Match the calculation to the study.
A subgroup's usable base may be much smaller than the full survey. Repeated respondents, rare outcomes, and multiple comparisons also change what a study can establish.
Discuss your sample and comparisonsWhy is a margin of error different?
With 400 independent observations and percentages near 50%, a simple 95% margin of error is about ±4.9 points for one percentage. For the difference between two groups of 400, it is about ±6.9 points. Detecting an increase from 50% with 80% power needs about 9.84 points under this planner's model.
These answer different questions: precision around an estimate and the chance of detecting a specified true difference. They are not interchangeable sample-size requirements. Plan precision for a single percentage.
What if the same people answer twice?
The responses are paired. For means, variability in the within-person differences matters. For percentages, the shares changing in each direction matter. Counting each person in two independent groups gives the wrong model.
Partial overlap also needs the covariance between groups. The calculator withholds independent-group results when you select these designs. Mean-comparison assumptions.
Does 80% power prove an improvement exceeds the business threshold?
No. This planner tests against zero. Power to detect a true five-point difference does not mean power to establish that the difference exceeds five points. That needs a different null boundary and a plausible true effect beyond it. Likewise, showing similarity requires equivalence bounds; a nonsignificant difference does not establish equivalence.
How should we plan several audiences or concepts?
Specify the comparisons that will support the decision. Four cells produce six possible pairs, or three comparisons if each alternative is tested only against one control. The Bonferroni option allocates the family significance threshold across that declared count.
The power result applies to the pair entered here. It is not the probability of detecting every effect, and this pair's total is not the whole study budget when comparisons share a control. Evaluate the actual contrasts and reuse shared groups in the recruitment plan. Plan concept and message allocation. Multiple-testing methods.
Can a larger sample correct an unrepresentative one?
It can reduce random uncertainty under a model, but it does not correct coverage, selection, measurement, or missing-data bias. Random assignment helps evaluate treatments within the study; it does not automatically make participants representative of the wider market. Report the recruitment and analysis assumptions alongside any precision claims. AAPOR disclosure guidance.
Methods and research
Named calculations. Explicit assumptions.
Percentages: a large-sample two-proportion calculation with a pooled variance under no difference and separate variances under the assumed alternative. Both rejection tails are included for a two-sided test. There is no continuity correction. Expected yes/no counts below 10 trigger a review note; they need a suitable finite-sample test or simulation.
Means: a Welch-Satterthwaite noncentral-t power approximation using each group's assumed standard deviation. It includes a t critical value and uncertainty in the variance estimate, rather than substituting a normal cutoff. It remains an approximation to Welch's test; skew, outliers, and complex survey designs need further assessment.
Required sample sizes are solved numerically, rounded to usable people, and checked again for achieved power. Searches require an effective base of at least 10 and stop at 1,000,000 usable people per group. A failed search reports that limit. Optional design effects reduce the effective bases for approximate sensitivity; they do not validate weighting or replace cluster-specific degrees of freedom.
The distribution calculations are checked against SciPy and Statsmodels under matching settings. The sources below were reviewed September 29, 2026.
If you already have respondent weights, the weighting diagnostic shows their dispersion and concentration. Its weight-only effect is not automatically a valid design effect for this power calculation.
- Lakens (2022). Sample Size Justification. Collabra: Psychology.
Distinguishes power-based planning, sensitivity analysis, and the smallest effect that matters. A budget alone does not define a meaningful difference.
- Statsmodels. Power for two independent proportions.
Documents the pooled variance under the null and separate variances under the alternative used by this planner. Matching numerical settings are essential when comparing tools.
- NCSS. Two-Sample T-Tests Allowing Unequal Variance.
Explains the unequal-variance mean comparison and its distributional assumptions. Our named noncentral-t planning approximation is not a claim to reproduce every PASS option or an exact Welch power integral.
- SciPy. Noncentral Student’s t distribution.
Defines the distribution used in the mean-power approximation. Independent SciPy calculations check the browser engine's rejection probabilities.
- Cook and colleagues (2018). DELTA² guidance on choosing the target difference. Trials.
Supports justifying the target difference and testing sensitivity to uncertain baseline rates and variability. Its planning principles are adapted here; clinical thresholds are not imported into market research.
- R documentation. Adjust P-values for Multiple Comparisons.
Distinguishes Bonferroni from other multiple-testing methods. This planner divides the declared family threshold by the number of planned comparisons.
- AAPOR. Transparency Initiative.
Precision claims need disclosed methods and assumptions. Numerical power does not measure selection bias or establish that an opt-in sample represents a population.