Statistical power is the probability that a test will detect an effect of a given size when that effect truly exists. It is the sensitivity of a study — the odds that the fire alarm rings when there is an actual fire. A test with 80 percent power against a meaningful improvement will find it four times out of five; the fifth time, a real gain gets dismissed as noise. Power is where honest studies are made, or quietly broken.
How it works
Power lives in a four-way tradeoff with sample size, effect size and significance level. Larger samples raise power. Larger effects are easier to detect. A stricter significance level — fewer false alarms — lowers power unless data are added to compensate. Practitioners conventionally plan for at least 80 percent power against the smallest effect worth acting on, then solve for the sample size that delivers it. Done in that order, before collection, the study is an instrument built for its question. Done never — the common case — the study’s sensitivity is an accident, and its negative results are uninterpretable.
That last point deserves emphasis: a non-significant result from an underpowered test is not evidence that the change failed. It is evidence that the study could not see. The distinction saves good improvements from bad verdicts.
A worked example
Suppose a fabricator pilots a new fixture expected to reduce drilling defects modestly but meaningfully. The team runs a small pilot, and the improvement — visible in the raw numbers — is not statistically significant. Before declaring failure, a Black Belt computes the power of the test that was actually run: against an improvement of the expected size, it had well under a coin flip’s chance of detecting anything. The verdict was baked in before the first part was drilled. The pilot is redesigned with a sample size that gives the fixture a fair hearing. A generic example of an endemic mistake.
- Compute power before the study, while it can still change the design.
- Never read “not significant” from an underpowered test as “no effect” — the test may simply have been blind.
- Size the study against the smallest effect that matters, not the most optimistic one.
- If the affordable sample cannot reach adequate power, say so before running — an honest maybe beats a false no.
- Post-hoc power computed from the observed effect adds nothing the p-value did not already say.
An underpowered test is a verdict written before the evidence arrives.
Power analysis is a Black Belt discipline in our curriculum — part of 60 hours of advanced statistics for $499, examined in a 150-question proctored test with a 70 percent pass mark. Green Belts meet the idea earlier, when they learn why sample size decides what a study can see.
Put it into practice
Ready to make it official?
Our Six Sigma belt programs — White through Black — are self-paced, 100% online, and end in a proctored exam and a credential you can verify and share.