The p-value is the probability of obtaining data at least as extreme as the data you observed, assuming the null hypothesis is true — assuming, that is, that nothing has changed. It is a measure of surprise: how implausible would this result be if the process were behaving exactly as before? What it is not, despite decades of misuse, is the probability that your conclusion is correct.
How it’s used
In practice, the p-value is compared against a significance level chosen in advance, most commonly 0.05. If the p-value falls below that threshold, the result is called statistically significant and the null hypothesis is rejected: the data are too surprising to blame on noise. If it falls above, the team withholds judgment — the evidence is insufficient, which is not the same as evidence of absence. Six Sigma teams lean on this machinery throughout Analyze, where suspected causes are tested against data, and again in Improve, where a pilot must prove it actually changed the process.
Three misreadings cause most of the trouble. A p-value of 0.03 does not mean there is a 3 percent chance the null hypothesis is true. It does not mean the improvement has a 97 percent chance of being real. And it says nothing about size: a trivial difference can earn a tiny p-value if the sample is large enough, while an important difference can miss the threshold if the sample is small. The p-value reports surprise, never importance.
A worked example
Picture a packaging plant comparing seal strength from two adhesive suppliers. Samples from each are pulled and tested, and the averages differ. The test returns a p-value of 0.03. The correct reading: if the two suppliers truly performed identically, samples this far apart would occur only about 3 percent of the time — so the difference is probably real. The team then asks the question the p-value cannot answer: is the difference large enough to justify switching suppliers? The figures are illustrative, but the two-step reasoning — real first, then meaningful — is exactly how trained practitioners work.
- A p-value of 0.06 is not proof of no effect, and 0.04 is not proof of a large one — the threshold is a convention, not a cliff.
- Always report an effect size alongside the p-value; readers need to know how big, not just how surprising.
- Never run many tests and publicize only the one that cleared 0.05 — that is how phantom causes get institutionalized.
- Large samples make everything significant; ask whether the difference matters in the customer’s units.
- Decide the significance level before the data arrive, and leave it alone.
A p-value measures how surprised you should be — deciding whether to care is still your job.
P-values appear the moment hypothesis testing does: in the Analyze phase of our Green Belt curriculum ($299, 35 hours), where they are drilled through the tools that generate them — t-tests, ANOVA, chi-square and regression. Black Belt training pushes further into power and sample-size design, the machinery that makes p-values trustworthy.
Put it into practice
Ready to make it official?
Our Six Sigma belt programs — White through Black — are self-paced, 100% online, and end in a proctored exam and a credential you can verify and share.