AveronInstitute

Methods · The journal

Hypothesis Testing Cheat Sheet: Choosing the Right Test as a Green Belt

By the Averon Institute editorial team · September 30, 2026 · 8 min read

There's a specific moment in most Green Belt projects where progress stalls: the data is collected, the question is clear, and the practitioner freezes in front of a menu of tests they half-remember from a slide deck. T-test? ANOVA? Chi-square? The concepts made sense one at a time. Picking between them under real deadline pressure is a different skill, and it's the one this guide is for.

This isn't a re-explanation of what a hypothesis test is or what a p-value means — our guide to hypothesis testing (/blog/hypothesis-testing-without-fear) covers that ground already. This is the decision layer that sits on top: given the data in front of you, which test actually fits.

The first fork: what kind of data do you have?

Every test-selection decision starts with one question, and it isn't about your hypothesis — it's about your data. Is it continuous (a measurement on a scale: cycle time, weight, temperature) or is it categorical (a count sorted into buckets: pass/fail, defect type, shift)? Get this wrong and everything downstream is wrong, no matter how carefully you run the arithmetic. Continuous data leads toward t-tests and ANOVA. Categorical, count-based data leads toward chi-square. Keep that fork in mind before you open any statistical software.

Continuous data: how many groups are you comparing?

Once you know the data is continuous, the next question is simply arithmetic: how many groups are in the comparison?

  1. 01One group against a fixed target — use a one-sample t-test. Example: does the average fill weight differ from the 500g label claim?
  2. 02Two independent groups — use a two-sample t-test. Example: is average handling time different between two shifts, where the people measured are different in each group?
  3. 03Two groups that are the same items measured twice — use a paired t-test. Example: cycle time for the same twenty orders, before and after a process change. Pairing removes the noise that comes from order-to-order differences, which is why treating paired data as independent throws away precision you already paid for.
  4. 04Three or more groups — use ANOVA, not a string of pairwise t-tests. Running repeated t-tests inflates your false-alarm rate with every extra comparison; ANOVA asks one honest question of the whole dataset instead of several lucky-sounding ones.

That's the entire decision tree for continuous data: count your groups, note whether they're paired, and the test is already chosen. The hard part was never the arithmetic — it's noticing which of these four situations you're actually in, which is exactly why so many analyses default to "run a t-test" when the data plainly involves three or more groups.

Categorical data: chi-square, almost by default

When the data are counts rather than measurements — defects by category, complaints by reason code, pass/fail outcomes by shift — the chi-square test is usually the right tool. It compares the counts you actually observed against the counts you'd expect if nothing interesting were happening, and asks whether the mismatch is too large to be chance. The one requirement worth remembering: it needs raw counts, not percentages, and expected counts that aren't too thin — a common rule of thumb wants at least five expected observations per cell, or the result gets unreliable.

When the data won't behave: non-parametric alternatives

T-tests and ANOVA lean on an assumption: the data is roughly bell-shaped. Real process data is often skewed, littered with outliers, or genuinely ordinal — a 1-to-5 satisfaction scale isn't a true numeric measurement, whatever the spreadsheet pretends. In any of those cases, reach for a non-parametric alternative, which typically works on the ranks of the data instead of the raw values, trading a bit of sensitivity for a lot of robustness against messy real-world distributions.

The practical habit that prevents most of this: plot the data before you test it. A histogram or dot plot takes thirty seconds and tells you immediately whether you're looking at something bell-shaped enough for a t-test, or something skewed enough that the parametric assumption would be doing more harm than the extra statistical power is worth.

The test rarely changes your conclusion by much. Picking the wrong one changes whether anyone should trust it.

The one-page version

  • Comparing one average to a fixed number → one-sample t-test
  • Comparing two independent groups' averages → two-sample t-test
  • Comparing the same items measured twice (before/after) → paired t-test
  • Comparing three or more groups' averages → ANOVA, not repeated t-tests
  • Comparing counts across categories → chi-square test
  • Data is skewed, outlier-heavy, or ordinal → a non-parametric alternative to the t-test or ANOVA

Before you trust any of it, qualify the measurement

Choosing the right test is wasted effort if the numbers feeding it are noise from the measurement system rather than signal from the process. Before you baseline anything with a formal test, it's worth confirming the data can be trusted in the first place — our guide to running a first Gage R&R study (/blog/first-gage-rr-study) walks through exactly that check, and it belongs earlier in the Measure phase than any of the tests above.

A short checklist before you report a result

  • State the practical question and choose your significance level before collecting data, not after the numbers arrive
  • Confirm which of the situations above you're actually in — group count and pairing decide the test, not preference
  • Plot the data first; skew and outliers are the signal to reach for a non-parametric test
  • Report the effect size next to the p-value — a statistically significant result can still be too small to matter
  • Treat a non-significant result as "not enough evidence," not as proof that nothing is happening

None of this replaces understanding why a test works — that intuition is what keeps you from misreading the output once you have it, and it's worth building deliberately rather than picking up secondhand. It's core Analyze-phase material in our Green Belt program (/courses/green-belt), built around a full simulated DMAIC project where you practice choosing between these tests on realistic data before it's your own process on the line. If the vocabulary above is still new, our free White Belt (/courses/white-belt) covers the foundational statistical thinking first, and every level uses the same timed, closed-book exam format, with unlimited free retakes while you build the reflex this cheat sheet is meant to shortcut.

Put it into practice

Ready to make it official?

Our Six Sigma belt programs — White through Black — are self-paced, 100% online, and end in a timed, closed-book exam and a credential you can verify and share.