AveronInstitute

Methods · The journal

Design of Experiments: A Gentle First Factorial

By the Averon Institute editorial team · October 2, 2026 · 9 min read

Most process improvement starts with changing one thing and watching what happens. Adjust the temperature, run the batch, check the yield. Adjust it again, run another batch, check again. It feels rigorous because it's methodical, but it quietly wastes most of the information the process is willing to give you — and it can lead you to the wrong answer entirely when two factors interact. Design of experiments (DOE) is the fix: a way of changing several factors at once, in a planned pattern, so you learn more from fewer runs and catch interactions a one-factor-at-a-time approach will never see.

This isn't a statistics lecture. It's the smallest useful version of DOE — the two-level full factorial — explained well enough that you could plan one this week.

The problem with changing one factor at a time

Say a process has two factors you suspect matter: temperature and pressure. The one-factor-at-a-time (OFAT) instinct is to hold pressure fixed, try two temperatures, pick the better one, then hold that temperature fixed and try two pressures. Four runs, a clear winner, done.

The trouble is interaction. What if the best temperature depends on which pressure you're running at — if high temperature helps at low pressure but hurts at high pressure? OFAT testing can easily miss this, because it never tests the combination where the effect actually shows up. It reports a "best" setting that only looks best under the one pressure it happened to test. A factorial design tests every combination on purpose, so an interaction like this becomes visible instead of invisible.

What a two-level full factorial actually is

Take each factor you want to study and pick two settings for it — a "low" and a "high." These don't have to be extreme; they just need to be far enough apart that a real effect would show up above the process's normal noise. With two factors at two levels each, there are four possible combinations. With three factors, there are eight. The pattern is 2 raised to the number of factors, which is exactly why this family is called a "2^k factorial" — k is just the factor count.

For the temperature/pressure example, a full factorial means running all four combinations: low temp/low pressure, low temp/high pressure, high temp/low pressure, and high temp/high pressure. Ideally each combination gets run more than once (replicated) so you can see how much the result varies even when nothing changes — without that, you can't tell a real effect from ordinary process noise.

  1. 01Pick 2-3 factors you believe affect the outcome you care about — not everything you can think of, just your strongest suspects.
  2. 02Choose a low and high setting for each factor, spaced wide enough to matter but still inside safe, realistic operating limits.
  3. 03List every combination — 4 runs for two factors, 8 for three — and randomize the order you actually run them in.
  4. 04Replicate each combination at least twice if the process and budget allow it, so you can separate signal from noise.
  5. 05Measure the same output for every run, using a measurement system you already trust.

Reading the results without a statistics degree

Once the runs are in, the first thing to look at isn't a p-value — it's a simple average. For each factor, compare the average output when it was at its high setting against the average output when it was at its low setting. A big gap between those averages means the factor likely matters; a small gap means it probably doesn't, at least not within the range you tested.

The interaction check follows the same logic one level up: does the effect of temperature change depending on which pressure setting it's paired with? If the high-temperature advantage is large at low pressure but small or reversed at high pressure, that's an interaction, and it's information OFAT testing structurally cannot give you. Software will eventually quantify how confident you can be in these effects, but you can usually eyeball which factors and interactions are worth following up on before you ever open a statistics package.

A factorial design doesn't just ask whether each factor matters. It asks whether the factors have an opinion about each other — and that second question is the one OFAT testing can't answer at all.

A small worked example

Imagine a bakery troubleshooting a crust that's inconsistent. Two suspects: oven temperature (325°F vs. 375°F) and bake time (18 minutes vs. 24 minutes). That's a 2^2 design — four combinations, each baked three times for a replicated, 12-run experiment. The average crust score at each of the four combinations gets compared: if the jump from 325°F to 375°F is small at 18 minutes but large at 24 minutes, time and temperature are interacting, and the real recommendation is a specific pairing — not "always use the higher temperature" or "always use the longer time" in isolation. That's the kind of nuanced, defensible answer a factorial design produces and a one-at-a-time test would have missed.

Where this fits in a DMAIC project

DOE belongs in the Improve phase, after root cause analysis has narrowed down a reasonably small list of candidate factors — running a factorial on ten suspected causes gets expensive fast, since the run count doubles with every added factor. A Pareto chart, a fishbone diagram, or an FMEA (/blog/fmea-step-by-step-scoring-rpn-without-guesswork) is usually what gets you from a long list of theories down to the two or three factors worth testing deliberately. Once you have that short list, understanding variation (/blog/understanding-variation) first matters too — a factorial design is wasted on a process so noisy that no signal could show through it regardless of what you change.

  • Narrow your factor list before designing the experiment — DOE tests suspects, it doesn't generate them
  • Keep low/high settings realistic and safe; you can always widen the range in a follow-up experiment
  • Randomize run order to keep drift, warm-up effects, or shift changes from masquerading as a factor effect
  • Replicate when you can — a single run per combination can't separate a real effect from noise
  • Treat the interaction check as the main event, not an afterthought; it's the result OFAT testing can't give you

A two-level factorial with two or three factors is deliberately the simplest entry point into DOE — there's a much larger toolkit beyond it (fractional factorials, response surface designs, blocking) for when the factor list grows or the relationship isn't a straight line between your two settings. But most teams get real, defensible answers out of a design this small, and it's exactly where our Black Belt program (/courses/black-belt) picks up, building from this first factorial into the fuller DOE toolkit used on harder improvement problems. If root cause tools and process baselining are still new ground, our Green Belt program (/courses/green-belt) covers that foundation first, and our free White Belt (/courses/white-belt) is the place to start from zero — every level shares the same timed, closed-book exam format with unlimited free retakes while the ideas settle in.

Put it into practice

Ready to make it official?

Our Six Sigma belt programs — White through Black — are self-paced, 100% online, and end in a timed, closed-book exam and a credential you can verify and share.