AveronInstitute

Methods · The journal

Running Your First Gage R&R Study

By the Averon Institute editorial team · November 24, 2025 · 5 min read

Every Six Sigma project runs on data, and every piece of data passes through a measurement system on its way to you. A caliper, a scale, a stopwatch, a reviewer deciding pass or fail — each one adds its own variation to the numbers. Before you analyze a process, you owe yourself an honest answer to an uncomfortable question: how much of the variation in my data comes from the process, and how much comes from the act of measuring it?

Gage R&R is the standard way to answer that question for continuous, measured data. The name is short for repeatability and reproducibility — the two ways a measurement system betrays you. The study belongs to the Measure phase of DMAIC for a blunt reason: analysis performed through a noisy measurement system is analysis of the noise.

A first Gage R&R intimidates people, mostly because the output arrives dressed in statistics. The study itself is simple: a structured exercise in which several people measure the same items several times, and the arithmetic sorts out who — or what — is responsible for the disagreement.

Repeatability and reproducibility, in plain terms

Repeatability asks: if the same person measures the same item with the same instrument several times, how much do the results differ? Ideally not at all; in practice, always somewhat. Poor repeatability points at the equipment or the method — a loose fixture, an ambiguous procedure, an instrument too coarse for the tolerance it polices.

Reproducibility asks: when different people measure the same item, does the answer depend on who measured? If one inspector reads consistently high and another consistently low, every measurement carries the signature of its operator. Poor reproducibility usually points at technique and definitions — where exactly to place the probe, when to take the reading, what counts as the edge of the part. The study separates the two components, and the separation matters because the remedies differ: you fix repeatability with better equipment and method, and reproducibility with better operational definitions and training.

When you need one

Run a Gage R&R before you baseline a critical metric, before you charter a project on data whose quality nobody has checked, and after anything about the measurement changes — new instrument, new procedure, new people. The most common discovery in Measure is not that the process is worse than believed. It is that the measurement system cannot reliably tell good parts from bad ones, which silently invalidates every chart built on it. An hour of study protects a month of analysis.

Designing your first crossed study

The workhorse design is the crossed study, so called because every appraiser measures every part. A common and sensible starting layout uses about ten parts, three appraisers, and two to three trials each — enough structure to separate the sources of variation without turning the exercise into a career.

  1. 01Select about ten parts that span the real operating range of the process — including some near the specification limits, not ten near-identical items from one good batch.
  2. 02Recruit the people who actually take this measurement in daily work — not the engineer who designed the gage, and not the most careful person in the room.
  3. 03Number the parts where appraisers cannot see the labels, and have each person measure every part in random order.
  4. 04Repeat the full round two or three times per appraiser, re-randomizing the order each time, so nobody can recognize a part and remember their previous answer.
  5. 05Record every reading as it happens, and resist the urge to remeasure anything that looks odd — the odd readings are the point.

The blinding and the randomization are not ceremony. If appraisers know which part is which, they remember answers, and the study ends up measuring memory instead of the gage.

Reading the results

The analysis — from statistical software or a spreadsheet template — splits total observed variation into part-to-part variation and measurement-system variation, then expresses the measurement share as a percentage. Widely used rules of thumb, published in the automotive industry’s measurement system analysis guidance, treat a system consuming under ten percent of the variation as acceptable, ten to thirty percent as marginal and usable only with justification, and over thirty percent as unacceptable for the decision at hand.

Look also at the number of distinct categories, which estimates how many groups of parts the system can genuinely tell apart. A system that can only sort parts into two or three buckets may be adequate for crude screening but cannot support the fine comparisons an improvement project depends on. And read the components separately: a study that fails on repeatability sends you to the instrument and the method, while one that fails on reproducibility sends you to definitions and training.

When the system fails the study

A failed Gage R&R feels like bad news and is actually a gift: you found the problem before it fabricated your conclusions. The fixes follow the diagnosis. For repeatability, improve the fixturing, tighten the measurement procedure, or move to a finer instrument. For reproducibility, write an operational definition so specific that two strangers would measure the same way, train against it, and re-run the study. And when the measurement is a human judgment — pass or fail, defect or acceptable — the same logic applies through an attribute agreement study, which checks whether appraisers agree with themselves, with each other, and with a known standard.

Before you ask what the data says, ask whether the data can speak at all.

Where measurement fits in the method

Gage R&R is the moment a project stops taking its numbers on faith — the difference between decorating a decision with data and actually resting a decision on it. The reasoning belongs to the Measure phase, and it is taught, not absorbed. Our free White Belt covers the DMAIC roadmap and the vocabulary of measurement in about six hours with a 30-question exam. The Green Belt — 35 hours, a full simulated project, and a proctored 100-question exam — puts the tools in your hands, measurement system analysis included. Run the study once on a gage you have always trusted. The result, either way, will change how you read every number after it.

Put it into practice

Ready to make it official?

Our Six Sigma belt programs — White through Black — are self-paced, 100% online, and end in a proctored exam and a credential you can verify and share.