AveronInstitute

Methods · The journal

Building a Data Collection Plan You Can Defend

By the Averon Institute editorial team · November 2, 2025 · 6 min read

Every improvement project reaches the moment when someone challenges the numbers. A department head disputes the defect count. An operator points out that the two shifts record downtime differently. A sponsor asks, reasonably, how you can be sure. In that moment you are holding either a data collection plan you can defend — or a spreadsheet you have to apologize for.

The difference is made weeks earlier, before the first data point is recorded. A defensible plan is not longer or more sophisticated than an indefensible one; it is simply built by answering a series of questions in the right order, in writing, before collection starts. This article walks through those questions in the order they should be answered.

Begin with the question, not the spreadsheet

Weak data collection starts with what is easy to get: the system already exports this report, so the project uses this report. Strong collection starts from the decision the data must support. What question are we answering — how often does the defect occur, where does the delay concentrate, did the change work? Each question implies what must be measured, at what point in the process, over what period. Write the question at the top of the plan, literally, because every later argument about the data gets settled by pointing at it. Data collected without a question attached tends to answer nothing while consuming everyone’s patience.

Operational definitions: the heart of the plan

An operational definition turns a word into a procedure — a rule so specific that two people observing the same event will record it the same way. What exactly counts as a defect? When precisely does the clock start on cycle time: when the email arrives, when it is opened, or when it is logged? Does a form corrected by the same clerk who wrote it count as rework, or only one sent back from downstream? Every ambiguous term in your metric is a place where two collectors will silently diverge, and their disagreement will surface later as data no one trusts.

The test of an operational definition is brutal and simple: hand it to someone unfamiliar with the project along with a handful of real examples, including the awkward borderline ones, and see if their classifications match yours. Where they diverge, the definition — not the person — needs work. Definitions earn trust at the borderline, never at the obvious cases.

Choose the data type deliberately

Data comes in two broad kinds. Continuous data measures on a scale — minutes, millimeters, dollars — while attribute data counts categories: pass or fail, late or on time, defect type A, B, or C. The practical rule is to take continuous data when you can get it at reasonable cost, because it carries far more information per observation. Recording that a delivery was late tells you it missed the mark; recording that it was late by forty minutes tells you how badly, whether things are drifting, and how much improvement would move it inside the limit. Attribute data has its place — some things genuinely are categories — but many teams count pass/fail out of habit when the underlying measurement was available all along.

Sampling and stratification

Measuring everything is rarely necessary and often impossible, so most plans sample — and sampling is where hidden bias loves to live. A sample of invoices pulled from the top of the pile overweights the newest ones. Data collected only on day shift says nothing about nights. The guiding principle is that the sample must be drawn in a way that gives the process’s full variety a chance to appear: across shifts, days of the week, product types, locations, and people, over a period long enough to include the process’s normal rhythms.

Stratification is the planning habit that makes analysis possible later: recording the categories each observation belongs to — which shift, which machine, which clerk, which customer type — alongside the measurement itself. Those tags cost seconds at collection time. Added afterward, they cost weeks, and usually cannot be reconstructed at all. Most of the analytic power in the Analyze phase comes from comparisons across strata, and those comparisons are only available if the plan captured the tags.

Trust the measurement before you trust the data

Before relying on any measurement, ask how much of the variation you see comes from the measuring rather than the process. Hand the same items to two people with your operational definition and compare their records. Have the same person re-measure the same items later without knowing their first answers. If the same event, measured twice, produces different numbers, your data will faithfully record noise. Formal measurement-system studies exist for high-stakes work, but even the informal version — a structured agreement check before launch — catches the worst problems, and skipping it is the single most common hole in otherwise careful plans.

The plan on one page

A data collection plan is typically one table. For each metric, a row answers a fixed set of questions — and a row that cannot be completed is a warning worth heeding before launch, not after.

  • What is measured, with its operational definition
  • The data type — continuous or attribute — and the unit of measure
  • Where in the process the measurement is taken
  • Who records it, using what form, gauge, or system
  • When and how often, and the sampling rule if not every unit
  • The stratification tags recorded alongside each observation
  • How the measurement itself was checked before launch

Pilot before you commit

Run the plan for a day or two before trusting it for weeks. A pilot surfaces everything the conference room missed: the form with no room for the common case, the definition that collapses on a borderline item, the step where collection takes so long that people quietly skip it. Fix the plan, brief the collectors on what changed and why their accuracy matters more than flattering numbers, and only then start the real baseline. And resist the strongest temptation in the Measure phase: keep the pilot data as a lesson, not as evidence. Its job was to test the instrument, and mixing it into the baseline contaminates both.

Data does not become trustworthy when it is analyzed — it becomes trustworthy when it is collected.

From careful data to confident conclusions

A defensible plan is the foundation the whole DMAIC sequence stands on: Analyze can only interrogate what Measure captured honestly. Our free White Belt program introduces measurement thinking and the core vocabulary in about six hours, with a 30-question exam and a verifiable certificate. Green Belt teaches the full Measure-phase craft — definitions, sampling, baselines — inside a 35-hour simulated project, with a 100-question proctored exam, one free retake, and lifetime access. Build the habit now, and the next time someone challenges your numbers, the answer will already be written at the top of the plan.

Put it into practice

Ready to make it official?

Our Six Sigma belt programs — White through Black — are self-paced, 100% online, and end in a proctored exam and a credential you can verify and share.