Run Two Changes at Once, Prove Neither
Someone calls out.
A position goes unfilled and the line keeps running
Someone calls out. The station sits empty for the shift. The work gets absorbed by whoever is nearest, and the line runs to the end of the day without stopping.
That is free evidence, and most operations throw it away. It is also thinner evidence than it looks. An unplanned absence tells you the operation survived one day, at one volume, on one product mix, with one supervisor covering and one crew that knew it was short. It is enough to justify a controlled test. It is not enough to remove the position.
What usually happens next is the problem. The observation becomes a proposal, the proposal gets scheduled, and the test starts running in a window where three other things are already moving.
Every change moves the baseline for the next one
Walk the sources of change on a single line over a typical quarter. Operations has already taken headcount out of an adjacent area on its own judgment. A piece of ergonomic or handling equipment is on order and will land sometime inside the next month. A new line is starting up and will pull trained people off this one when it does. Volume moves. The mix moves with it. And somewhere in that window, a test is scheduled to prove whether a position is needed.
Each of those is a legitimate change. Together they are a measurement problem, because the baseline is not a fixed thing the test is compared against. The baseline is whatever the operation was doing at the moment the step started, and every one of those changes rewrites it.
The failure is not that the line performs badly. Usually it performs fine. The failure is that the number at the end of the window belongs to no one. Output per paid hour improved, and the equipment install, the crew change, a softer mix, and the removed position are all sitting in the same interval with equal claim to it.
This is the part that costs money later. An improvement nobody can attribute gets counted by everyone who touched the window, so the same benefit shows up in an operations plan and a project case at the same time. A decline nobody can attribute gets blamed on whichever change is easiest to reverse, which is almost always the newest one, which is almost always the test. Good changes get pulled back and mediocre ones survive, and the record that would have settled it was never written.
My rule is that a change you cannot attribute is a change you cannot keep. I would rather run four steps over four weeks and be able to defend each one than run all four at once, get a better looking quarter, and have nothing to say when someone asks which step to repeat at the next plant.
Run one change per interval
Start by writing the change log before the first step, not after the last one. List every planned change to the operation, who owns it, and when it lands. Include the ones you do not control: equipment deliveries, a startup elsewhere that draws from this crew, a scheduled sanitation change, a seasonal mix shift. That list is the calendar the test has to fit inside, and it is usually the first time anyone has seen all of it in one place.
Then give each step its own interval, with one metric and one denominator named in advance. Output per paid hour and labor hours per unit of good product are not the same claim, and picking the denominator after you see the result is how a case stops being evidence. An interval is long enough when it has covered the conditions that could invalidate the change: the heavy days, the difficult product family, the short crew, the changeover. Calendar length is a proxy for that, not a substitute.
Protect the interval. When a delivery lands mid-step or a crew gets pulled, stop the step and restart it rather than averaging across the disturbance. The averaged number is the expensive one, because it looks like data and is not.
Separate what is already decided from what is being tested. A reduction operations has already made and a piece of equipment already purchased are captured changes. They belong in the baseline and in the change log, not in the test queue and not in the savings column a second time. Re-testing a decision that has already been made wastes an interval, and counting it again in a case is how a proposal loses credibility with the people who made the original call.
Step in the direction that preserves the ability to reverse. Move one increment, hold, read the metric, then decide whether the next increment is supported. A sequence built that way gives the supervisor a stopping point at every step, which is what makes the crew willing to run it honestly.
What a controlled sequence looks like on the floor
The supervisor can name the current step and what it is testing. The metric for the step was written down before it started, with its denominator. There is a single change log for the operation, and equipment dates and crew moves are on it next to the test steps. Each interval contains one change and has covered the hard days, not just the quiet ones. When a step is stopped early, the reason is recorded and the interval restarts rather than being averaged. At the end, each step has a result attached to it by name, and the changes that were already decided are marked as decided rather than counted as findings.
Before the next operating change goes on the schedule, ask what else is landing on that line inside the same window, and which of the two you are willing to move.