8D, 5 Whys, Ishikawa, Is/Is-Not, and Shainin methods, followed by comparative selection, maturity, roadmap, and measurement architecture.
4.22 The Structured Problem-Solving Family: 8D, 5 Whys, Ishikawa, Is/Is-Not and the Shainin System
At a glance |
|
Layer |
L3 — Variation and quality engineering |
Evidence grade |
Mixed: B for the Shainin System and for the critical literature on root cause analysis; C for 8D; the critical evidence on 5 Whys is strong and negative |
Primary lever |
A graded set of causal-analysis techniques of differing power, suited to problems of differing structure |
Time to measurable effect |
Days to months depending on method and problem |
Regulatory friction (device plant) |
High for 8D, which in device plants is generally the corrective-action vehicle and therefore sits inside the quality system. |
4.22.1 The family and its ordering principle
These methods are usually presented as an undifferentiated toolbox, which is the main reason they are misapplied. They are better understood as differing along one axis — how much causal structure they assume — and the correct selection is the least powerful method whose assumptions the problem actually satisfies.
Method |
Causal structure assumed |
Appropriate problem class |
Characteristic misuse |
5 Whys |
A single linear causal chain terminating in one root cause |
Simple, well-bounded problems with an obvious causal path; a teaching device for causal reasoning |
Applied to multifactorial or systemic problems, where it produces a single plausible cause and stops |
Ishikawa (cause-and-effect) diagram |
Multiple parallel candidate causes across defined categories, unranked |
Structured generation of candidate causes before any narrowing |
Treated as analysis rather than as brainstorming; causes accepted without evidence |
Is / Is-Not (Kepner-Tregoe style specification) |
A distinguishing difference exists between where the problem occurs and where it does not |
Problems that are present in some units, times or locations and absent in others |
Skipped in favour of immediate cause listing, discarding the most informative contrast |
8D |
A multi-cause problem requiring containment, verified root cause and verified corrective action |
Customer-facing or regulatory non-conformances requiring formal, auditable resolution |
Used as a documentation template with D4 (root cause) asserted rather than demonstrated |
Shainin System |
A dominant cause exists and can be isolated by progressive elimination using contrast between extreme parts |
Chronic, high-volume, measurable variation problems with unknown cause |
Applied where no dominant cause exists, or without the disciplined component-swapping logic |
DMAIC (4.4) |
Multiple interacting causes requiring statistical estimation and designed experimentation |
Chronic variation problems with quantified financial exposure and available data |
Applied to ill-structured problems for which the goal cannot be specified in advance |
Table 4.1 The problem-solving family ordered by assumed causal structure. The correct choice is the least powerful method whose assumptions the problem satisfies — using a more powerful method than the problem requires wastes capacity; using a less powerful one produces a confident wrong answer.
4.22.2 5 Whys: what the evidence actually says
The 5 Whys technique — asking why iteratively until a root cause is reached — is the most widely taught and most heavily criticised method in this family, and the criticism is well founded and comes from the peer-reviewed literature rather than from rival consultancies.
Card (2017), writing in BMJ Quality & Safety, sets out the structural objections. Users are limited to one root cause per causal pathway, which misrepresents problems with multiple contributing causes. The convention of stopping at the fifth why and treating the most distal cause as the target has no logical basis — there is no reason to assume the most distal cause is the most effective or most efficient point of intervention, and in practice distal causes (“inadequate training culture”) are frequently less actionable than proximal ones. And the technique tends to oversimplify complex problems, limiting understanding of how processes actually fail. Card credits the technique’s value as a teaching tool for causal reasoning while questioning its use as a root cause analysis method.
The broader critique of root cause analysis is developed by Peerally, Carr, Waring and Dixon-Woods (2017), also in BMJ Quality & Safety, who identify a set of interlocking problems: the fixation on a single root cause when multiple factors interact; variable and often poor quality of investigations; the political and social context in which investigations are conducted, which shapes what is findable; and — the most operationally important — action plans that generate weak or poorly designed risk controls. That last point deserves emphasis because it is the one most directly under a plant’s control: an investigation of excellent quality that terminates in retraining and a procedure update has produced two of the weakest controls available, and the recurrence that follows is not a failure of the analysis but of the countermeasure design.
The practical position: retain 5 Whys as a first-line technique for simple, bounded problems and as a training device for causal reasoning; prohibit it as the sole method for recurring problems, for problems with safety or regulatory consequence, and for any problem that has previously been “solved”. In a device plant, a corrective action whose entire causal analysis is a 5 Whys chain terminating in operator error is a finding waiting to happen, and it should be treated as one internally before an inspector treats it as one externally.
4.22.3 8D: structure, value and the D4 problem
The eight disciplines method originated at Ford Motor Company in the 1980s as Team Oriented Problem Solving, and has become the de facto format for supplier and customer non-conformance resolution across automotive and, increasingly, medical device supply chains. Its disciplines run: D1 form the team; D2 describe the problem; D3 implement and verify interim containment; D4 identify and verify root causes and escape point; D5 choose and verify permanent corrective actions; D6 implement and validate; D7 prevent recurrence; D8 recognise the team.
Two design features distinguish 8D from the rest of the family and both are genuinely valuable. Containment as an explicit, verified discipline (D3) separates protecting the customer from understanding the problem — a separation that prevents the common and damaging pattern of leaving product exposed while an investigation proceeds. And the escape point concept in D4 requires identifying not only why the defect was created but why the control system failed to detect it, which is a second, independent causal question that most methods never ask. In a plant whose quality assurance rests on automated optical inspection, the escape-point question is frequently the more consequential of the two.
The characteristic failure is that D4 is asserted rather than demonstrated. The discipline says identify and verify; verification requires showing that the proposed cause can turn the effect on and off, which is an experimental claim. In practice D4 is commonly populated with a plausible narrative agreed by the team. Published research on 8D effectiveness exists but is dominated by single-organisation case studies; grade C is appropriate. The method’s real value is as an auditable structure that forces containment, escape-point analysis and effectiveness verification — not as a causal-analysis technique, which it does not itself supply. 8D is a container; it must be filled with an actual analytical method, whether that is Is/Is-Not, Shainin, or designed experimentation.
4.22.4 The Shainin System: progressive elimination of a dominant cause
The Shainin System, developed by Dorian Shainin, is the most technically distinctive member of this family and the least well known to engineers trained in the Six Sigma tradition. Steiner, MacKay and Ramberg (2008), in Quality Engineering, provide the reference overview and critical assessment, noting that much of the system had previously been neither well documented nor adequately discussed in peer-reviewed journals, and that Shainin himself called it statistical engineering — a term now used differently (4.4.6), which is a source of ongoing confusion.
The system’s organising assumption is the existence of a dominant cause — a single input whose variation accounts for the majority of output variation — and its strategy is progressive elimination: rather than estimating the effect of many candidate causes, successively partition the candidate set until the dominant cause is isolated. The characteristic tools follow from this logic. Best-of-the-best / worst-of-the-worst selects extreme parts to maximise contrast. Component search progressively swaps components between a good and a bad assembly to localise the cause. Multi-vari studies partition variation by family (within-part, part-to-part, time-to-time) before any cause is hypothesised. Paired comparison and isoplot provide low-cost discrimination.
Where a dominant cause exists, this approach is dramatically more efficient than factorial experimentation, because elimination requires far fewer runs than estimation. Where one does not — where variation is genuinely distributed across many small contributors — the assumption fails and the method will isolate something that is not dominant. Steiner and colleagues make exactly this point in their critical assessment, and the honest statement is that the Shainin System and designed experimentation are complements selected on problem structure, not competitors.
For an ophthalmic plant the fit is unusually good: high volume, measurable continuous outputs, multi-cavity tooling and multi-lane processes generate precisely the family-structured variation that multi-vari analysis is designed to decompose, and the component-search logic maps naturally onto tooling and fixture investigations. Grade B, on the strength of the peer-reviewed methodological assessment and a long industrial record.
4.22.5 Countermeasure strength: the step most methods omit
Every method in this family terminates in a countermeasure, and the Peerally critique identifies countermeasure design as the weakest link in practice. A countermeasure hierarchy should therefore be applied explicitly and recorded, ordered by reliance on human vigilance:
Rank |
Countermeasure class |
Ophthalmic example |
Reliance on human performance |
1 (strongest) |
Eliminate by design |
Redesign the filter housing so particulate ingress at filter change is physically impossible |
None |
2 |
Physical error-proofing (poka-yoke) |
Asymmetric tray keying so a tray cannot be loaded in the wrong orientation |
None |
3 |
Automated detection with automatic stop |
Ionic-strength interlock that halts bath transfer outside the validated band |
Minimal |
4 |
Automated detection with alarm |
Ionic-strength alarm requiring operator response |
Moderate |
5 |
Visual control or checklist |
Standard work sheet with a verification step at start-up |
High |
6 (weakest) |
Training and procedure revision |
Retrain operators on the existing method; revise the work instruction |
Total |
Table 4.2 Countermeasure strength hierarchy. Ranks 5 and 6 are the most frequently selected and the least effective. A corrective action portfolio dominated by training and procedure revision predicts recurrence, and the prediction is testable against the plant’s own recurrence data.
This table is worth using as a live audit instrument. Classify the last fifty corrective actions by rank. If the distribution is concentrated in ranks 5 and 6 — which it usually is — the plant has identified its dominant improvement opportunity in problem solving, and it is a countermeasure-design problem rather than an analysis problem.
4.22.6 Worked example — ophthalmic
A customer complaint reports lens edge chips in a specific lot. The response is structured as an 8D because it is customer-facing and regulatorily reportable, with the analytical content selected on problem structure rather than habit.
D3 containment: affected lots quarantined, 100 percent inspection instituted on stock, escape assessment completed within 48 hours. Containment is verified — a sample of quarantined stock is re-inspected to confirm the containment inspection detects the defect — because unverified containment is the most dangerous item on an 8D.
D4 analysis: the team explicitly declines a 5 Whys chain, on the grounds that edge chipping is a chronic, measurable, multi-lane variation problem rather than a bounded event. Instead an Is/Is-Not specification is built first: the defect is present on lanes 3 and 7, is not present on lanes 1, 2, 4, 5, 6, 8; is concentrated in the second half of production runs, is not present at run start. That contrast is the highest-information observation available and it is obtained before any cause is proposed. A multi-vari study then partitions variation and confirms that the dominant family is cavity-to-cavity within lane rather than time-to-time. Component search — swapping the demould fixtures between an affected and an unaffected lane — reproduces the defect on the previously good lane, which is the verification that D4 requires and that D4 usually lacks: the proposed cause has been shown to turn the effect on and off.
Escape point: the automated optical inspection recipe’s edge-region sensitivity was set below the level required to detect chips of this size class. This is an independent finding of equal importance to the cause, and it would not have been surfaced by any method that asks only why the defect occurred.
D5–D7 countermeasures, classified against Table 4.2: fixture redesign eliminating the wear geometry (rank 1); inspection recipe sensitivity revised and requalified (rank 3); fixture wear added to the predictive maintenance model of 4.11 (rank 3). No rank-5 or rank-6 countermeasure is used, which is deliberate and is recorded as such. Effectiveness verification is defined statistically in advance — the criterion, the sample size and the review date are written before implementation, not chosen afterwards.
4.22.7 Limitations and failure modes
Method selected by habit rather than by problem structure. The dominant failure across this family. Table 4.1 exists to make the selection explicit.
Unverified root cause. A cause is verified when it can be shown to turn the effect on and off, not when the team agrees it is plausible.
Weak countermeasures. Concentration in ranks 5 and 6 of Table 4.2 predicts recurrence. This is measurable in any plant today.
Escape point ignored. Solving why the defect was made while leaving intact the reason it was not detected preserves half the failure.
Effectiveness check as documentation. Under ISO 13485 and the QMSR, corrective action effectiveness must be verified. A check with no pre-defined statistical criterion cannot fail, and a check that cannot fail is not a check.
5 Whys used as the sole method on serious problems. Structurally incapable of representing multifactorial causation (Card 2017).
4.22.8 Primary metrics
Recurrence rate of closed corrective actions at 12 months (the single most informative measure of problem-solving quality); distribution of countermeasures across the Table 4.2 hierarchy; proportion of investigations with an experimentally verified rather than asserted root cause; proportion with an identified escape point; proportion with a pre-defined effectiveness criterion; investigation cycle time and rework-loop frequency (from the process mining of 4.14).
5 Comparative Analysis
5.1 The comparison problem
Methods are commonly compared on a single axis — usually claimed benefit — which is the least informative axis available, because claimed benefit is a function of publication incentives rather than of method properties. A useful comparison must span at least four independent dimensions: what the method acts on, how strong the evidence is, what it costs to acquire, and under what conditions it works. Tables 5.1 and 5.2 present the twenty-two methods across fourteen dimensions organised on that basis. Section 5.5 then compares the structured problem-solving methods head to head, since those are the ones practitioners most often treat as interchangeable.
Method |
Layer |
Primary operational outcome |
Time to effect |
Evidence |
Hard prerequisites |
4.1 Hoshin Kanri / XPS |
L1 |
Improvement capacity aligned and protected; selectivity enforced |
2–4 quarters |
B |
Executive commitment; a real capacity-allocation decision |
4.2 Toyota Kata |
L1 |
Scientific problem-solving capability distributed to line level |
2–3 quarters |
C |
Experienced coaches; daily cadence; managerial willingness not to answer |
4.3 Lean production system |
L2 |
Lead time, work-in-process, first-pass yield, flexibility |
2–4 quarters |
A |
Stable availability; adequate measurement; the soft bundle deployed concurrently |
4.4 Six Sigma / DMAIC |
L3 |
Reduced variation in critical characteristics; reduced indirect cost |
3–9 months per project |
A |
Measurement system adequacy; trained practitioners; project governance |
4.4.6 Statistical engineering |
L3 |
Resolution of large, ill-structured, multi-study problems |
6–18 months |
B |
Senior statistical competence; tolerance for non-linear project structure |
4.5 DFSS / design space |
L3 |
Capability designed in; post-launch change-control load reduced |
1–3 years |
B |
Development-manufacturing integration; experimental capacity at development scale |
4.6 Theory of Constraints |
L2 |
Throughput, due-date performance, inventory |
Weeks to 1 quarter |
B |
Identifiable constraint; measurement system that tolerates subordination |
4.7 QRM / POLCA |
L2 |
Manufacturing critical-path time in high-variety environments |
2–4 quarters |
B/C |
High-mix low-volume context; accounting reform to permit planned slack |
4.8 TPM / OEE |
L4 |
Availability, performance rate, start-up quality |
2–4 quarters |
B |
Protected operator time; honest loss measurement |
4.9 Zero Defect Manufacturing |
L3 |
Defect rate; detection lag; quantity at risk per event |
3–6 quarters |
B |
Part-level traceability joining process data to inspection outcome |
4.10 Quality 4.0 / ML-SPC |
L3/L5 |
Multivariate observability; predictive quality; soft sensing |
4–8 quarters |
B/C |
Process in statistical control; time-synchronised data layer; model governance |
4.11 Predictive maintenance |
L4 |
Unplanned downtime; maintenance cost; life utilisation |
3–6 quarters |
B/C |
Observable degradation signal with usable lead time; failure history |
4.12 Digital twin |
L5 |
Physical experiments avoided; faster capacity decisions |
4–8 quarters |
B/C |
Tractable physics or rich data; a stated decision the twin will support |
4.13 RL scheduling |
L2 |
Schedule adherence under dynamic conditions |
4–8 quarters |
D |
Validated simulation; tuned rule baseline; hard-constraint architecture |
4.14 Process mining |
L5→L2 |
Administrative cycle time; conformance evidence |
1–3 quarters |
B |
Event logs with reliable timestamps and case identity |
4.15 LSS4.0 / DMAIC 4.0 |
All |
Compressed improvement-project cycle time |
4–8 quarters |
B |
Functioning classical LSS capability; measured phase-level bottleneck |
4.16 Human-centric / I5.0 |
L1/L5 |
Intervention quality; knowledge retention; sustained engagement |
Continuous |
B/D |
Willingness to design work, not only technology |
4.17 Standardised and leader standard work |
L0 |
Abnormality made visible; managerial attention protected; gains retained |
1–2 quarters |
B / C |
Standards built with operators; separation of standardised work from the controlled work instruction |
4.18 Skills matrix / cross-training |
L0/L1 |
Usable workforce flexibility; competence depth on critical operations |
2–4 quarters |
B |
A stated mechanism for the flexibility; observable competence criteria |
4.19 Gemba walks |
L0 |
Unfiltered process information; workarounds surfaced |
1–2 quarters |
C |
Scheduled into leader standard work; findings visibly closed |
4.20 Kaizen (events and daily) |
L0/L2 |
Concentrated change; high-volume small improvement |
Days to 6 quarters |
B |
Loss-based target selection; pre-classified change control; designed follow-through |
4.21 PDCA/PDSA and A3 |
L0/L3 |
Disciplined experimental cycle; coached, visible reasoning |
Immediate to 6 quarters |
B (fidelity poor in practice) |
Iteration and small-scale testing actually practised; a mentor per A3 author |
4.22 8D, Shainin and the PS family |
L3 |
Causal analysis matched to problem structure; verified root cause |
Days to months |
B / C |
Method selected by problem structure; root cause verified experimentally |
Table 5.1 Comparative summary, dimensions 1–6. Evidence grades follow the scheme of Table 2.2.
Method |
Capital cost |
Capability residue |
Regulatory friction |
Best-fit context |
Dominant failure mode |
4.1 Hoshin Kanri / XPS |
Low |
High |
Low |
Any plant with more improvement ideas than capacity |
Ritualisation into an annual artefact |
4.2 Toyota Kata |
Low |
Very high |
Low |
Organisations with a long horizon and coachable supervision |
Coach scarcity; vocabulary without practice |
4.3 Lean production system |
Low–Medium |
High |
Moderate |
Repetitive discrete manufacture at moderate to high volume |
Hard bundle only; buffers removed without variability reduction |
4.4 Six Sigma / DMAIC |
Low–Medium |
Medium |
Moderate–High |
Chronic variation problems with quantified financial exposure |
Measurement system neglect; solution written into the charter |
4.4.6 Statistical engineering |
Low |
High |
Moderate |
Large ill-structured problems that resist DMAIC |
Treated as a synonym for advanced statistics |
4.5 DFSS / design space |
Medium |
High |
High then negative |
New product platforms; processes facing frequent post-launch change |
Under-investment where development and manufacturing are separate |
4.6 Theory of Constraints |
Low |
Medium |
Low |
Capacity-constrained plants with a stable constraint |
Local-efficiency measures silently defeating subordination |
4.7 QRM / POLCA |
Medium |
Medium |
Low–Moderate |
High-variety, made-to-order, long-lead-time operations |
Rejected by cost accounting before it is tested |
4.8 TPM / OEE |
Low–Medium |
High |
Low–Moderate |
Equipment-intensive processes with unplanned downtime |
Effectiveness figure managed rather than losses eliminated |
4.9 Zero Defect Manufacturing |
High |
Medium |
High |
High-speed, high-volume processes with large quantity at risk |
Traceability prerequisite discovered after commitment |
4.10 Quality 4.0 / ML-SPC |
High |
Medium |
High |
Multivariate, correlated, sensor-rich processes |
Modelling an out-of-control process; optimistic offline evaluation |
4.11 Predictive maintenance |
Medium–High |
Low–Medium |
Moderate |
Critical assets with high-variance failure distributions |
Predictions that never reach the maintenance schedule |
4.12 Digital twin |
High |
Medium |
Moderate–High |
Expensive physical experimentation; frequent capacity decisions |
Level inflation; use outside the validated region |
4.13 RL scheduling |
High |
Low |
Low–Moderate |
Extreme scheduling complexity with dynamic arrival |
Comparison against an untuned baseline |
4.14 Process mining |
Low–Medium |
Medium |
Low (often positive) |
Administrative and quality-system processes with event logs |
Analysis that terminates in a presentation |
4.15 LSS4.0 / DMAIC 4.0 |
Medium–High |
Medium |
Moderate |
Mature LSS organisations with a measured phase bottleneck |
Integration attempted without a functioning foundation |
4.16 Human-centric / I5.0 |
Low |
High |
Low–Moderate |
Highly automated environments with supervisory operators |
Rhetoric with no change to how automation is specified |
4.17 Standardised and leader standard work |
Very low |
High |
Low (often favourable) |
Any plant where problems surface late |
Standards imposed; leader standard work audited for ticks not findings |
4.18 Skills matrix / cross-training |
Low |
High |
Low (often favourable) |
Variable demand, deep competence requirements, retirement exposure |
Matrix as decoration; universal cross-training as the default target |
4.19 Gemba walks |
Very low |
Medium |
Low |
Report-mediated organisations with a large floor-to-office distance |
Inspection tourism; walk experienced as audit |
4.20 Kaizen (events and daily) |
Low |
Medium–High |
Moderate–High |
Areas needing rapid change; mature plants needing volume of small change |
Event addiction; decay by default; change-control collision |
4.21 PDCA/PDSA and A3 |
Very low |
Very high |
Low |
Any organisation building distributed problem-solving capability |
Single-pass PDCA; A3 completed alone as a form |
4.22 8D, Shainin and the PS family |
Low |
Medium–High |
High (corrective action is quality-system work) |
Complaint-driven and chronic-variation problems |
Root cause asserted; countermeasures at ranks 5–6 of Table 4.2 |
Table 5.2 Comparative summary, dimensions 7–14. Capital cost is relative and excludes the shared data-layer investment, which several methods share.
5.2 Reading the evidence distribution
Two patterns in the grades deserve comment. First, the highest grades attach to the oldest methods. Lean practice bundles and Six Sigma adoption carry grade A because they have been studied for two decades with large samples and control groups. The digitally enabled methods carry B, C and D not because they are inferior but because the evidence has not yet been generated. Confusing novelty with efficacy — or, symmetrically, confusing an established evidence base with a superior method — are both errors, and both are common.
Second, the methods with the highest capability residue have the weakest evidence. Toyota Kata is the clearest case: it plausibly has the largest long-run effect of any method in this review and it carries grade C, because capability is slow, diffuse and hard to attribute. Organisations that allocate resource strictly by evidence grade will systematically underinvest in capability and overinvest in tooling. The correct response is to hold capability investment to a different standard of proof, made explicitly, rather than to pretend the evidence is stronger than it is.
5.3 A selection procedure
Method selection should be driven by a measured diagnosis, not by a maturity fashion. The following procedure derives the method set from the plant’s own data.
Establish where the loss is. Decompose the gap between current and entitlement performance into cost of poor quality, capacity loss, inventory carrying cost, lead-time-driven lost revenue, and improvement-throughput constraint. Express each in currency per year. This decomposition, not a maturity assessment, is the entry point.
Classify the dominant loss. Quality-dominant (cost of poor quality above roughly 4 percent of cost of goods sold) → L3 methods. Capacity-dominant (demand exceeds capacity, or capital investment is being contemplated) → L2, beginning with 4.6. Flow-dominant (lead time uncompetitive, inventory high, service poor) → L2, choosing 4.3 or 4.7 by mix. Availability-dominant (unplanned downtime above roughly 8 percent) → L4. Improvement-throughput-dominant (a change-control queue longer than the improvement pipeline) → 4.14 on the quality-system processes, then 4.5.
Test the prerequisites. Against Table 5.1, confirm each hard prerequisite. A failed prerequisite is a project in itself and must be sequenced first, not assumed away.
Check the layer coupling. If the time from abnormality to detection is measured in shifts, the binding constraint is at L0 and no method above it will perform. If availability is below roughly 85 percent, no L2 flow method will hold. If measurement system agreement is inadequate, no L3 method will produce trustworthy results. If improvement capacity is unprotected, no method will survive a bad month.
Select the minimum sufficient method. Prefer the lowest-cost method that addresses the diagnosed loss. Discrete-event simulation before digital twin; tuned dispatching rules before reinforcement learning; classical multivariate control before machine learning; inspection improvement before process improvement where measurement is implicated.
Compute the change-control load of the proposed portfolio and compare it against demonstrated change-control throughput. If the portfolio exceeds throughput, either reduce the portfolio or treat change-control throughput as the first project.
The most common selection error Selecting a method because it is advanced rather than because it addresses the measured dominant loss. In device manufacture the specific form this takes is investment in predictive quality analytics while the automated inspection system’s attribute agreement is unquantified — building a sophisticated inference layer on a measurement system of unknown accuracy. The diagnostic question is simple and it is almost never asked first: what is the false-discovery rate of the inspection system that generates every quality number in this plant? |
5.4 Sequencing and dependency
The methods are not independent, and several have hard predecessors. The dependency structure below is expressed as a partial order: a method should not be started until its predecessors are functioning.
Method |
Hard predecessors |
Rationale |
Lean flow control (4.3) |
Measurement adequacy; TPM availability recovery (4.8) |
Pull systems on unstable capacity oscillate; the pull system is then blamed for the maintenance deficit |
Six Sigma projects (4.4) |
Measurement system analysis; project governance (4.1) |
Without measurement adequacy the analysis measures the gauge; without governance the portfolio drifts to easy problems |
Zero Defect Manufacturing (4.9) |
Defect taxonomy; part-level traceability; inspection MSA |
Prediction requires labelled data joined at part level; models inherit inspection error |
ML-augmented SPC (4.10) |
Classical SPC; process in statistical control; data layer |
Machine learning on an out-of-control process models the special causes as normal |
Predictive maintenance (4.11) |
TPM loss decomposition; failure history; criticality ranking |
Conditioning is only worthwhile where failure variance is high and the asset is critical |
Digital twin (4.12) |
A stated decision; validated physical data for model validation |
A twin without a decision to support becomes an orphaned asset |
RL scheduling (4.13) |
Validated discrete-event simulation; tuned rule baseline |
Policy quality is bounded by simulation fidelity; gains must be measured against a tuned baseline |
LSS4.0 (4.15) |
Functioning classical DMAIC capability; phase-time measurement |
Technology amplifies an existing capability and cannot substitute for its absence |
Everything above L0 |
Standardised work; visual controls; tiered accountability; one working problem-solving method (4.17, 4.21) |
A problem that is not detected cannot be solved, and a gain that is not standardised is not retained |
Kaizen events (4.20) |
Loss decomposition; pre-event change-control classification; designed follow-through |
Events without follow-through decay by default; events generating unimplementable changes destroy credibility |
A3 practice (4.21) |
Mentors with capacity to iterate over drafts |
Without mentoring the A3 is a template; the coaching is the method |
Predictive and analytical methods (4.9–4.13) |
A daily management system that already detects and resolves problems |
These methods multiply an existing improvement rate rather than creating one |
Table 5.3 Hard dependencies. Violating these is the most reliable way to spend a large budget for no operational effect.
5.5 Comparing the Problem-Solving Methods
Six methods in this review are, in some sense, all “structured problem solving”: PDCA/PDSA, the A3 process, DMAIC, 8D, the Improvement Kata, and the Shainin System — with the kaizen event functioning as a delivery format rather than a method. Practitioners routinely treat these as interchangeable, or as a maturity ladder in which 5 Whys is for beginners and DMAIC for experts. Both readings are wrong and both are expensive. They are distinguished by four properties: the problem structure they assume, the elapsed time and resource they require, the capability they build in the person using them, and whether they produce an auditable record.
Property |
PDCA / PDSA |
A3 process |
DMAIC |
8D |
Improvement Kata |
Shainin System |
Primary purpose |
Disciplined experimental cycle |
Coached reasoning made visible on one page |
Statistical variation reduction in a governed project |
Auditable resolution of a customer or regulatory non-conformance |
Deliberate practice of scientific thinking |
Isolation of a dominant cause by elimination |
Assumed problem structure |
Any; cause need not be known in advance |
Any; suits problems needing shared understanding |
Multiple interacting causes, measurable, data available |
Multi-cause defect requiring containment and escape analysis |
Unknown path to a known target condition |
A dominant cause exists and is isolable |
Typical elapsed time |
Hours to weeks per cycle |
2–8 weeks over 3–5 drafts |
3–9 months |
4–12 weeks |
Continuous, 2–4 week target conditions |
1–6 weeks |
Resource intensity |
Very low |
Low |
High |
Moderate–high |
Low but sustained |
Low–moderate |
Capability built in the user |
Moderate (if the study step is real) |
High — the mentoring is the method |
Moderate — mostly in the specialist, not the line |
Low — often experienced as documentation |
Very high — capability is the output |
High in causal reasoning; narrow in scope |
Auditable record produced |
Not inherently |
Yes, one page |
Yes, heavyweight |
Yes, and the format regulators and customers expect |
Experiment record, not an audit artefact |
Study reports, not a governance artefact |
Requires a coach or specialist |
No |
Yes — a mentor |
Yes — a Black Belt or equivalent |
No, but needs an analytical method inside it |
Yes — a trained coach |
Yes — trained practitioner |
Dominant failure mode |
Single pass, full scale, no prediction |
Completed alone as a form |
Ill-structured problems forced into it; MSA skipped |
Root cause asserted, not verified |
Coach scarcity; vocabulary without practice |
Applied where no dominant cause exists |
Evidence position |
Fidelity of application demonstrably poor (Taylor et al. 2014) |
C — canonical practitioner-scholarly sources |
A for adoption effects (Swink and Jacobs 2012) |
C — single-organisation case studies |
C — action research |
B — peer-reviewed methodological assessment |
Table 5.4 Head-to-head comparison of the structured problem-solving methods.
5.5.1 PDCA and A3: the comparison most often requested, and most often mis-stated
The question “should we use PDCA or A3?” is malformed, and understanding why is more useful than any answer to it. They are not alternatives at the same level of abstraction. PDCA is the logic; the A3 is a container for the logic plus a coaching protocol. Every well-executed A3 is a PDCA cycle — background and current condition constitute the grasp of the situation, target and countermeasures constitute plan, implementation is do, verification is study, and standardisation is act. The A3 adds three things the bare cycle does not have.
A space constraint that forces selection. One page is not a stylistic preference; it is a mechanism. An author who must fit the analysis on a single sheet is compelled to distinguish what is load-bearing from what is merely collected, and that compulsion is where much of the thinking happens.
A fixed left-to-right logical order that makes the reasoning inspectable. A mentor can point at the precise joint where the argument fails: this countermeasure does not address that cause; that cause does not explain this current condition; this target does not close the stated gap. Bare PDCA offers no such surface.
A mentoring protocol. Shook’s central claim is that the A3 is a mechanism for a senior person to develop a junior one through iteration over drafts. Remove the mentor and the A3 degrades to a form — which is the most common way organisations adopt it.
The practical consequences of the distinction are specific. Use bare PDCA where the problem is small, the owner is competent, and the overhead of a written artefact exceeds its value — daily kaizen (4.20) runs on bare PDCA, and should. Use the A3 where the problem needs shared understanding across functions, where the reasoning must be transmissible, or where the primary objective is to develop the author. Use the Improvement Kata where the objective is to build the underlying scientific-thinking habit rather than to solve any particular problem; Kata and A3 are complementary and are frequently run together, with the Kata supplying the experimental cadence and the A3 supplying the documentary structure.
One caution about the comparison as it is usually taught. Many organisations adopt the A3 template and report improved communication, which is real but is the least important of its three mechanisms. If the A3 is not being iterated with a mentor, the organisation has bought a page layout. The diagnostic is simple and worth asking directly: how many drafts does a typical A3 go through before it is accepted? If the answer is one, the method is not being practised.
5.5.2 When each method is the right choice
Problem situation |
Recommended method |
Why |
A small, local, obvious problem within one person’s control |
Bare PDCA, one cycle, immediately |
Written artefact overhead exceeds its value; speed dominates |
A recurring problem whose causes are disputed across functions |
A3 with a mentor |
Shared visible reasoning is the actual obstacle, not the analysis |
A chronic measurable variation problem with quantified financial exposure and data available |
DMAIC, with statistical engineering framing if the problem is ill-structured |
Requires measurement system assessment, designed experimentation and governance |
A chronic variation problem where a dominant cause is plausible and physical parts can be contrasted |
Shainin System (multi-vari, component search) |
Elimination requires far fewer runs than estimation; suits multi-cavity and multi-lane processes |
A customer complaint or regulatory non-conformance |
8D, with Is/Is-Not and an experimental verification inside D4 |
Containment and escape-point analysis are required; the format is what customers and regulators expect |
A problem where the path to the goal is genuinely unknown |
Improvement Kata |
Designed for operating at the threshold of knowledge; other methods assume the path is findable by analysis |
An area needing rapid, concentrated change with cross-functional agreement |
Kaizen event, with pre-classified change control |
A delivery format, not a method; still needs a method inside it |
A simple bounded problem, or a teaching occasion in causal reasoning |
5 Whys |
Adequate for its assumptions; strong evidence against using it beyond them (Card 2017) |
Table 5.5 Method selection by problem situation. The governing rule is to use the least powerful method whose assumptions the problem satisfies.
5.5.3 What the comparison implies for a device plant
Three implications follow, and they are not the ones most improvement programmes act on.
First, the plant’s corrective action system is almost certainly its highest-volume problem-solving process, and it is usually the least methodologically disciplined. A plant may run twelve DMAIC projects a year and four hundred corrective actions. If the four hundred are populated with 5 Whys chains terminating in operator error and closed with rank-5 or rank-6 countermeasures (Table 4.2), the aggregate effect of the twelve well-run projects is small by comparison. Improving the method inside the corrective action system has more leverage than adding projects, and it is rarely where improvement resource is directed.
Second, the effectiveness check is the enforcement mechanism, and it is usually unenforced. Both ISO 13485 and the QMSR require verification that corrective action was effective. A check with no pre-defined statistical criterion, no defined observation period and no defined sample size cannot fail. Defining these three things in advance — at the point the corrective action is written, not when it is closed — converts the effectiveness check from documentation into a real test, and the recurrence rate will fall for that reason alone.
Third, capability-building methods and delivery methods should be balanced deliberately. DMAIC and 8D deliver results and build modest capability, concentrated in specialists. A3 and Kata build capability distributed across the line and deliver results more slowly. A portfolio consisting only of the former produces an organisation that can solve problems only when a specialist is assigned — which caps the number of problems that can be solved per year at the number of specialists, and is precisely the ceiling most plants are operating against without recognising it.
5.6 The Daily Management System as the Substrate
Methods 4.17 to 4.22 do not sit alongside the methods of Sections 4.1 to 4.16; they sit beneath them. The daily management system — standardised work, leader standard work, visual controls, tiered accountability, the skills matrix, gemba, and a working problem-solving method — is the substrate that determines whether anything else functions.
The dependency is concrete rather than rhetorical. Detection: without standardised work there is no baseline, so abnormality is invisible and the improvement system has no input. Cadence: without tiered accountability, a detected problem waits for a meeting, and problem-solving cycle time is bounded below by meeting frequency. Capacity: without leader standard work, managerial attention is consumed by escalation and no observation occurs. Competence: without a real skills matrix, an improvement cannot be reliably deployed because the ability to execute the new standard is unknown. Retention: without standardisation following countermeasure, every gain decays to the mean.
This produces an uncomfortable but useful corollary. A plant with an excellent daily management system and only 5 Whys will out-improve a plant with a Six Sigma programme and no daily management, because the first plant detects and resolves a large volume of small problems quickly while the second detects few and resolves them slowly and expensively. The advanced analytical methods of Sections 4.9 to 4.15 multiply an existing improvement rate; they do not create one. If the multiplicand is near zero, the multiplier is irrelevant — which is the single most common reason expensive operational excellence programmes produce disappointing results.
Revision to the sequencing rule of Section 3.2 Layer L0 — the daily management system — precedes everything. Establish standardised work, visual controls, tiered accountability, leader standard work and one working problem-solving method before committing to any L2–L4 programme and long before any L5 investment. The diagnostic is a single question asked on the floor: when something goes wrong on this line, how long is it before someone whose job it is to fix it knows about it? If the answer is measured in shifts rather than minutes, the plant’s binding constraint is at L0, whatever its improvement roadmap says. |
6 An Integrated Maturity Model
Maturity models are frequently criticised, with justification, for encouraging organisations to pursue a level rather than a result. The model below is offered as a diagnostic instrument for identifying which capability is currently binding, not as a target. Its levels are defined by observable evidence rather than by self-assessment, and each level specifies what must be true, not what must be documented.
Level |
Descriptor |
Observable evidence |
Binding constraint at this level |
1 |
Reactive |
Improvement is event-driven and follows complaints, audit findings or crises. Loss data exist only in financial aggregate. Overall equipment effectiveness is not measured or is measured inconsistently. No standardised work, so abnormality is not detectable; problems surface in shifts rather than minutes. Deviation investigations recur on the same modes. |
Absence of detection and of loss visibility. Nothing can be prioritised because nothing is decomposed, and nothing is noticed in time to act. |
2 |
Measured |
Loss decomposition exists by cost of poor quality category, by six-loss category and by process step. Measurement systems have been formally assessed. Standardised work and visual controls are in place with a tiered accountability cadence, so abnormality is raised within the shift. A baseline capability study exists for each critical characteristic. |
Absence of protected improvement capacity. The plant knows what is wrong and cannot get to it. |
3 |
Systematic |
Improvement capacity is committed and defended through strategy deployment. Projects are selected from the risk file and the loss decomposition. DMAIC and lean practice are standard. Problem-solving method is chosen by problem structure, root causes are verified experimentally, and countermeasures are classified by strength. Control mechanisms are verified at 180 days. |
Change-control throughput. The improvement pipeline exceeds the plant’s ability to validate and implement. |
4 |
Capable |
Line-level personnel run structured experiments as routine. Design spaces exist for critical processes, reducing post-launch change events. Administrative and quality-system processes are mined and improved. Improvement rate is stable and predictable. |
Observability. Further gain requires information the current instrumentation cannot supply. |
5 |
Anticipatory |
Process state is observed continuously and multivariately; defect modes are addressed by prediction and prevention rather than by detection; maintenance is condition-driven within validated envelopes; experimentation is substantially in-silico. |
Model governance and organisational learning rate. The technical limits move outward faster than the organisation can absorb them. |
Table 6.1 Five-level operational excellence maturity model with explicit binding constraints.
The final column is the substance of the model. Each level is characterised not by what the organisation has achieved but by what will stop it next, and the correct investment at any level is the one that relieves that constraint. A level-2 plant buying predictive analytics is relieving a level-4 constraint while its level-2 constraint — no protected improvement capacity — remains untouched. This is the most common and most expensive misallocation in the field, and the model exists chiefly to make it visible.
A note on levels 4 and 5 in regulated manufacture. The transition from level 3 to level 4 is where the change-control constraint becomes decisive, and it is why design-space methods (4.5) and process mining of quality-system processes (4.14) are disproportionately valuable at that transition. Plants that attempt to reach level 4 by adding improvement resource without addressing change-control throughput simply lengthen the queue.
7 Implementation Roadmap for a Regulated Ophthalmic Plant
The roadmap below assumes a plant at maturity level 2 — losses measured, capacity unprotected — which is the modal starting position. It is expressed in four waves over thirty-six months. Durations are indicative; the sequencing is not, and it follows the dependency structure of Table 5.3.
7.1 Wave 1 (months 0–6): diagnosis, capacity and measurement
Complete the loss decomposition in currency terms: cost of poor quality by category and by defect mode; six-loss decomposition on all constraint and near-constraint equipment; manufacturing critical-path time by product family with process, queue, move and decision components separated; change-control throughput and queue length.
Conduct attribute agreement analysis on every automated inspection station against a truth panel, and gauge studies on every critical continuous measurement. Publish false-discovery and escape rates. Expect this to change the improvement priority list materially.
Establish strategy deployment (4.1): three to five breakthrough objectives, means-ends decomposition, catchball, and — the operative step — a named, protected improvement capacity with an exception procedure at plant-manager level.
Identify the constraint (4.6) and complete exploitation analysis before any capital request proceeds.
Begin process mining (4.14) on deviation management and change control. These are typically the fastest-returning projects in the entire programme and they raise the ceiling for everything that follows.
Establish the daily management substrate (4.17): standardised work built with operators on the constraint and near-constraint areas, visual controls designed for seconds-to-detect, a tiered accountability cadence, and leader standard work with an explicit gradient by level. Measure time from abnormality to detection as a programme metric from week one.
Rebuild the skills matrix against operational criteria (4.18): criticality-weighted depth targets, observable competence levels, and single-point-of-failure operations entered on the risk register rather than the training plan.
Audit the last fifty corrective actions against the countermeasure strength hierarchy (Table 4.2) and against whether root cause was verified or asserted. This costs a week and usually reframes the entire quality improvement agenda.
Wave 1 gate: loss decomposition complete and independently reviewed; measurement adequacy quantified for all critical measurements; improvement capacity committed in writing; constraint identified with exploitation options costed; tiered accountability running with time-to-detection measured and falling; corrective action countermeasure profile baselined.
7.2 Wave 2 (months 4–15): stabilise and flow
Total Productive Maintenance (4.8) on constraint and near-constraint equipment: initial cleaning and inspection, contamination-source elimination, operator standards with granted time, and planned maintenance intervals rebuilt from failure data.
Constraint exploitation (4.6): changeover reduction, load-density improvement, relocation of inspection gates so that constraint capacity is not spent on product that will be scrapped, and buffer management with coded penetration causes.
Lean flow (4.3) once availability is stable: setup reduction, work-in-process caps sized from measured variability, levelling, and — concurrently, not later — the soft bundle of structured small-group problem solving with protected time and a skills matrix.
Begin Toyota Kata (4.2) with two to four coaches on one value stream. Start now, because the coach-development lead time is long and it gates everything at wave 4.
Open the first Six Sigma projects (4.4) against the two largest quantified defect modes, selected from the loss decomposition and cross-referenced to the ISO 14971 risk file.
Install a working problem-solving discipline (4.21, 4.22): A3 practice with named mentors and a minimum draft count; PDCA enforced with written predictions, small-scale testing and iteration; method selection by problem structure using Table 4.1; pre-defined statistical effectiveness criteria on every corrective action.
Run two to four kaizen events (4.20) on loss-decomposition targets with change-control classification completed before the event, and begin building the daily kaizen route so that events decline in frequency over time.
Establish gemba walks (4.19) with separated purposes, fixed routes and visible closure of findings, with explicit attention to workarounds.
Wave 2 gate: availability on constraint equipment above 90 percent; work-in-process capped and holding; two Six Sigma projects closed with verified 180-day sustainment; coaching cycles running at planned frequency; median A3 draft count above two; corrective actions at ranks 1–3 of Table 4.2 rising as a share of the total; 12-month recurrence rate baselined.
7.3 Wave 3 (months 12–26): design in, and build the data layer
Design-space characterisation (4.5) on the two highest-change-frequency validated processes. Target a measurable reduction in post-launch change-control events, and report that reduction as the primary benefit metric rather than reporting capability alone.
Build the shared data layer: time-synchronised process traces, part- or batch-level identity, and a reliable join to inspection outcomes. Build it once, for the plant, not per project. This is the prerequisite for 4.9, 4.10 and 4.12 and it is the item most often under-scoped.
Construct the defect taxonomy and assign Zero Defect Manufacturing strategies (4.9). Move at least two modes from detection to prevention — the strategy migration, not the defect rate alone, is the maturity signal.
Deploy classical multivariate statistical process control on the sensor-rich process before any machine learning. Establish what conventional multivariate methods achieve, so that the incremental value of learned models can be measured rather than assumed.
Build a discrete-event model of the moulding-to-sterilisation chain (4.12, digital model level) and use it for buffer and capacity decisions. Defer any sensor-integrated twin until a specific decision requires one.
Wave 3 gate: design space validated on at least one critical process with a demonstrated reduction in change events; data layer operational with verified part-level joins; multivariate monitoring in production with quantified lead time over univariate detection.
7.4 Wave 4 (months 24–36): predict and prevent
Predictive quality models (4.10) on the two defect modes with the strongest process signature, deployed in advisory mode with prospective performance measurement against the offline estimate.
Predictive maintenance (4.11) on assets with a high coefficient of variation in failure interval, structured as condition-based scheduling inside a retained validated maximum interval.
Model governance framework: version control, performance monitoring, retraining triggers, change control on model updates, rollback path, and periodic review. Establish this before, not after, any model acquires disposition authority.
Extend Toyota Kata to further value streams as coach capacity allows. Coach capacity, not enthusiasm, is the governor.
Formalise the integration architecture (4.15) against the plant’s own measured phase-level bottleneck rather than against a reference framework.
Wave 4 gate: at least one predictive model with prospective performance within a defined tolerance of its offline estimate; model governance operating; documented reduction in unplanned downtime on target assets; improvement rate sustained without additional headcount.
What is deliberately absent from this roadmap Reinforcement-learning scheduling (4.13) does not appear. At evidence grade D, with scarce production deployment and unresolved verification difficulties, it does not belong in a committed thirty-six-month plan for a regulated plant. It belongs on a watch list, revisited annually, and should be entered only through a bounded evaluation against a properly tuned dispatching baseline. Sensor-integrated digital twins in the strict Kritzinger sense are similarly deferred: the digital model level answers most of the questions at a small fraction of the cost. |
8 Measurement Architecture
8.1 Design principles
Measurement systems in improvement programmes fail in characteristic ways: they proliferate, they drift toward what is easy to collect, they aggregate away the variation that carries the information, and they are reported as means when the operational consequence lives in the tail. Four principles counter this.
Report distributions, not means. Lead time, cycle time and deviation closure time are right-skewed. The median and the 90th percentile carry the operational content; the mean carries almost none. A plant that reports mean deviation closure time is reporting a number that describes no actual deviation.
Preserve the decomposition. Every composite metric — overall equipment effectiveness, cost of poor quality, manufacturing critical-path time — must be reported with its components. Composites without components become targets to be managed rather than measurements to be acted on.
Separate leading from lagging. Lagging metrics establish whether the programme worked. Leading metrics establish whether it will. A review consisting only of lagging metrics is an autopsy.
Instrument the improvement system itself. Project cycle time by phase, change-control throughput, coaching cycles delivered, experiments run — these measure the capability, which is the actual object of investment under the definition adopted in Section 2.1.
8.2 The metric architecture
Tier |
Purpose |
Representative metrics |
Cadence and audience |
Tier 1 — Outcome |
Establish whether the programme is producing value |
Cost of poor quality as percentage of cost of goods sold (with four-category split); dock-to-dock lead time (median and P90); overall equipment effectiveness on constraint assets; on-time in-full; unit cost |
Monthly; plant leadership and corporate |
Tier 2 — Process |
Establish whether the operational system is behaving as designed |
Process capability by critical characteristic; first-pass yield by stage; work-in-process turns; buffer penetration by zone; unplanned downtime; schedule adherence; every-part-every-interval |
Weekly; value stream and department level |
Tier 3 — Improvement system |
Establish whether the capability is growing |
Project cycle time by DMAIC phase; change-control throughput and queue; experiments per learner per week; coaching cycles delivered against planned; improvement hours delivered against committed; 180-day sustainment rate |
Monthly; improvement governance |
Tier 4 — Measurement integrity |
Establish whether the numbers can be trusted |
Inspection false-discovery and escape rates; gauge study currency by instrument; model performance drift against qualification baseline; data-layer join completeness |
Quarterly; quality engineering |
Table 8.1 Four-tier metric architecture. Tier 4 is the one most often absent and the one on which the credibility of tiers 1 to 3 entirely depends.
8.3 Linking to financial outcome
The evidence on financial linkage is instructive and should discipline the business case. Swink and Jacobs (2012) found that the return-on-assets improvement associated with Six Sigma adoption arose predominantly from indirect cost reduction, with direct cost and asset productivity effects not significant. A benefit model that projects savings principally from direct labour and material is therefore projecting the effect that the best available controlled evidence did not find, and should expect to under-deliver.
For a device plant the implication is specific and actionable: the largest realistically capturable financial benefit is likely to lie in the quality and engineering overhead — deviation investigation labour, batch record review, rework administration, complaint handling, validation effort per change — rather than in the conversion cost of the product. This is precisely the domain that process mining (4.14) addresses, and it is a reason to sequence that method early despite its unglamorous subject matter.
Two governance rules follow. First, require finance validation of realised benefit at 180 days for every closed project, and track the ratio of validated to claimed benefit as a programme metric; a ratio persistently below about 0.7 indicates a benefit-modelling problem, not an execution problem. Second, account for the change-control cost of each improvement explicitly in the business case. An improvement worth 120,000 currency units annually that consumes 400 hours of validation effort has a materially different profile from one that consumes 40, and in a plant where validation capacity is the binding constraint the second is worth more than the first even at half the nominal benefit.
Talk to us →