Advanced Methods · Part 5 of 6 · ~40 min read

Structured Problem Solving and Method Selection

8D, 5 Whys, Ishikawa, Is/Is-Not, and Shainin methods, followed by comparative selection, maturity, roadmap, and measurement architecture.

A technical monograph series · Heron Operational Excellence

Numbered citations refer to the complete references and supporting material in Part 6.

In this article · approximately 40 min read
Share on LinkedIn

8D, 5 Whys, Ishikawa, Is/Is-Not, and Shainin methods, followed by comparative selection, maturity, roadmap, and measurement architecture.

4.22 The Structured Problem-Solving Family: 8D, 5 Whys, Ishikawa, Is/Is-Not and the Shainin System

At a glance


Layer

L3 — Variation and quality engineering

Evidence grade

Mixed: B for the Shainin System and for the critical literature on root cause analysis; C for 8D; the critical evidence on 5 Whys is strong and negative

Primary lever

A graded set of causal-analysis techniques of differing power, suited to problems of differing structure

Time to measurable effect

Days to months depending on method and problem

Regulatory friction (device plant)

High for 8D, which in device plants is generally the corrective-action vehicle and therefore sits inside the quality system.



4.22.1 The family and its ordering principle

These methods are usually presented as an undifferentiated toolbox, which is the main reason they are misapplied. They are better understood as differing along one axis — how much causal structure they assume — and the correct selection is the least powerful method whose assumptions the problem actually satisfies.

Method

Causal structure assumed

Appropriate problem class

Characteristic misuse

5 Whys

A single linear causal chain terminating in one root cause

Simple, well-bounded problems with an obvious causal path; a teaching device for causal reasoning

Applied to multifactorial or systemic problems, where it produces a single plausible cause and stops

Ishikawa (cause-and-effect) diagram

Multiple parallel candidate causes across defined categories, unranked

Structured generation of candidate causes before any narrowing

Treated as analysis rather than as brainstorming; causes accepted without evidence

Is / Is-Not (Kepner-Tregoe style specification)

A distinguishing difference exists between where the problem occurs and where it does not

Problems that are present in some units, times or locations and absent in others

Skipped in favour of immediate cause listing, discarding the most informative contrast

8D

A multi-cause problem requiring containment, verified root cause and verified corrective action

Customer-facing or regulatory non-conformances requiring formal, auditable resolution

Used as a documentation template with D4 (root cause) asserted rather than demonstrated

Shainin System

A dominant cause exists and can be isolated by progressive elimination using contrast between extreme parts

Chronic, high-volume, measurable variation problems with unknown cause

Applied where no dominant cause exists, or without the disciplined component-swapping logic

DMAIC (4.4)

Multiple interacting causes requiring statistical estimation and designed experimentation

Chronic variation problems with quantified financial exposure and available data

Applied to ill-structured problems for which the goal cannot be specified in advance

Table 4.1 The problem-solving family ordered by assumed causal structure. The correct choice is the least powerful method whose assumptions the problem satisfies — using a more powerful method than the problem requires wastes capacity; using a less powerful one produces a confident wrong answer.

4.22.2 5 Whys: what the evidence actually says

The 5 Whys technique — asking why iteratively until a root cause is reached — is the most widely taught and most heavily criticised method in this family, and the criticism is well founded and comes from the peer-reviewed literature rather than from rival consultancies.

Card (2017), writing in BMJ Quality & Safety, sets out the structural objections. Users are limited to one root cause per causal pathway, which misrepresents problems with multiple contributing causes. The convention of stopping at the fifth why and treating the most distal cause as the target has no logical basis — there is no reason to assume the most distal cause is the most effective or most efficient point of intervention, and in practice distal causes (“inadequate training culture”) are frequently less actionable than proximal ones. And the technique tends to oversimplify complex problems, limiting understanding of how processes actually fail. Card credits the technique’s value as a teaching tool for causal reasoning while questioning its use as a root cause analysis method.

The broader critique of root cause analysis is developed by Peerally, Carr, Waring and Dixon-Woods (2017), also in BMJ Quality & Safety, who identify a set of interlocking problems: the fixation on a single root cause when multiple factors interact; variable and often poor quality of investigations; the political and social context in which investigations are conducted, which shapes what is findable; and — the most operationally important — action plans that generate weak or poorly designed risk controls. That last point deserves emphasis because it is the one most directly under a plant’s control: an investigation of excellent quality that terminates in retraining and a procedure update has produced two of the weakest controls available, and the recurrence that follows is not a failure of the analysis but of the countermeasure design.

The practical position: retain 5 Whys as a first-line technique for simple, bounded problems and as a training device for causal reasoning; prohibit it as the sole method for recurring problems, for problems with safety or regulatory consequence, and for any problem that has previously been “solved”. In a device plant, a corrective action whose entire causal analysis is a 5 Whys chain terminating in operator error is a finding waiting to happen, and it should be treated as one internally before an inspector treats it as one externally.

4.22.3 8D: structure, value and the D4 problem

The eight disciplines method originated at Ford Motor Company in the 1980s as Team Oriented Problem Solving, and has become the de facto format for supplier and customer non-conformance resolution across automotive and, increasingly, medical device supply chains. Its disciplines run: D1 form the team; D2 describe the problem; D3 implement and verify interim containment; D4 identify and verify root causes and escape point; D5 choose and verify permanent corrective actions; D6 implement and validate; D7 prevent recurrence; D8 recognise the team.

Two design features distinguish 8D from the rest of the family and both are genuinely valuable. Containment as an explicit, verified discipline (D3) separates protecting the customer from understanding the problem — a separation that prevents the common and damaging pattern of leaving product exposed while an investigation proceeds. And the escape point concept in D4 requires identifying not only why the defect was created but why the control system failed to detect it, which is a second, independent causal question that most methods never ask. In a plant whose quality assurance rests on automated optical inspection, the escape-point question is frequently the more consequential of the two.

The characteristic failure is that D4 is asserted rather than demonstrated. The discipline says identify and verify; verification requires showing that the proposed cause can turn the effect on and off, which is an experimental claim. In practice D4 is commonly populated with a plausible narrative agreed by the team. Published research on 8D effectiveness exists but is dominated by single-organisation case studies; grade C is appropriate. The method’s real value is as an auditable structure that forces containment, escape-point analysis and effectiveness verification — not as a causal-analysis technique, which it does not itself supply. 8D is a container; it must be filled with an actual analytical method, whether that is Is/Is-Not, Shainin, or designed experimentation.

4.22.4 The Shainin System: progressive elimination of a dominant cause

The Shainin System, developed by Dorian Shainin, is the most technically distinctive member of this family and the least well known to engineers trained in the Six Sigma tradition. Steiner, MacKay and Ramberg (2008), in Quality Engineering, provide the reference overview and critical assessment, noting that much of the system had previously been neither well documented nor adequately discussed in peer-reviewed journals, and that Shainin himself called it statistical engineering — a term now used differently (4.4.6), which is a source of ongoing confusion.

The system’s organising assumption is the existence of a dominant cause — a single input whose variation accounts for the majority of output variation — and its strategy is progressive elimination: rather than estimating the effect of many candidate causes, successively partition the candidate set until the dominant cause is isolated. The characteristic tools follow from this logic. Best-of-the-best / worst-of-the-worst selects extreme parts to maximise contrast. Component search progressively swaps components between a good and a bad assembly to localise the cause. Multi-vari studies partition variation by family (within-part, part-to-part, time-to-time) before any cause is hypothesised. Paired comparison and isoplot provide low-cost discrimination.

Where a dominant cause exists, this approach is dramatically more efficient than factorial experimentation, because elimination requires far fewer runs than estimation. Where one does not — where variation is genuinely distributed across many small contributors — the assumption fails and the method will isolate something that is not dominant. Steiner and colleagues make exactly this point in their critical assessment, and the honest statement is that the Shainin System and designed experimentation are complements selected on problem structure, not competitors.

For an ophthalmic plant the fit is unusually good: high volume, measurable continuous outputs, multi-cavity tooling and multi-lane processes generate precisely the family-structured variation that multi-vari analysis is designed to decompose, and the component-search logic maps naturally onto tooling and fixture investigations. Grade B, on the strength of the peer-reviewed methodological assessment and a long industrial record.

4.22.5 Countermeasure strength: the step most methods omit

Every method in this family terminates in a countermeasure, and the Peerally critique identifies countermeasure design as the weakest link in practice. A countermeasure hierarchy should therefore be applied explicitly and recorded, ordered by reliance on human vigilance:

Rank

Countermeasure class

Ophthalmic example

Reliance on human performance

1 (strongest)

Eliminate by design

Redesign the filter housing so particulate ingress at filter change is physically impossible

None

2

Physical error-proofing (poka-yoke)

Asymmetric tray keying so a tray cannot be loaded in the wrong orientation

None

3

Automated detection with automatic stop

Ionic-strength interlock that halts bath transfer outside the validated band

Minimal

4

Automated detection with alarm

Ionic-strength alarm requiring operator response

Moderate

5

Visual control or checklist

Standard work sheet with a verification step at start-up

High

6 (weakest)

Training and procedure revision

Retrain operators on the existing method; revise the work instruction

Total

Table 4.2 Countermeasure strength hierarchy. Ranks 5 and 6 are the most frequently selected and the least effective. A corrective action portfolio dominated by training and procedure revision predicts recurrence, and the prediction is testable against the plant’s own recurrence data.

This table is worth using as a live audit instrument. Classify the last fifty corrective actions by rank. If the distribution is concentrated in ranks 5 and 6 — which it usually is — the plant has identified its dominant improvement opportunity in problem solving, and it is a countermeasure-design problem rather than an analysis problem.

4.22.6 Worked example — ophthalmic

A customer complaint reports lens edge chips in a specific lot. The response is structured as an 8D because it is customer-facing and regulatorily reportable, with the analytical content selected on problem structure rather than habit.

D3 containment: affected lots quarantined, 100 percent inspection instituted on stock, escape assessment completed within 48 hours. Containment is verified — a sample of quarantined stock is re-inspected to confirm the containment inspection detects the defect — because unverified containment is the most dangerous item on an 8D.

D4 analysis: the team explicitly declines a 5 Whys chain, on the grounds that edge chipping is a chronic, measurable, multi-lane variation problem rather than a bounded event. Instead an Is/Is-Not specification is built first: the defect is present on lanes 3 and 7, is not present on lanes 1, 2, 4, 5, 6, 8; is concentrated in the second half of production runs, is not present at run start. That contrast is the highest-information observation available and it is obtained before any cause is proposed. A multi-vari study then partitions variation and confirms that the dominant family is cavity-to-cavity within lane rather than time-to-time. Component search — swapping the demould fixtures between an affected and an unaffected lane — reproduces the defect on the previously good lane, which is the verification that D4 requires and that D4 usually lacks: the proposed cause has been shown to turn the effect on and off.

Escape point: the automated optical inspection recipe’s edge-region sensitivity was set below the level required to detect chips of this size class. This is an independent finding of equal importance to the cause, and it would not have been surfaced by any method that asks only why the defect occurred.

D5–D7 countermeasures, classified against Table 4.2: fixture redesign eliminating the wear geometry (rank 1); inspection recipe sensitivity revised and requalified (rank 3); fixture wear added to the predictive maintenance model of 4.11 (rank 3). No rank-5 or rank-6 countermeasure is used, which is deliberate and is recorded as such. Effectiveness verification is defined statistically in advance — the criterion, the sample size and the review date are written before implementation, not chosen afterwards.

4.22.7 Limitations and failure modes

4.22.8 Primary metrics

Recurrence rate of closed corrective actions at 12 months (the single most informative measure of problem-solving quality); distribution of countermeasures across the Table 4.2 hierarchy; proportion of investigations with an experimentally verified rather than asserted root cause; proportion with an identified escape point; proportion with a pre-defined effectiveness criterion; investigation cycle time and rework-loop frequency (from the process mining of 4.14).



5 Comparative Analysis

5.1 The comparison problem

Methods are commonly compared on a single axis — usually claimed benefit — which is the least informative axis available, because claimed benefit is a function of publication incentives rather than of method properties. A useful comparison must span at least four independent dimensions: what the method acts on, how strong the evidence is, what it costs to acquire, and under what conditions it works. Tables 5.1 and 5.2 present the twenty-two methods across fourteen dimensions organised on that basis. Section 5.5 then compares the structured problem-solving methods head to head, since those are the ones practitioners most often treat as interchangeable.

Method

Layer

Primary operational outcome

Time to effect

Evidence

Hard prerequisites

4.1 Hoshin Kanri / XPS

L1

Improvement capacity aligned and protected; selectivity enforced

2–4 quarters

B

Executive commitment; a real capacity-allocation decision

4.2 Toyota Kata

L1

Scientific problem-solving capability distributed to line level

2–3 quarters

C

Experienced coaches; daily cadence; managerial willingness not to answer

4.3 Lean production system

L2

Lead time, work-in-process, first-pass yield, flexibility

2–4 quarters

A

Stable availability; adequate measurement; the soft bundle deployed concurrently

4.4 Six Sigma / DMAIC

L3

Reduced variation in critical characteristics; reduced indirect cost

3–9 months per project

A

Measurement system adequacy; trained practitioners; project governance

4.4.6 Statistical engineering

L3

Resolution of large, ill-structured, multi-study problems

6–18 months

B

Senior statistical competence; tolerance for non-linear project structure

4.5 DFSS / design space

L3

Capability designed in; post-launch change-control load reduced

1–3 years

B

Development-manufacturing integration; experimental capacity at development scale

4.6 Theory of Constraints

L2

Throughput, due-date performance, inventory

Weeks to 1 quarter

B

Identifiable constraint; measurement system that tolerates subordination

4.7 QRM / POLCA

L2

Manufacturing critical-path time in high-variety environments

2–4 quarters

B/C

High-mix low-volume context; accounting reform to permit planned slack

4.8 TPM / OEE

L4

Availability, performance rate, start-up quality

2–4 quarters

B

Protected operator time; honest loss measurement

4.9 Zero Defect Manufacturing

L3

Defect rate; detection lag; quantity at risk per event

3–6 quarters

B

Part-level traceability joining process data to inspection outcome

4.10 Quality 4.0 / ML-SPC

L3/L5

Multivariate observability; predictive quality; soft sensing

4–8 quarters

B/C

Process in statistical control; time-synchronised data layer; model governance

4.11 Predictive maintenance

L4

Unplanned downtime; maintenance cost; life utilisation

3–6 quarters

B/C

Observable degradation signal with usable lead time; failure history

4.12 Digital twin

L5

Physical experiments avoided; faster capacity decisions

4–8 quarters

B/C

Tractable physics or rich data; a stated decision the twin will support

4.13 RL scheduling

L2

Schedule adherence under dynamic conditions

4–8 quarters

D

Validated simulation; tuned rule baseline; hard-constraint architecture

4.14 Process mining

L5→L2

Administrative cycle time; conformance evidence

1–3 quarters

B

Event logs with reliable timestamps and case identity

4.15 LSS4.0 / DMAIC 4.0

All

Compressed improvement-project cycle time

4–8 quarters

B

Functioning classical LSS capability; measured phase-level bottleneck

4.16 Human-centric / I5.0

L1/L5

Intervention quality; knowledge retention; sustained engagement

Continuous

B/D

Willingness to design work, not only technology

4.17 Standardised and leader standard work

L0

Abnormality made visible; managerial attention protected; gains retained

1–2 quarters

B / C

Standards built with operators; separation of standardised work from the controlled work instruction

4.18 Skills matrix / cross-training

L0/L1

Usable workforce flexibility; competence depth on critical operations

2–4 quarters

B

A stated mechanism for the flexibility; observable competence criteria

4.19 Gemba walks

L0

Unfiltered process information; workarounds surfaced

1–2 quarters

C

Scheduled into leader standard work; findings visibly closed

4.20 Kaizen (events and daily)

L0/L2

Concentrated change; high-volume small improvement

Days to 6 quarters

B

Loss-based target selection; pre-classified change control; designed follow-through

4.21 PDCA/PDSA and A3

L0/L3

Disciplined experimental cycle; coached, visible reasoning

Immediate to 6 quarters

B (fidelity poor in practice)

Iteration and small-scale testing actually practised; a mentor per A3 author

4.22 8D, Shainin and the PS family

L3

Causal analysis matched to problem structure; verified root cause

Days to months

B / C

Method selected by problem structure; root cause verified experimentally

Table 5.1 Comparative summary, dimensions 1–6. Evidence grades follow the scheme of Table 2.2.

Method

Capital cost

Capability residue

Regulatory friction

Best-fit context

Dominant failure mode

4.1 Hoshin Kanri / XPS

Low

High

Low

Any plant with more improvement ideas than capacity

Ritualisation into an annual artefact

4.2 Toyota Kata

Low

Very high

Low

Organisations with a long horizon and coachable supervision

Coach scarcity; vocabulary without practice

4.3 Lean production system

Low–Medium

High

Moderate

Repetitive discrete manufacture at moderate to high volume

Hard bundle only; buffers removed without variability reduction

4.4 Six Sigma / DMAIC

Low–Medium

Medium

Moderate–High

Chronic variation problems with quantified financial exposure

Measurement system neglect; solution written into the charter

4.4.6 Statistical engineering

Low

High

Moderate

Large ill-structured problems that resist DMAIC

Treated as a synonym for advanced statistics

4.5 DFSS / design space

Medium

High

High then negative

New product platforms; processes facing frequent post-launch change

Under-investment where development and manufacturing are separate

4.6 Theory of Constraints

Low

Medium

Low

Capacity-constrained plants with a stable constraint

Local-efficiency measures silently defeating subordination

4.7 QRM / POLCA

Medium

Medium

Low–Moderate

High-variety, made-to-order, long-lead-time operations

Rejected by cost accounting before it is tested

4.8 TPM / OEE

Low–Medium

High

Low–Moderate

Equipment-intensive processes with unplanned downtime

Effectiveness figure managed rather than losses eliminated

4.9 Zero Defect Manufacturing

High

Medium

High

High-speed, high-volume processes with large quantity at risk

Traceability prerequisite discovered after commitment

4.10 Quality 4.0 / ML-SPC

High

Medium

High

Multivariate, correlated, sensor-rich processes

Modelling an out-of-control process; optimistic offline evaluation

4.11 Predictive maintenance

Medium–High

Low–Medium

Moderate

Critical assets with high-variance failure distributions

Predictions that never reach the maintenance schedule

4.12 Digital twin

High

Medium

Moderate–High

Expensive physical experimentation; frequent capacity decisions

Level inflation; use outside the validated region

4.13 RL scheduling

High

Low

Low–Moderate

Extreme scheduling complexity with dynamic arrival

Comparison against an untuned baseline

4.14 Process mining

Low–Medium

Medium

Low (often positive)

Administrative and quality-system processes with event logs

Analysis that terminates in a presentation

4.15 LSS4.0 / DMAIC 4.0

Medium–High

Medium

Moderate

Mature LSS organisations with a measured phase bottleneck

Integration attempted without a functioning foundation

4.16 Human-centric / I5.0

Low

High

Low–Moderate

Highly automated environments with supervisory operators

Rhetoric with no change to how automation is specified

4.17 Standardised and leader standard work

Very low

High

Low (often favourable)

Any plant where problems surface late

Standards imposed; leader standard work audited for ticks not findings

4.18 Skills matrix / cross-training

Low

High

Low (often favourable)

Variable demand, deep competence requirements, retirement exposure

Matrix as decoration; universal cross-training as the default target

4.19 Gemba walks

Very low

Medium

Low

Report-mediated organisations with a large floor-to-office distance

Inspection tourism; walk experienced as audit

4.20 Kaizen (events and daily)

Low

Medium–High

Moderate–High

Areas needing rapid change; mature plants needing volume of small change

Event addiction; decay by default; change-control collision

4.21 PDCA/PDSA and A3

Very low

Very high

Low

Any organisation building distributed problem-solving capability

Single-pass PDCA; A3 completed alone as a form

4.22 8D, Shainin and the PS family

Low

Medium–High

High (corrective action is quality-system work)

Complaint-driven and chronic-variation problems

Root cause asserted; countermeasures at ranks 5–6 of Table 4.2

Table 5.2 Comparative summary, dimensions 7–14. Capital cost is relative and excludes the shared data-layer investment, which several methods share.

5.2 Reading the evidence distribution

Two patterns in the grades deserve comment. First, the highest grades attach to the oldest methods. Lean practice bundles and Six Sigma adoption carry grade A because they have been studied for two decades with large samples and control groups. The digitally enabled methods carry B, C and D not because they are inferior but because the evidence has not yet been generated. Confusing novelty with efficacy — or, symmetrically, confusing an established evidence base with a superior method — are both errors, and both are common.

Second, the methods with the highest capability residue have the weakest evidence. Toyota Kata is the clearest case: it plausibly has the largest long-run effect of any method in this review and it carries grade C, because capability is slow, diffuse and hard to attribute. Organisations that allocate resource strictly by evidence grade will systematically underinvest in capability and overinvest in tooling. The correct response is to hold capability investment to a different standard of proof, made explicitly, rather than to pretend the evidence is stronger than it is.

5.3 A selection procedure

Method selection should be driven by a measured diagnosis, not by a maturity fashion. The following procedure derives the method set from the plant’s own data.

  1. Establish where the loss is. Decompose the gap between current and entitlement performance into cost of poor quality, capacity loss, inventory carrying cost, lead-time-driven lost revenue, and improvement-throughput constraint. Express each in currency per year. This decomposition, not a maturity assessment, is the entry point.

  2. Classify the dominant loss. Quality-dominant (cost of poor quality above roughly 4 percent of cost of goods sold) → L3 methods. Capacity-dominant (demand exceeds capacity, or capital investment is being contemplated) → L2, beginning with 4.6. Flow-dominant (lead time uncompetitive, inventory high, service poor) → L2, choosing 4.3 or 4.7 by mix. Availability-dominant (unplanned downtime above roughly 8 percent) → L4. Improvement-throughput-dominant (a change-control queue longer than the improvement pipeline) → 4.14 on the quality-system processes, then 4.5.

  3. Test the prerequisites. Against Table 5.1, confirm each hard prerequisite. A failed prerequisite is a project in itself and must be sequenced first, not assumed away.

  4. Check the layer coupling. If the time from abnormality to detection is measured in shifts, the binding constraint is at L0 and no method above it will perform. If availability is below roughly 85 percent, no L2 flow method will hold. If measurement system agreement is inadequate, no L3 method will produce trustworthy results. If improvement capacity is unprotected, no method will survive a bad month.

  5. Select the minimum sufficient method. Prefer the lowest-cost method that addresses the diagnosed loss. Discrete-event simulation before digital twin; tuned dispatching rules before reinforcement learning; classical multivariate control before machine learning; inspection improvement before process improvement where measurement is implicated.

  6. Compute the change-control load of the proposed portfolio and compare it against demonstrated change-control throughput. If the portfolio exceeds throughput, either reduce the portfolio or treat change-control throughput as the first project.

The most common selection error

Selecting a method because it is advanced rather than because it addresses the measured dominant loss. In device manufacture the specific form this takes is investment in predictive quality analytics while the automated inspection system’s attribute agreement is unquantified — building a sophisticated inference layer on a measurement system of unknown accuracy. The diagnostic question is simple and it is almost never asked first: what is the false-discovery rate of the inspection system that generates every quality number in this plant?



5.4 Sequencing and dependency

The methods are not independent, and several have hard predecessors. The dependency structure below is expressed as a partial order: a method should not be started until its predecessors are functioning.

Method

Hard predecessors

Rationale

Lean flow control (4.3)

Measurement adequacy; TPM availability recovery (4.8)

Pull systems on unstable capacity oscillate; the pull system is then blamed for the maintenance deficit

Six Sigma projects (4.4)

Measurement system analysis; project governance (4.1)

Without measurement adequacy the analysis measures the gauge; without governance the portfolio drifts to easy problems

Zero Defect Manufacturing (4.9)

Defect taxonomy; part-level traceability; inspection MSA

Prediction requires labelled data joined at part level; models inherit inspection error

ML-augmented SPC (4.10)

Classical SPC; process in statistical control; data layer

Machine learning on an out-of-control process models the special causes as normal

Predictive maintenance (4.11)

TPM loss decomposition; failure history; criticality ranking

Conditioning is only worthwhile where failure variance is high and the asset is critical

Digital twin (4.12)

A stated decision; validated physical data for model validation

A twin without a decision to support becomes an orphaned asset

RL scheduling (4.13)

Validated discrete-event simulation; tuned rule baseline

Policy quality is bounded by simulation fidelity; gains must be measured against a tuned baseline

LSS4.0 (4.15)

Functioning classical DMAIC capability; phase-time measurement

Technology amplifies an existing capability and cannot substitute for its absence

Everything above L0

Standardised work; visual controls; tiered accountability; one working problem-solving method (4.17, 4.21)

A problem that is not detected cannot be solved, and a gain that is not standardised is not retained

Kaizen events (4.20)

Loss decomposition; pre-event change-control classification; designed follow-through

Events without follow-through decay by default; events generating unimplementable changes destroy credibility

A3 practice (4.21)

Mentors with capacity to iterate over drafts

Without mentoring the A3 is a template; the coaching is the method

Predictive and analytical methods (4.9–4.13)

A daily management system that already detects and resolves problems

These methods multiply an existing improvement rate rather than creating one

Table 5.3 Hard dependencies. Violating these is the most reliable way to spend a large budget for no operational effect.

5.5 Comparing the Problem-Solving Methods

Six methods in this review are, in some sense, all “structured problem solving”: PDCA/PDSA, the A3 process, DMAIC, 8D, the Improvement Kata, and the Shainin System — with the kaizen event functioning as a delivery format rather than a method. Practitioners routinely treat these as interchangeable, or as a maturity ladder in which 5 Whys is for beginners and DMAIC for experts. Both readings are wrong and both are expensive. They are distinguished by four properties: the problem structure they assume, the elapsed time and resource they require, the capability they build in the person using them, and whether they produce an auditable record.

Property

PDCA / PDSA

A3 process

DMAIC

8D

Improvement Kata

Shainin System

Primary purpose

Disciplined experimental cycle

Coached reasoning made visible on one page

Statistical variation reduction in a governed project

Auditable resolution of a customer or regulatory non-conformance

Deliberate practice of scientific thinking

Isolation of a dominant cause by elimination

Assumed problem structure

Any; cause need not be known in advance

Any; suits problems needing shared understanding

Multiple interacting causes, measurable, data available

Multi-cause defect requiring containment and escape analysis

Unknown path to a known target condition

A dominant cause exists and is isolable

Typical elapsed time

Hours to weeks per cycle

2–8 weeks over 3–5 drafts

3–9 months

4–12 weeks

Continuous, 2–4 week target conditions

1–6 weeks

Resource intensity

Very low

Low

High

Moderate–high

Low but sustained

Low–moderate

Capability built in the user

Moderate (if the study step is real)

High — the mentoring is the method

Moderate — mostly in the specialist, not the line

Low — often experienced as documentation

Very high — capability is the output

High in causal reasoning; narrow in scope

Auditable record produced

Not inherently

Yes, one page

Yes, heavyweight

Yes, and the format regulators and customers expect

Experiment record, not an audit artefact

Study reports, not a governance artefact

Requires a coach or specialist

No

Yes — a mentor

Yes — a Black Belt or equivalent

No, but needs an analytical method inside it

Yes — a trained coach

Yes — trained practitioner

Dominant failure mode

Single pass, full scale, no prediction

Completed alone as a form

Ill-structured problems forced into it; MSA skipped

Root cause asserted, not verified

Coach scarcity; vocabulary without practice

Applied where no dominant cause exists

Evidence position

Fidelity of application demonstrably poor (Taylor et al. 2014)

C — canonical practitioner-scholarly sources

A for adoption effects (Swink and Jacobs 2012)

C — single-organisation case studies

C — action research

B — peer-reviewed methodological assessment

Table 5.4 Head-to-head comparison of the structured problem-solving methods.

5.5.1 PDCA and A3: the comparison most often requested, and most often mis-stated

The question “should we use PDCA or A3?” is malformed, and understanding why is more useful than any answer to it. They are not alternatives at the same level of abstraction. PDCA is the logic; the A3 is a container for the logic plus a coaching protocol. Every well-executed A3 is a PDCA cycle — background and current condition constitute the grasp of the situation, target and countermeasures constitute plan, implementation is do, verification is study, and standardisation is act. The A3 adds three things the bare cycle does not have.

  1. A space constraint that forces selection. One page is not a stylistic preference; it is a mechanism. An author who must fit the analysis on a single sheet is compelled to distinguish what is load-bearing from what is merely collected, and that compulsion is where much of the thinking happens.

  2. A fixed left-to-right logical order that makes the reasoning inspectable. A mentor can point at the precise joint where the argument fails: this countermeasure does not address that cause; that cause does not explain this current condition; this target does not close the stated gap. Bare PDCA offers no such surface.

  3. A mentoring protocol. Shook’s central claim is that the A3 is a mechanism for a senior person to develop a junior one through iteration over drafts. Remove the mentor and the A3 degrades to a form — which is the most common way organisations adopt it.

The practical consequences of the distinction are specific. Use bare PDCA where the problem is small, the owner is competent, and the overhead of a written artefact exceeds its value — daily kaizen (4.20) runs on bare PDCA, and should. Use the A3 where the problem needs shared understanding across functions, where the reasoning must be transmissible, or where the primary objective is to develop the author. Use the Improvement Kata where the objective is to build the underlying scientific-thinking habit rather than to solve any particular problem; Kata and A3 are complementary and are frequently run together, with the Kata supplying the experimental cadence and the A3 supplying the documentary structure.

One caution about the comparison as it is usually taught. Many organisations adopt the A3 template and report improved communication, which is real but is the least important of its three mechanisms. If the A3 is not being iterated with a mentor, the organisation has bought a page layout. The diagnostic is simple and worth asking directly: how many drafts does a typical A3 go through before it is accepted? If the answer is one, the method is not being practised.

5.5.2 When each method is the right choice

Problem situation

Recommended method

Why

A small, local, obvious problem within one person’s control

Bare PDCA, one cycle, immediately

Written artefact overhead exceeds its value; speed dominates

A recurring problem whose causes are disputed across functions

A3 with a mentor

Shared visible reasoning is the actual obstacle, not the analysis

A chronic measurable variation problem with quantified financial exposure and data available

DMAIC, with statistical engineering framing if the problem is ill-structured

Requires measurement system assessment, designed experimentation and governance

A chronic variation problem where a dominant cause is plausible and physical parts can be contrasted

Shainin System (multi-vari, component search)

Elimination requires far fewer runs than estimation; suits multi-cavity and multi-lane processes

A customer complaint or regulatory non-conformance

8D, with Is/Is-Not and an experimental verification inside D4

Containment and escape-point analysis are required; the format is what customers and regulators expect

A problem where the path to the goal is genuinely unknown

Improvement Kata

Designed for operating at the threshold of knowledge; other methods assume the path is findable by analysis

An area needing rapid, concentrated change with cross-functional agreement

Kaizen event, with pre-classified change control

A delivery format, not a method; still needs a method inside it

A simple bounded problem, or a teaching occasion in causal reasoning

5 Whys

Adequate for its assumptions; strong evidence against using it beyond them (Card 2017)

Table 5.5 Method selection by problem situation. The governing rule is to use the least powerful method whose assumptions the problem satisfies.

5.5.3 What the comparison implies for a device plant

Three implications follow, and they are not the ones most improvement programmes act on.

First, the plant’s corrective action system is almost certainly its highest-volume problem-solving process, and it is usually the least methodologically disciplined. A plant may run twelve DMAIC projects a year and four hundred corrective actions. If the four hundred are populated with 5 Whys chains terminating in operator error and closed with rank-5 or rank-6 countermeasures (Table 4.2), the aggregate effect of the twelve well-run projects is small by comparison. Improving the method inside the corrective action system has more leverage than adding projects, and it is rarely where improvement resource is directed.

Second, the effectiveness check is the enforcement mechanism, and it is usually unenforced. Both ISO 13485 and the QMSR require verification that corrective action was effective. A check with no pre-defined statistical criterion, no defined observation period and no defined sample size cannot fail. Defining these three things in advance — at the point the corrective action is written, not when it is closed — converts the effectiveness check from documentation into a real test, and the recurrence rate will fall for that reason alone.

Third, capability-building methods and delivery methods should be balanced deliberately. DMAIC and 8D deliver results and build modest capability, concentrated in specialists. A3 and Kata build capability distributed across the line and deliver results more slowly. A portfolio consisting only of the former produces an organisation that can solve problems only when a specialist is assigned — which caps the number of problems that can be solved per year at the number of specialists, and is precisely the ceiling most plants are operating against without recognising it.

5.6 The Daily Management System as the Substrate

Methods 4.17 to 4.22 do not sit alongside the methods of Sections 4.1 to 4.16; they sit beneath them. The daily management system — standardised work, leader standard work, visual controls, tiered accountability, the skills matrix, gemba, and a working problem-solving method — is the substrate that determines whether anything else functions.

The dependency is concrete rather than rhetorical. Detection: without standardised work there is no baseline, so abnormality is invisible and the improvement system has no input. Cadence: without tiered accountability, a detected problem waits for a meeting, and problem-solving cycle time is bounded below by meeting frequency. Capacity: without leader standard work, managerial attention is consumed by escalation and no observation occurs. Competence: without a real skills matrix, an improvement cannot be reliably deployed because the ability to execute the new standard is unknown. Retention: without standardisation following countermeasure, every gain decays to the mean.

This produces an uncomfortable but useful corollary. A plant with an excellent daily management system and only 5 Whys will out-improve a plant with a Six Sigma programme and no daily management, because the first plant detects and resolves a large volume of small problems quickly while the second detects few and resolves them slowly and expensively. The advanced analytical methods of Sections 4.9 to 4.15 multiply an existing improvement rate; they do not create one. If the multiplicand is near zero, the multiplier is irrelevant — which is the single most common reason expensive operational excellence programmes produce disappointing results.

Revision to the sequencing rule of Section 3.2

Layer L0 — the daily management system — precedes everything. Establish standardised work, visual controls, tiered accountability, leader standard work and one working problem-solving method before committing to any L2–L4 programme and long before any L5 investment.

The diagnostic is a single question asked on the floor: when something goes wrong on this line, how long is it before someone whose job it is to fix it knows about it? If the answer is measured in shifts rather than minutes, the plant’s binding constraint is at L0, whatever its improvement roadmap says.

6 An Integrated Maturity Model

Maturity models are frequently criticised, with justification, for encouraging organisations to pursue a level rather than a result. The model below is offered as a diagnostic instrument for identifying which capability is currently binding, not as a target. Its levels are defined by observable evidence rather than by self-assessment, and each level specifies what must be true, not what must be documented.

Level

Descriptor

Observable evidence

Binding constraint at this level

1

Reactive

Improvement is event-driven and follows complaints, audit findings or crises. Loss data exist only in financial aggregate. Overall equipment effectiveness is not measured or is measured inconsistently. No standardised work, so abnormality is not detectable; problems surface in shifts rather than minutes. Deviation investigations recur on the same modes.

Absence of detection and of loss visibility. Nothing can be prioritised because nothing is decomposed, and nothing is noticed in time to act.

2

Measured

Loss decomposition exists by cost of poor quality category, by six-loss category and by process step. Measurement systems have been formally assessed. Standardised work and visual controls are in place with a tiered accountability cadence, so abnormality is raised within the shift. A baseline capability study exists for each critical characteristic.

Absence of protected improvement capacity. The plant knows what is wrong and cannot get to it.

3

Systematic

Improvement capacity is committed and defended through strategy deployment. Projects are selected from the risk file and the loss decomposition. DMAIC and lean practice are standard. Problem-solving method is chosen by problem structure, root causes are verified experimentally, and countermeasures are classified by strength. Control mechanisms are verified at 180 days.

Change-control throughput. The improvement pipeline exceeds the plant’s ability to validate and implement.

4

Capable

Line-level personnel run structured experiments as routine. Design spaces exist for critical processes, reducing post-launch change events. Administrative and quality-system processes are mined and improved. Improvement rate is stable and predictable.

Observability. Further gain requires information the current instrumentation cannot supply.

5

Anticipatory

Process state is observed continuously and multivariately; defect modes are addressed by prediction and prevention rather than by detection; maintenance is condition-driven within validated envelopes; experimentation is substantially in-silico.

Model governance and organisational learning rate. The technical limits move outward faster than the organisation can absorb them.

Table 6.1 Five-level operational excellence maturity model with explicit binding constraints.

The final column is the substance of the model. Each level is characterised not by what the organisation has achieved but by what will stop it next, and the correct investment at any level is the one that relieves that constraint. A level-2 plant buying predictive analytics is relieving a level-4 constraint while its level-2 constraint — no protected improvement capacity — remains untouched. This is the most common and most expensive misallocation in the field, and the model exists chiefly to make it visible.

A note on levels 4 and 5 in regulated manufacture. The transition from level 3 to level 4 is where the change-control constraint becomes decisive, and it is why design-space methods (4.5) and process mining of quality-system processes (4.14) are disproportionately valuable at that transition. Plants that attempt to reach level 4 by adding improvement resource without addressing change-control throughput simply lengthen the queue.

7 Implementation Roadmap for a Regulated Ophthalmic Plant

The roadmap below assumes a plant at maturity level 2 — losses measured, capacity unprotected — which is the modal starting position. It is expressed in four waves over thirty-six months. Durations are indicative; the sequencing is not, and it follows the dependency structure of Table 5.3.

7.1 Wave 1 (months 0–6): diagnosis, capacity and measurement

Wave 1 gate: loss decomposition complete and independently reviewed; measurement adequacy quantified for all critical measurements; improvement capacity committed in writing; constraint identified with exploitation options costed; tiered accountability running with time-to-detection measured and falling; corrective action countermeasure profile baselined.

7.2 Wave 2 (months 4–15): stabilise and flow

Wave 2 gate: availability on constraint equipment above 90 percent; work-in-process capped and holding; two Six Sigma projects closed with verified 180-day sustainment; coaching cycles running at planned frequency; median A3 draft count above two; corrective actions at ranks 1–3 of Table 4.2 rising as a share of the total; 12-month recurrence rate baselined.

7.3 Wave 3 (months 12–26): design in, and build the data layer

Wave 3 gate: design space validated on at least one critical process with a demonstrated reduction in change events; data layer operational with verified part-level joins; multivariate monitoring in production with quantified lead time over univariate detection.

7.4 Wave 4 (months 24–36): predict and prevent

Wave 4 gate: at least one predictive model with prospective performance within a defined tolerance of its offline estimate; model governance operating; documented reduction in unplanned downtime on target assets; improvement rate sustained without additional headcount.

What is deliberately absent from this roadmap

Reinforcement-learning scheduling (4.13) does not appear. At evidence grade D, with scarce production deployment and unresolved verification difficulties, it does not belong in a committed thirty-six-month plan for a regulated plant. It belongs on a watch list, revisited annually, and should be entered only through a bounded evaluation against a properly tuned dispatching baseline. Sensor-integrated digital twins in the strict Kritzinger sense are similarly deferred: the digital model level answers most of the questions at a small fraction of the cost.

8 Measurement Architecture

8.1 Design principles

Measurement systems in improvement programmes fail in characteristic ways: they proliferate, they drift toward what is easy to collect, they aggregate away the variation that carries the information, and they are reported as means when the operational consequence lives in the tail. Four principles counter this.

8.2 The metric architecture

Tier

Purpose

Representative metrics

Cadence and audience

Tier 1 — Outcome

Establish whether the programme is producing value

Cost of poor quality as percentage of cost of goods sold (with four-category split); dock-to-dock lead time (median and P90); overall equipment effectiveness on constraint assets; on-time in-full; unit cost

Monthly; plant leadership and corporate

Tier 2 — Process

Establish whether the operational system is behaving as designed

Process capability by critical characteristic; first-pass yield by stage; work-in-process turns; buffer penetration by zone; unplanned downtime; schedule adherence; every-part-every-interval

Weekly; value stream and department level

Tier 3 — Improvement system

Establish whether the capability is growing

Project cycle time by DMAIC phase; change-control throughput and queue; experiments per learner per week; coaching cycles delivered against planned; improvement hours delivered against committed; 180-day sustainment rate

Monthly; improvement governance

Tier 4 — Measurement integrity

Establish whether the numbers can be trusted

Inspection false-discovery and escape rates; gauge study currency by instrument; model performance drift against qualification baseline; data-layer join completeness

Quarterly; quality engineering

Table 8.1 Four-tier metric architecture. Tier 4 is the one most often absent and the one on which the credibility of tiers 1 to 3 entirely depends.

8.3 Linking to financial outcome

The evidence on financial linkage is instructive and should discipline the business case. Swink and Jacobs (2012) found that the return-on-assets improvement associated with Six Sigma adoption arose predominantly from indirect cost reduction, with direct cost and asset productivity effects not significant. A benefit model that projects savings principally from direct labour and material is therefore projecting the effect that the best available controlled evidence did not find, and should expect to under-deliver.

For a device plant the implication is specific and actionable: the largest realistically capturable financial benefit is likely to lie in the quality and engineering overhead — deviation investigation labour, batch record review, rework administration, complaint handling, validation effort per change — rather than in the conversion cost of the product. This is precisely the domain that process mining (4.14) addresses, and it is a reason to sequence that method early despite its unglamorous subject matter.

Two governance rules follow. First, require finance validation of realised benefit at 180 days for every closed project, and track the ratio of validated to claimed benefit as a programme metric; a ratio persistently below about 0.7 indicates a benefit-modelling problem, not an execution problem. Second, account for the change-control cost of each improvement explicitly in the business case. An improvement worth 120,000 currency units annually that consumes 400 hours of validation effort has a materially different profile from one that consumes 40, and in a plant where validation capacity is the binding constraint the second is worth more than the first even at half the nominal benefit.