Industry 5.0, standardized work, skills architecture, Gemba, Kaizen, and PDCA/A3 as practical systems for daily improvement.
4.16 Human-Centric Operations and the Industry 5.0 Frame
At a glance |
|
Layer |
L1/L5 — Capability and socio-technical design |
Evidence grade |
B for the socio-technical evidence base; D for Industry 5.0 as a named programme |
Primary lever |
Design of work and technology around human capability, wellbeing and judgement rather than around automation alone |
Time to measurable effect |
Continuous; effects observable over 4–8 quarters |
Regulatory friction (device plant) |
Low to moderate. Human factors and usability engineering are already regulated design inputs under IEC 62366 for device design. |
4.16.1 Theoretical basis
Industry 5.0 is a policy-originated framing — human-centricity, sustainability and resilience — advanced as a corrective to the productivity-and-technology emphasis of Industry 4.0. Lu and colleagues (2022), in the Journal of Manufacturing Systems, set out the human-centric manufacturing outlook underpinning it. The literature is candid that Industry 5.0 is socio-politically pulled rather than technology-pushed, which distinguishes it from its predecessor and also explains why its empirical base is thin: it is a normative programme first and a research programme second.
Beneath the label, however, sits a substantial and much older evidence base. The soft-practice findings of Bortolotti and colleagues (2015), the critical-success-factor rankings of Netland (2016), the capability argument of Anand and colleagues (2009) and the coaching mechanism of Toyota Kata all point at the same conclusion from different directions: the human element is not a complement to the technical system but the binding constraint on it. Industry 5.0 is best treated as a contemporary label for a well-supported socio-technical proposition, and evaluated on the older evidence rather than on its own.
4.16.2 Mechanism
Three mechanisms are operationally relevant. Cognitive load management: operators in highly automated environments are asked to supervise systems they rarely intervene in, which degrades situation awareness and produces poor intervention when it is finally required — a well-documented ironies-of-automation effect that applies directly to automated optical inspection oversight. Augmentation rather than replacement: systems that present the operator with the model’s reasoning and permit informed override outperform both full automation and unaided human judgement in ambiguous cases. Retention of tacit process knowledge: in plants with long-tenure operators, a substantial fraction of process knowledge is undocumented, and both automation programmes and workforce turnover destroy it silently unless it is deliberately elicited.
4.16.3 Evidence
Grade B for the underlying socio-technical propositions, which rest on the large-sample lean and quality management literature cited throughout this review, and grade D for Industry 5.0 as a distinct programme, whose empirical corpus is dominated by conceptual frameworks, roadmaps and review articles. Recent reviews assemble the three-pillar framework and its supply-chain implications but do not yet supply controlled evidence of operational effect. Practitioners should adopt the mechanisms and be sceptical of the label.
4.16.4 Implementation protocol
Include operators in the design of automation, particularly at the interfaces they will supervise. This is a design requirement, not a change-management courtesy.
Design for informed override: where a model or automated system makes a call, present the basis for that call in a form the operator can evaluate, and record override decisions and reasons as data.
Maintain intervention competence deliberately. Where automation has removed routine manual involvement, schedule periodic manual practice so that the capability exists when it is required.
Elicit tacit knowledge systematically before automation programmes and ahead of known retirements, using structured methods rather than exit interviews.
Measure workload and situation awareness alongside throughput. Deteriorating human performance is a leading indicator of quality events and is not visible in production metrics.
Apply usability engineering methods (IEC 62366 in the device context) to internal manufacturing systems, not only to the device itself. The reasoning transfers directly.
4.16.5 Worked example — ophthalmic
The automated optical inspection system rejects at a rate that includes a known false-discovery component. Historically, borderline rejects were re-examined manually by experienced inspectors; a productivity programme removed that step, reducing labour cost and raising scrap.
The human-centric redesign restores the human to the loop, but at a different point and with a different tool. Instead of manually re-examining all borderline rejects, inspectors review a model-selected subset — those where the classifier’s confidence is low — with the classifier’s evidence displayed alongside the image. Their dispositions are recorded as labelled data and used to retrain the classifier on a defined cadence. The operators are thereby engaged in improving the system that supervises the process, rather than merely supervising it, and the arrangement produces three effects: recovered yield from correctly re-dispositioned parts, a continuously improving classifier, and retained inspection competence for the cases where it is required. Under the device quality system the re-disposition step is a documented, competence-controlled activity, which is straightforwardly accommodated.
4.16.6 Limitations and failure modes
Rhetoric without design change. Industry 5.0 language adopted with no alteration to how automation is specified or how work is designed.
Automation complacency. Operators asked to supervise reliable systems stop attending to them; the failure appears only at the rare intervention.
Measurement gap. Human-factors outcomes are harder to quantify than throughput, and consequently lose in resource competition unless leading indicators are instituted.
Weak programme-level evidence. Justify these interventions on the socio-technical and lean literature, not on the Industry 5.0 corpus.
4.16.7 Primary metrics
Operator override rate and disposition accuracy; time-to-intervention on abnormal conditions; skills matrix coverage on critical operations; suggestions per employee and implementation rate; tacit-knowledge elicitation coverage against retirement risk; safety and ergonomic incident rate; retention on critical roles.
4.17 Standardised Work and Leader Standard Work
At a glance |
|
Layer |
L0 — Daily management |
Evidence grade |
B for standardised work as part of the lean bundle; C for leader standard work specifically |
Primary lever |
Conversion of both operator work and managerial work from discretionary activity into a defined, auditable process |
Time to measurable effect |
1–2 quarters for adherence; 3–4 quarters for the sustainment effect |
Regulatory friction (device plant) |
Low to moderate, and largely favourable. Standardised work is already a quality-system expectation; leader standard work generates the management-review evidence ISO 13485 clause 5.6 requires. |
4.17.1 Theoretical basis
Standardised work is the documented current best-known method for performing a task, comprising three elements: takt time, work sequence, and standard work-in-process. Its function is frequently misunderstood. It is not principally a productivity device; it is the baseline against which abnormality becomes visible. A process without a standard cannot deviate from anything, so no abnormality can be detected, so no problem can be raised. Standardised work is therefore the precondition for every improvement method in this document — a point Taiichi Ohno made repeatedly and that organisations rediscover expensively.
Leader standard work extends the same logic upward. Mann (2014) argues that lean systems decay not because the technical changes fail but because the management system that sustains them was never designed. His prescription is a four-element management architecture — leader standard work, visual controls, a daily accountability process, and discipline — in which each layer of management has a documented set of recurring activities, at defined frequencies, with defined outputs. The theoretical claim is that managerial attention is the scarce resource that determines whether improvements persist, and that attention allocated by discretion drifts toward the urgent and away from the process.
The percentage of a leader’s time that is standardised varies systematically by level: conventionally around 80 percent for a team leader, falling to perhaps 25 percent for a value-stream manager and 10 percent for a plant manager. The gradient matters. Standardising a plant manager’s day would destroy the responsiveness the role requires; standardising nothing at team-leader level guarantees that first-line management collapses into firefighting.
4.17.2 Mechanism
Three mechanisms operate, and they are distinct. Abnormality detection: a standard converts an unbounded observation problem (is this process alright?) into a bounded comparison (does this match the standard?), which a person can perform in seconds rather than minutes. Attention protection: leader standard work reserves managerial time for process observation before the day’s escalations consume it — structurally the same mechanism as the improvement-capacity protection of Hoshin Kanri (4.1), applied at daily rather than annual cadence. Knowledge retention: the standard is the artefact in which improvement is recorded. An improvement not written into the standard is an improvement that leaves with the person who made it, which is why standardisation must follow every accepted countermeasure rather than precede the next problem.
The relationship between standardisation and improvement is frequently posed as a tension. It is better understood as a ratchet: the standard holds the gain so that the next experiment starts from the improved position rather than from the original one. Without the ratchet, improvement is Sisyphean, and the organisation experiences continuous activity with no cumulative movement.
4.17.3 Evidence
Grade B for standardised work, which appears within the just-in-time and total quality management bundles whose joint performance contribution is established at grade A (Shah and Ward 2003, 2007), though its individual contribution is not separately isolated. Grade C for leader standard work as a distinct construct: Mann’s framework is widely adopted and the supporting literature is predominantly practitioner and case-based rather than controlled.
The strongest adjacent evidence concerns visual management, which is the companion element in Mann’s architecture. Bateman, Philp and Warrender (2016), in a two-year longitudinal study of communication boards in a British lock manufacturer published in the International Journal of Production Research, found that applying visual management principles to board design measurably improved team leaders’ ability to engage their teams in problem solving and continuous improvement. That is a rare instance of a daily-management artefact studied over sufficient time to observe its effect on behaviour. Tezel, Koskela and Tzortzopoulos (2016) provide the corresponding literature synthesis and note — a caution that applies across this whole layer — that the field suffers from inconsistent terminology and weak conceptual definition.
4.17.4 Implementation protocol
Establish operator standardised work first, and establish it with the operators. Standards written by engineers for operators are followed until the engineer leaves the area. Document takt, sequence and standard work-in-process, and post them at the workplace, not in a binder.
Verify that the standard is executable in the time available. A standard that cannot be performed within takt will be silently abandoned, and the abandonment will not be reported.
Define leader standard work by level, with an explicit gradient of standardised time — high at team leader, low at plant manager. Specify frequency, location and required output for each recurring item.
Make the leader standard work document itself a visual, checked artefact. If nobody ever looks at whether the leader completed their standard work, it is a suggestion.
Install the daily accountability process: a short tiered meeting cadence in which each tier reviews its own visual controls, escalates only what it cannot resolve, and assigns owners and dates for gaps. Tier 1 at the line, tier 2 at value stream, tier 3 at plant; typically 10–15 minutes each, at the board, standing.
Design visual controls so that abnormality is detectable within a few seconds and by someone who does not work in the area. A board that requires interpretation is a report, not a control.
Close the loop to standardisation: every accepted countermeasure produces a revised standard, and the revision is trained out and confirmed. In a device plant this means the work instruction, the training record and the process control plan move together.
4.17.5 Worked example — ophthalmic
A packaging and inspection area operates with work instructions written for regulatory compliance — accurate, complete, and unused, because they describe requirements rather than method and run to eleven pages. Observed cycle times across four operators on the same operation range from 41 to 68 seconds, and the tray-loading sequence differs between all four. Nobody experiences this as a problem, because there is no standard from which to deviate.
Intervention. The team constructs standardised work with the operators over three shifts: takt of 52 seconds, a nine-step sequence, and a defined standard work-in-process of two trays at the station. The distinction from the existing work instruction is made explicit and is the key to acceptance — the work instruction remains the controlled quality-system document specifying what must be achieved, and the standardised work sheet specifies the current best method for achieving it. The two are cross-referenced, and only the former is subject to full change control. This separation is what makes rapid method improvement possible inside a validated environment, and it is worth designing deliberately rather than discovering by accident.
Leader standard work is then defined. The team leader’s day becomes roughly 75 percent standardised: a start-of-shift board review, three standardised-work adherence observations, an hourly production check against plan, a mid-shift gemba pass on the inspection station, and a fifteen-minute tier-1 accountability meeting. The supervisor’s day is roughly 50 percent standardised and includes a weekly audit of the team leader’s observations — the layered audit that prevents the first layer from decaying.
The measurable effects within two quarters are cycle-time variation reduced to a range of 49 to 56 seconds, and — the more consequential result — the mean time between an abnormality occurring and its being raised falling from roughly a shift to under an hour. That second number is what makes every other method in this document faster, because problem-solving cycle time is bounded below by problem-detection time.
4.17.6 Limitations and failure modes
Standard imposed rather than built. The most reliable way to produce a document nobody follows. Operators must construct the standard; the engineer’s role is to challenge and to time.
Confusion of standardised work with the controlled work instruction. In regulated plants this conflation makes every method improvement a change-control event and therefore stops method improvement. Separate the documents deliberately.
Leader standard work as a compliance checklist. When completion is audited but content is not, leaders produce ticks. Audit the output of the standardised activity — what abnormalities were found — not its completion.
Over-standardisation at senior level. A plant manager whose day is 70 percent standardised cannot respond to anything, and the system will be discarded wholesale rather than corrected.
Boards that report rather than control. If the board shows month-to-date performance rather than hour-by-hour deviation, it cannot trigger action within the shift and is therefore not a control.
4.17.7 Primary metrics
Standardised-work adherence (observations conforming as a proportion of observations made); cycle-time variation by operation; time from abnormality occurrence to abnormality raised; leader standard work completion and abnormalities detected per observation cycle; escalations resolved at tier 1 as a proportion of total; standards revised per quarter following countermeasures.
4.18 Skills Matrix, Cross-Training and Competence Architecture
At a glance |
|
Layer |
L0 — Daily management; L1 — Capability |
Evidence grade |
B — cross-training has a substantial analytical and empirical operations literature; the skills matrix as an artefact is grade C |
Primary lever |
Deliberate design of workforce flexibility and competence depth against demand variability and risk |
Time to measurable effect |
2–4 quarters |
Regulatory friction (device plant) |
Moderate, and largely favourable. ISO 13485 clause 6.2 requires competence determination, training and effectiveness evaluation; a skills matrix is the natural evidence. |
4.18.1 Theoretical basis
The skills matrix — a grid of people against operations, with a graded symbol at each intersection — is one of the most widely used and least well theorised artefacts in operations. Its usual four-level scale (commonly rendered as a quartered circle, or the ILUO convention) distinguishes aware, can perform with support, can perform independently, and can train others. The final level is the one that carries the system, because an organisation’s ability to grow competence is bounded by its stock of people who can transmit it — precisely the constraint identified for coaching in 4.2.
The underlying operations theory is cross-training, and it is well developed. Hopp and Van Oyen (2004), in IIE Transactions, set out the Agile Workforce Evaluation framework, which structures the strategic question (what mechanism is cross-training supposed to serve — capacity balancing, absorption of absence, quality through ownership, or job enrichment?) and the tactical question (which workforce architecture and which worker-coordination policy fit the process). Their central contribution is the demonstration that cross-training is not a single decision but a design space with distinct architectures, and that choosing the architecture without first identifying the mechanism produces flexibility that is expensive and unused.
A robust and non-obvious result from this literature is that flexibility exhibits strong diminishing returns: chaining — where each worker is trained on two adjacent operations, forming a closed loop across the line — captures most of the benefit of full cross-training at a small fraction of the training cost. Full cross-training of everyone on everything is almost never the economic optimum, and it is the default that untutored skills-matrix programmes converge on.
4.18.2 Mechanism
Four mechanisms, which should be distinguished because they imply different matrix targets. Capacity pooling: flexible workers can be redeployed to the momentary bottleneck, which reduces the variability the flow system must buffer — directly relevant to 4.3 and 4.6. Absence absorption: coverage depth ensures an operation is never single-manned, which converts a random absence from a line stoppage into a reassignment. Quality through ownership: workers who understand adjacent operations detect upstream-caused defects that a narrowly trained worker cannot interpret. Capability propagation: level-four (can-train) competence is the transmission mechanism for every standard and every improvement.
In a regulated plant a fifth mechanism appears that has no analogue in general manufacturing: competence as a validated control. Where a process risk control is “operator performs X correctly”, competence is not a human-resources matter but a risk control under ISO 14971, and the skills matrix becomes part of the evidence that the control is effective. This reframing is worth making explicit, because it moves the matrix from a training-department artefact to a quality-system one and changes who cares about its accuracy.
4.18.3 Evidence
Grade B. Human resource management is one of the four practice bundles in Shah and Ward (2003), and training and multi-skilling sit within the soft lean practices that Bortolotti, Boscari and Danese (2015) identify as discriminating successful from unsuccessful lean plants — which is the strongest available evidence that competence development is causally implicated rather than merely correlated. The cross-training design literature, of which Hopp and Van Oyen (2004) is the reference synthesis, is analytically rigorous and supported by queueing and simulation results, though field validation of specific architectures is thinner. The skills matrix as a specific artefact has no dedicated controlled evidence base; grade C is appropriate for the artefact and B for the underlying practice.
4.18.4 Implementation protocol
State the mechanism before building the matrix. Which of capacity pooling, absence absorption, quality through ownership, capability propagation, or risk control is the flexibility for? Different answers imply different target patterns, and a matrix built without this question produces an aspiration to universal competence that will not be funded.
Define the competence scale operationally, with an observable criterion for each level. “Can perform independently” must mean a specific, assessed demonstration, not a supervisor’s impression. Ambiguous scales make the matrix unfalsifiable and therefore useless as a control.
Map criticality first: which operations are risk controls, which are constraint operations, which have long competence acquisition times. Depth requirements follow criticality, not headcount convenience.
Set a minimum-depth rule per operation — commonly at least two independent performers per shift for any operation, and at least two can-train performers per operation across the plant — and treat violations as open risks with owners.
Prefer a chained pattern over universal cross-training. Train each person on their own operation plus one or two adjacent ones, arranged so that the coverage graph is connected. This captures most of the pooling benefit at a fraction of the cost.
Link the matrix to the training record, the work instruction version and the competence assessment date, so that a work-instruction revision automatically flags the affected competences for re-verification. Absent this link, the matrix and the quality system drift apart and the matrix becomes decorative.
Review at the daily accountability meeting (4.17), not annually. Coverage gaps are an operational risk with a same-week resolution, not a training-plan item.
4.18.5 Worked example — ophthalmic
An inspection and packaging area of 34 operators covers eleven distinct operations. The initial matrix, built for audit purposes, shows nominal competence broadly distributed and reveals nothing. Rebuilt against operational criteria, it exposes three findings that were previously invisible.
First, two operations — automated optical inspection system set-up and the sterilisation load-pattern verification — have exactly one can-train performer each, both on day shift, and both within four years of retirement. These are also risk-control operations under the device risk file. This is a single-point-of-failure on a validated control, and it is a finding that belongs in the risk register, not the training plan.
Second, competence depth is inversely correlated with operation criticality: the easily learned operations have deep coverage because they are easy to train, and the critical operations have shallow coverage because they are not. This is the natural equilibrium of an unmanaged training system and it is exactly backwards.
Third, of 34 operators, 21 are trained on three or more operations, but the coverage graph is not connected — two clusters exist with no operator bridging them, so the nominal flexibility cannot actually be used to rebalance across the area. Restructuring the training plan to connect the graph, at a cost of eleven cross-training assignments, delivers more usable flexibility than the previous forty-plus assignments did.
Countermeasures: a can-train depth target of two per critical operation with a dated plan; competence assessment for risk-control operations moved to an assessed demonstration against defined acceptance criteria rather than a training sign-off; and the coverage gap on the two critical operations entered on the risk register with an owner. The third of these is what changes behaviour, because it moves the issue from a queue that is reviewed annually to one that is reviewed weekly.
4.18.6 Limitations and failure modes
Matrix as decoration. Built for audit, coloured in optimistically, never used for a staffing decision. Diagnostic: has the matrix ever caused a shift assignment to change?
Unfalsifiable competence levels. Without observable criteria the levels record familiarity, not capability, and the matrix systematically overstates coverage.
Universal cross-training as the default target. Expensive, slow to achieve, and dominated by chained patterns on both cost and benefit.
Skill decay unmodelled. Competence on infrequently performed operations degrades. Without a currency rule — a required minimum practice frequency — the matrix records historical rather than present capability, which is worse than recording nothing.
Disconnection from document control. When a work instruction is revised and the matrix does not flag re-verification, the plant holds training records against superseded methods.
4.18.7 Primary metrics
Coverage depth per operation (independent performers, and can-train performers, per shift); number of single-point-of-failure operations, weighted by criticality; connectivity of the coverage graph; competence currency (proportion of recorded competences practised within the currency window); time to independent competence for new starters by operation; proportion of risk-control operations with assessed rather than signed-off competence.
4.19 Gemba Walks and Genchi Genbutsu
At a glance |
|
Layer |
L0 — Daily management |
Evidence grade |
C — widely practised, sparsely evidenced as an isolated intervention |
Primary lever |
Direct observation of the actual process at the actual place, replacing report-mediated management |
Time to measurable effect |
1–2 quarters for information quality; 3–4 quarters for behavioural effect |
Regulatory friction (device plant) |
Low. Observation is not a controlled process; findings entering the quality system follow normal routes. |
4.19.1 Theoretical basis
Gemba (現場, the actual place) denotes where value is created. Imai (1997) advanced gemba as the primary locus of management attention, and the associated Toyota principle of genchi genbutsu — go and see for yourself — is presented by Liker (2004) as one of the system’s foundational management behaviours. The theoretical content is epistemic rather than motivational, and the distinction matters: the argument is not that visiting the floor improves morale but that reports are lossy in a systematically biased way. Aggregation removes variance, which is where the diagnostic information lives; summarisation removes the anomalies that do not fit the reporting categories; and every reporting layer applies a filter shaped by the reporter’s interests. Direct observation is the only channel that does not have these properties.
A gemba walk is therefore best understood as a designed sampling procedure for an information channel with different bias characteristics from the reporting channel — not as a substitute for it. The two are complementary, and the walk is valuable precisely on the dimensions where the report is weakest: variation, exception handling, workarounds, and the gap between the documented process and the executed one.
4.19.2 Mechanism
Three mechanisms. Bias correction: observing what actually happens corrects the systematic optimism of upward reporting, particularly regarding workarounds — the informal adaptations by which operators make an unworkable process work, which never appear in any report and which are simultaneously the plant’s largest source of process knowledge and its largest source of undocumented risk. Latency reduction: an issue observed today is raised today rather than appearing in next month’s metrics. Signalling: the questions a leader asks on the floor define, more reliably than any policy statement, what the organisation attends to. A leader who asks only about output teaches the organisation that only output matters, whatever the posters say.
4.19.3 Evidence
Grade C, and the grade should be stated honestly because the practice is often presented as self-evidently effective. Academic literature treating the gemba walk as an isolated intervention is sparse; it is almost always studied as a component of a broader lean programme, which makes attribution impossible. Reviews of lean implementation in healthcare identify gemba walks among the interventions used in a minority of studies, with generally positive reported outcomes on efficiency, quality, cost and satisfaction — but the confound with co-deployed practices is complete. Manufacturing case reports show substantial local improvements attributed to gemba-based observation, with the usual limitations of single-case evidence.
The defensible position is that the epistemic argument is sound and does not depend on the empirical literature — direct observation of a process genuinely does carry information that aggregated reporting cannot — while claims about the magnitude of the operational effect should be treated as unestablished.
4.19.4 Implementation protocol
Define the purpose of each walk type and do not mix them. A process-confirmation walk (is the standard being followed and does it work?), an improvement-coaching walk (4.2), a problem-solving walk on a specific issue, and a leadership-presence walk are four different activities with different questions and different outputs. Mixing them produces a walk that achieves none of them.
Schedule walks into leader standard work (4.17) with defined frequency, route and duration. An unscheduled walk is a walk that does not happen in a difficult week — which is exactly the week it is most needed.
Prepare: know the area’s current performance, its open countermeasures and its standards before arriving. A leader who has to be briefed on the floor consumes the very attention the walk was meant to supply.
Ask about the process, not about the person. “Show me how you know this is running to standard” is a process question; “why is your output down?” is an interrogation, and it reliably terminates the flow of information for months.
Observe the work, not the board. The board is a report and is already available. The value of being present is access to the work.
Record findings and close them visibly. Findings that vanish teach the area that raising problems is pointless, which is a durable and expensive lesson.
Attend explicitly to workarounds. When an operator has adapted the documented method, the adaptation encodes real process knowledge — and, in a device plant, a potential deviation from a validated method. Both facts require action; suppressing the disclosure loses the knowledge and retains the risk.
4.19.5 Worked example — ophthalmic
Weekly process-confirmation walks are established on the hydration and inspection lines, scheduled into the value-stream manager’s standard work, with a fixed route and a standard question set. Over one quarter the walks surface three classes of finding, and their distribution is itself the most informative result.
Workarounds (6 of 14 findings). Operators at the tray-transfer station were pre-staging trays out of sequence to compensate for an ergonomic reach problem, in a way that broke the first-in-first-out assumption on which the hydration dwell-time control depended. This had persisted for an estimated eighteen months, appeared in no report, and was disclosed within four minutes of a leader asking how the operator decided which tray to take next. The dwell-time exposure was quantified, the ergonomic cause was corrected, and the sequence control was made physical rather than procedural.
Standard-versus-practice gaps (5 findings). Documented cleaning frequencies for the inspection optics were being met, but the method varied between operators in a way that plausibly affected the false-discovery rate of the inspection system — which, as established in 4.9 and 4.10, propagates into every downstream quality number.
Genuine improvement ideas (3 findings). Raised by operators who had held them for some time and had no route by which to raise them.
The instructive point is the ratio. Nearly half the findings were workarounds — process knowledge that existed in the organisation, was operationally significant, was invisible to every reporting channel, and was available for the cost of asking a well-framed question in the right place. No analytical method in this document would have found them.
4.19.6 Limitations and failure modes
Inspection tourism. Walks that observe without changing anything consume floor time and produce cynicism. The test is whether findings are closed.
The walk as audit. Where the walk is experienced as evaluation of people, disclosure stops and the information channel closes — permanently, and without any visible signal that it has closed.
Solving on the spot. A leader who supplies answers prevents the area from developing the capability to find them, undoing the mechanism of 4.2.
Purpose conflation. Attempting process confirmation, coaching and presence in one walk achieves none of them well.
Workaround suppression. Punishing disclosed deviation guarantees that future deviations are concealed, which converts a knowledge problem into a compliance risk.
4.19.7 Primary metrics
Walks completed against scheduled; findings raised per walk (a decline over time is ambiguous and should be investigated, not celebrated); findings closed and mean days to closure; proportion of findings originating from operators rather than from the leader; workarounds identified and dispositioned; repeat findings (indicating closure without resolution).
4.20 Kaizen: Events and Daily Kaizen
At a glance |
|
Layer |
L0/L2 — Daily management and flow |
Evidence grade |
B — multi-organisation field studies with quantified factor analysis |
Primary lever |
Concentrated, time-boxed improvement (events) and continuous small-scale improvement by the people who do the work (daily kaizen) |
Time to measurable effect |
Days for the event; 3–6 quarters for daily kaizen to become self-sustaining |
Regulatory friction (device plant) |
Moderate to high for events. A three-day event that produces process changes on a validated process generates a change-control load that must be planned before the event, not after. |
4.20.1 Theoretical basis
Kaizen denotes change for the better, and in practice covers two quite different activities that should not be conflated. The kaizen event (also kaizen blitz, rapid improvement event) is a structured, time-boxed — typically three to five day — cross-functional team activity focused on a defined area with a defined objective. Daily kaizen is the continuous flow of small improvements generated and implemented by the people doing the work, without event structure. Imai (1997) treats the latter as the substance of the philosophy and the former as an intervention.
The theoretical justification for the event format is concentration: removing participants from normal duties for a bounded period overcomes the activation barrier that prevents improvement work from starting, and co-locating cross-functional participants collapses the handover delays that would otherwise extend the work over months. The justification for daily kaizen is different and, over a long horizon, stronger: the people performing the work possess process knowledge that no external team can acquire in a week, and the aggregate of many small improvements exceeds the yield of a few large ones while consuming far less change-control capacity per unit of benefit.
4.20.2 Mechanism
Events work by suspending the ordinary constraints — time, cross-functional coordination, decision authority — for a bounded period. That is also precisely why their results decay: the constraints return on Monday. The empirical literature accordingly distinguishes the technical outcomes of an event from its human resource outcomes (attitude, commitment, problem-solving capability), and finds that the latter are what predict whether the former survive.
Daily kaizen works by a different mechanism entirely: it lowers the transaction cost of a small improvement to near zero, so that improvements below the threshold worth organising a project for — which is the overwhelming majority of improvements available in a mature plant — actually get made.
4.20.3 Evidence
Grade B, and unusually well specified for a practice of this kind, thanks to a sustained programme of field research. Farris, Van Aken, Doolen and Worley (2009), studying 51 kaizen events across six manufacturing organisations and publishing in the International Journal of Production Economics, identified the input and process factors most strongly associated with employee attitudinal outcomes and problem-solving capability development. Glover, Farris, Van Aken and Doolen (2011), extending to 65 events across eight organisations, addressed the harder question of what makes event outcomes sustain — an issue the authors note had received limited empirical attention despite widespread reporting of positive short-run results.
The consistent finding across this programme is that the factors predicting durable outcomes are organisational and social — goal clarity, team composition and autonomy, management support, and the post-event follow-through mechanism — rather than technical. This mirrors the lean evidence of Bortolotti and colleagues (2015) and the critical-success-factor evidence of Netland (2016) closely enough that the convergence should be treated as a robust finding about improvement generally, not a fact about kaizen events specifically.
A caution the literature supports: event results decay by default. Sustainment is the exception achieved by deliberate design, not the normal case. Any programme that plans events without planning the follow-through mechanism is planning a temporary result.
4.20.4 Implementation protocol
Select the target from the loss decomposition, not from availability or enthusiasm. An event on a non-constraint is a well-organised way to spend a week.
Scope tightly. A defined area, a defined objective with a numeric target, and a defined boundary of authority. Events with open scope produce action lists rather than changes.
Compose the team deliberately: a majority of people who do the work, plus the functions whose agreement is required to change it — maintenance, quality, and in a device plant, regulatory affairs. The presence of the change-control decision-maker in the room is worth more than any facilitation technique.
Complete the change-control assessment before the event, not after. Classify in advance which candidate changes would be implementable within the week, which require verification, and which require revalidation, and set the event scope accordingly. Events that generate a queue of unimplementable changes damage credibility more than they help.
Grant real decision authority within the defined boundary. An event whose outputs are recommendations is a workshop.
Design the follow-through mechanism before the event begins: named owners, dated actions, a 30- and 60-day review, and the revised standard (4.17). The evidence identifies this as the sustainment determinant.
Measure human-resource outcomes as well as technical ones, since the former predict the durability of the latter.
Build daily kaizen in parallel and treat it as the primary system. A defined route for a small improvement — raise, assess, implement, standardise — with a target closure time measured in days, and with a pre-agreed risk-based classification that keeps low-risk method improvements out of the full change-control path.
4.20.5 Worked example — ophthalmic
A four-day event targets changeover on the moulding line, with the objective of reducing mean changeover from 96 to below 60 minutes (the target condition of the example in 4.2). The team comprises four operators, a setter, a maintenance technician, a process engineer, a quality engineer and — a deliberate inclusion — the change-control coordinator.
Pre-event classification proves decisive. Of eleven candidate changes identified in preparation, seven are external-preparation and sequencing changes with no effect on validated process parameters, implementable within the week; three require verification against existing validation, implementable within four weeks; one, a tool pre-heat, would alter a validated thermal profile and require revalidation. The event scope is set to the first two groups. The pre-heat item is routed instead to the design-space work of 4.5, where it belongs, rather than becoming an event output that stalls for two quarters and is remembered as a failure.
Outcome: mean changeover 58 minutes at event close, with standard deviation improving from 31 to 12 minutes — the variance reduction being the more valuable result, because it is what makes the every-part-every-interval calculation of 4.3 reliable. Follow-through: revised standardised work for the changeover, the setter trained to can-train level, a 30-day and 60-day confirmation, and the changeover time added to the tier-1 board (4.17).
At 90 days, mean changeover measures 63 minutes — some decay, as the literature predicts, but bounded by the standard and detected by the board rather than discovered a year later. The decay itself becomes the next problem, which is the correct outcome.
4.20.6 Limitations and failure modes
Event addiction. A plant that runs many events and has no daily kaizen has purchased an improvement function rather than built an improvement capability. Events should decline in frequency as daily kaizen matures.
Decay by default. Without designed follow-through, results regress. This is the modal outcome, not the exceptional one.
Change-control collision. Events that generate changes the plant cannot validate produce a visible backlog and a durable belief that improvement is futile.
Scope inflation. Broad scope converts the event into a planning workshop.
Recommendation-only authority. Removes the mechanism that makes the format work.
Participation without the work-doers. Events staffed by engineers and managers miss the process knowledge that is the format’s main advantage over a project.
4.20.7 Primary metrics
Target achievement at event close and at 30, 60 and 90 days (the decay profile is the metric that matters); proportion of event actions closed on date; daily kaizen ideas raised, implemented, and mean days to implementation; implementations per employee per year; ratio of daily-kaizen benefit to event benefit (a maturity indicator that should rise); change-control events generated per improvement, split by risk classification.
4.21 PDCA/PDSA and the A3 Process
At a glance |
|
Layer |
L0/L3 — The core improvement cycle and its documentary discipline |
Evidence grade |
B for the cycle as a framework; the strongest available evidence concerns poor fidelity of application rather than efficacy |
Primary lever |
A disciplined experimental cycle, and a one-page structure that makes the reasoning visible and coachable |
Time to measurable effect |
Immediate at the individual problem level; 3–6 quarters for organisational competence |
Regulatory friction (device plant) |
Low to moderate. The cycle is a reasoning discipline; only the resulting countermeasures enter change control. |
4.21.1 Theoretical basis and a necessary historical correction
The improvement cycle has a more tangled lineage than its ubiquity suggests, and the confusion has practical consequences. Moen and Norman trace the evolution: Shewhart presented a three-step cycle of specification, production and inspection in 1939, explicitly analogising it to the scientific method; Deming presented a four-step version — the Deming Wheel — in Japan in 1950; Japanese practice reformulated this as plan-do-check-act during the 1950s; and Deming himself, in 1986 and subsequently, returned to plan-do-study-act, objecting specifically to “check” on the grounds that it connotes holding back or inspecting rather than learning, and stating that PDSA rather than PDCA was the intended form.
This is not antiquarianism. The substitution of check for study marks precisely the degradation that the empirical evidence documents: a cycle in which the third step asks “did we do what we said?” rather than “what did the result teach us about our understanding?” loses the epistemic content and becomes a project-management checklist. Toyota Kata (4.2) can be read as an attempt to restore the study step by requiring an explicit prediction before the do step, since one cannot study a result without a prediction to compare it against.
The A3 process is the documentary and coaching discipline built around the cycle. The name refers to the paper size — a single A3 sheet, chosen precisely because the constraint forces selection and prevents the accumulation of undigested analysis. The conventional structure runs: background, current condition, goal or target condition, root cause analysis, countermeasures, implementation plan, follow-up and verification. Shook (2008) makes the essential point that the A3 is not a report format but a coaching mechanism: its value lies in the iterative dialogue between an author and a mentor over successive drafts, in which the mentor’s role is to interrogate the reasoning rather than supply the answer. Sobek and Smalley (2008), whose treatment received a Shingo Research and Professional Publication Prize in 2009, situate A3 thinking as a component of Toyota’s PDCA management system rather than as a standalone tool.
The distinction is the crux of the method. An A3 completed alone and submitted is a form. An A3 developed through four or five mentored iterations is a training intervention that happens to also solve a problem.
4.21.2 Mechanism
PDCA/PDSA works by imposing an experimental structure on action: state what you expect, act, compare, and revise your understanding. Its power is entirely in the comparison step, and its characteristic failure is the omission of that step.
The A3 adds three mechanisms the bare cycle lacks. Space constraint forces prioritisation and prevents analysis accumulation — a genuine and underrated design feature. Left-to-right logical flow makes the reasoning chain inspectable, so that a mentor can identify precisely where an argument breaks: a countermeasure that does not connect to a cause, a cause that does not explain the current condition, a target that does not close the stated gap. Iteration with a mentor converts problem solving into a teaching occasion. The one-page format is also, incidentally, an effective organisational communication device, but that is a by-product; organisations that adopt A3 for the communication benefit and skip the mentoring have kept the packaging and discarded the contents.
4.21.3 Evidence
The evidence here is unusual and deserves careful statement, because the most rigorous study of the cycle is a critique of how it is used rather than a test of whether it works.
Taylor, McNicholas, Nicolay and colleagues (2014), in a systematic review published in BMJ Quality & Safety, examined the application of the plan-do-study-act method across the healthcare improvement literature. Their central finding is stark: the principles of the method are frequently not followed. Fewer than 20 percent of reviewed articles reported conducting iterative cycles of change, and of those that did, only about 15 percent reported using initial small-scale tests with scale increasing as confidence developed. The authors conclude that reliable conclusions about the effectiveness of PDSA cannot presently be drawn, and call for systematic standards for the application and reporting of the method.
The correct inference is not that the cycle does not work. It is that what most organisations call PDCA is not PDCA — it is a linear plan-and-implement sequence wearing the label. The two defining features, iteration and small-scale testing before scaling, are precisely the features most commonly absent. Any assessment of a plant’s improvement capability should therefore test for those two features directly rather than accepting the presence of PDCA vocabulary as evidence.
For the A3 specifically, the evidence base is grade C: the canonical sources (Shook 2008; Sobek and Smalley 2008) are practitioner-scholarly rather than empirical, and controlled evaluation is absent. The mechanism is coherent and the practice is widely reported as effective by those who sustain it, with the same coach-scarcity constraint that limits Toyota Kata.
4.21.4 Implementation protocol
Train the cycle as an experimental discipline, not as a project template. The operational test of whether this has landed: does the plan step contain a written prediction of the result, and does the study step compare against it?
Enforce small-scale testing before scaling. One machine, one shift, one product family. This is the single most frequently omitted element and the one the evidence identifies most clearly.
Require iteration. A single pass through the cycle is not PDCA; it is a plan that was implemented. Set an expectation of multiple cycles per problem and treat a one-cycle resolution as a signal that the problem was trivial or that the cycle was skipped.
For A3s, assign a mentor for every author and require a minimum number of drafts — typically three to five — before the A3 is considered complete. If the first draft is accepted, no teaching has occurred.
Coach the left side before the right. Inexperienced authors rush to countermeasures; the mentor’s discipline is to refuse to discuss the right half of the page until the current condition is quantified and the cause chain is evidenced.
Require the follow-up section to be completed after implementation, with actual results against predicted. An A3 archived at the implementation step has discarded the learning.
Keep A3s visible and accessible. In a device plant, cross-reference the A3 to any resulting deviation, corrective action or change request, so that the reasoning and the quality-system record point at each other.
4.21.5 Worked example — ophthalmic
Problem: hydration bath ionic-strength excursions producing dimensional non-conformance, at a rate of roughly two events per month.
Weak practice (what usually happens). A team meets, concludes the cause is inadequate monitoring, installs continuous ionic-strength monitoring with an alarm across all four baths, closes the action, and reports the problem solved. This is a single-pass, full-scale implementation with no prediction and no comparison — the pattern Taylor and colleagues found dominant. Six months later the excursion rate is roughly unchanged, because the monitoring detects excursions it does not prevent, and the team has spent its capital.
A3 practice. Current condition: 14 excursions over seven months, stratified — 11 on baths 2 and 3, 3 on baths 1 and 4; 9 within four hours of a bath make-up; mean deviation magnitude 6.2 percent above the upper control limit. The stratification is done before any cause is proposed and it immediately narrows the problem from a plant-wide monitoring gap to a make-up-procedure question on two specific assets. Cause analysis: cause-and-effect development followed by evidence testing rather than by consensus, arriving at a chain — make-up concentrate dosing is volumetric, concentrate density varies with storage temperature, baths 2 and 3 draw from a storage location adjacent to an external wall with a wider ambient range. Each link is tested against data rather than accepted by plausibility. Target condition: excursions below two per year. Countermeasure: gravimetric rather than volumetric dosing, tested first on bath 3 only — a small-scale test on the worst-performing asset.
Cycle 1 result versus prediction: predicted elimination on bath 3; observed a reduction but not elimination, with one residual excursion. The prediction was wrong, and that is the useful outcome. Cycle 2 investigates the residual and identifies a second contributor — a dissolution-time assumption in the procedure that holds at 20 °C but not at 15 °C. Countermeasure revised to include a temperature-dependent dissolution hold. Cycle 3 confirms on bath 3, then scales to baths 2, 1 and 4 in sequence with confirmation at each step.
Two features distinguish the second account and both are the ones the evidence identifies as absent in practice: the countermeasure was tested at small scale before scaling, and a wrong prediction was treated as information rather than as a failure to be concealed. Note also that only the final converged countermeasure entered change control, once — whereas the weak practice generated a full-scale change on four validated assets on the strength of an untested hypothesis.
4.21.6 Limitations and failure modes
PDCA in name only. Single-pass, full-scale, no prediction, no comparison. Empirically the dominant form (Taylor et al. 2014). Test for iteration and small-scale testing explicitly; do not accept the vocabulary as evidence.
“Check” instead of “study”. Verifying that actions were completed rather than what the result taught. This is the specific degradation Deming objected to and it is the most common one.
A3 as a form. Completed alone, submitted, filed. Without mentored iteration the method is a template, and templates do not develop people.
Right-side rush. Countermeasures written before the current condition is quantified. The most common coaching failure and the reason the mentor must hold the line on the left half of the page.
Follow-up never completed. The verification section is the only part that closes the learning loop and it is the part most often left blank.
Coach scarcity. As with Toyota Kata, the binding constraint is the number of people who can mentor an A3 competently.
4.21.7 Primary metrics
Proportion of improvement work with a documented prior prediction; cycles per problem (a distribution centred on one indicates the method is not being practised); proportion of countermeasures tested at small scale before scaling; A3 drafts per completed A3; proportion of A3s with a completed follow-up section containing actual versus predicted results; recurrence rate of problems previously closed.
Talk to us →