Defining operational excellence, grading the evidence, building a layered method system, and applying Hoshin Kanri, Toyota Kata, and Lean production systems.
1 Executive Summary
Operational excellence is not a portfolio of tools. It is a designed control architecture in which strategy, flow, variation, asset reliability and data are coupled through explicit feedback loops. This review maps twenty-two methods that span that architecture, from the daily management substrate to the advanced analytical frontier, grades the empirical evidence behind each, and specifies how they compose in a regulated ophthalmic and medical device manufacturing environment.
Three findings organise the argument. First, the evidence for classical methods is strong but bounded and asymmetric. Large-sample studies establish that lean practice bundles explain roughly a quarter of the between-plant variance in operational performance (Shah and Ward 2003), and that Six Sigma adoption produces a measurable return-on-assets improvement — but one driven predominantly by indirect cost reduction rather than by direct cost or asset productivity gains (Swink and Jacobs 2012). Practitioners routinely promise the opposite. Method selection should therefore follow the demonstrated mechanism, not the marketing claim.
Second, the socio-technical component dominates the technical component. Bortolotti, Boscari and Danese (2015) show that so-called hard lean practices (pull, setup reduction, statistical process control) behave as order qualifiers: successful and unsuccessful plants adopt them at similar rates. What discriminates successful implementations is the extensive use of soft practices — small-group problem solving, training, supplier and customer involvement, leadership for continuous improvement — inside a supportive culture. Netland (2016), surveying 432 practitioners across 83 factories in two multinationals, reaches a compatible conclusion: management commitment and capability building outrank technical prerequisites among critical success factors, and the ranking is stable across factory size, national culture and implementation stage.
Third, digital methods are additive, not substitutive. Zero Defect Manufacturing, machine-learning-augmented process control, digital twins, process mining and reinforcement-learning scheduling all presuppose a stable, well-instrumented, standardised process. Applying them to an unstable process encodes the instability into the model. The correct sequencing — stabilise, then standardise, then instrument, then automate the inference — is the single most consequential design decision in an operational excellence programme, and the most frequently violated.
Fourth, and least fashionable, the daily management substrate governs everything above it. Standardised work, leader standard work, visual controls, tiered accountability, a real competence architecture and one working problem-solving method determine how quickly an abnormality is detected and resolved. Every analytical method in Sections 4.1 to 4.16 multiplies an existing improvement rate; none of them creates one. The evidence here is also sobering: a systematic review of plan-do-study-act application found that fewer than 20 percent of published applications reported iterative cycles at all, and only about 15 percent of those used small-scale testing before scaling (Taylor et al. 2014). What most organisations call structured problem solving is a single-pass plan-and-implement sequence wearing the label.
For a contact lens and ophthalmic device manufacturer, three domain constraints bind the design. Products are Class II or Class III devices produced at very high volume, very low unit cost and very tight optical tolerance; the dominant defect modes are cosmetic and dimensional, appear at parts-per-million rates, and are detected by automated optical inspection rather than by human judgement. Process changes are subject to design control, process validation and change control, which lengthens the improvement cycle materially: a documented Lean Six Sigma project in a medical device firm required nine months largely because of validation and regulatory approval activity that would not exist in an unregulated plant (McGrane et al. 2022). And since 2 February 2026 the FDA Quality Management System Regulation has incorporated ISO 13485:2016 by reference into 21 CFR Part 820, replacing the former Quality System Regulation and shifting inspection to a new compliance programme.
The practical consequence is that in this sector, improvement velocity is governed by the validation system, not by the analytical toolkit. Methods that reduce the number of validated changes required per unit of improvement — design-space thinking, statistical engineering, digital twin experimentation, risk-based change classification — are worth disproportionately more than methods that merely accelerate analysis. This document is organised around that insight.
How to read this document Section 3 gives the taxonomy and the layer model. Section 4 contains twenty-two method monographs — sixteen strategic, flow, quality, reliability and digital methods (4.1–4.16), then six daily-management and problem-solving methods (4.17–4.22) — each with theoretical basis, mechanism, graded evidence, an implementation protocol, a worked ophthalmic example, limitations and primary metrics. Section 5 provides the comparative analysis: a fourteen-dimension matrix, a head-to-head comparison of the problem-solving methods including PDCA against A3, an evidence-strength grading, and a selection decision procedure. Sections 6 to 8 give the maturity model, a thirty-six-month roadmap and the measurement architecture. Section 9 treats failure modes; Section 10 treats regulatory and validation constraints specific to medical devices. Every empirical claim attributed to a named study has been checked against the publisher or indexing record; full bibliographic details, including volume, pages and DOI where available, appear in Section 12. Claims not traceable to a verifiable source are marked as practitioner-reported or as the author’s synthesis. |
2 Introduction, Scope and Method
2.1 What operational excellence denotes, and what it does not
The term operational excellence has been used loosely enough that it risks becoming a synonym for “good management”. The research literature has converged on a narrower construct. Carvalho and colleagues, across a sustained programme of work linking operational excellence to organisational culture and agility, treat it as an organisational capability rather than a programme: the enduring capacity to improve performance across cost, quality, delivery and safety simultaneously, sustained through changes in leadership, product and market (Carvalho et al. 2023). The distinguishing property is not the level of performance but the rate and durability of improvement.
That framing has an important corollary for engineers. If operational excellence is a capacity, then the object of design is the improvement system itself — the loops by which deviations are detected, diagnosed, corrected and prevented from recurring — and not merely the production system. Anand, Ward, Tatikonda and Schilling (2009) make this explicit by treating continuous improvement infrastructure as a dynamic capability, and by identifying the infrastructure decision areas (organisational structure for improvement, project selection and governance, training and competence, resource allocation, knowledge capture) that determine whether improvement activity compounds or dissipates.
This review therefore evaluates every method on two axes: what it does to the production system, and what it does to the improvement system. Methods that improve the process once but leave no residual capability are graded accordingly.
2.2 The regulated manufacturing frame
Ophthalmic devices — soft and rigid contact lenses, intraocular lenses, lens care solutions, and the associated moulds, tooling and packaging — sit inside a quality-system perimeter that alters the economics of every method discussed here.
Instrument |
Scope and status |
Operational consequence |
21 CFR Part 820 (QMSR) |
The FDA Quality Management System Regulation took effect 2 February 2026, amending Part 820 to incorporate ISO 13485:2016 by reference in place of the former Quality System Regulation. FDA concurrently moved device inspections to Compliance Program 7382.850, retiring 7382.845 and 7383.001. |
Terminology, records structure and inspection narrative now align with ISO 13485. Firms already conformant to ISO 13485:2016 face mainly documentary and inspectional-readiness change; firms operating a legacy QSR-shaped system face structural change. |
ISO 13485:2016 |
Quality management systems for medical devices; risk-based process approach; design and development controls; process validation where output cannot be fully verified. |
Optical, dimensional and sterility attributes of lenses are only partially verifiable at 100 percent inspection, so the associated processes are validation-controlled. Improvement to those processes triggers revalidation. |
ISO 14971:2019 |
Application of risk management to medical devices across the lifecycle. |
Supplies the formal risk model that should drive improvement project selection. A risk-priority ranking derived from the device risk file is a defensible and auditable substitute for ad hoc project selection. |
EU MDR 2017/745 |
Conformity assessment, clinical evaluation, unique device identification, post-market surveillance. |
Post-market surveillance and vigilance data become a legitimate — and under-exploited — input to the improvement loop, closing the feedback path from field performance to process parameters. |
ISO 11607 / ISO 11137 (where applicable) |
Packaging for terminally sterilised devices; sterilisation by radiation. |
Seal integrity and dose mapping become validated process parameters with their own control strategies and revalidation triggers. |
Table 2.1 Regulatory instruments that shape improvement economics in ophthalmic device manufacture. Statuses stated as of August 2026.
The binding constraint is the change control tax. In an unregulated plant, the cost of a process change is the engineering effort plus the trial material. In a device plant it is that, plus risk assessment, plus impact analysis against the design history file, plus — for validation-controlled processes — installation, operational and performance qualification, plus documentation and, in some cases, notified body or agency notification. McGrane and colleagues, studying the deployment of a Lean Six Sigma project inside a medical device manufacturer, documented a nine-month project cycle attributable in substantial part to validation and regulatory approval activity (McGrane et al. 2022). The improvement rate of a device plant is therefore bounded above by its change-throughput capacity.
Design implication Because validated change is the scarce resource, the highest-leverage methods in a regulated plant are those that (a) increase the information yield per validated change — design of experiments, design space definition, statistical engineering; (b) allow experimentation without physical change — digital twin, discrete-event simulation, in-silico design of experiments; or (c) legitimately reduce the validation burden through risk-based classification of change. Methods that simply generate more improvement ideas add load to an already saturated change-control queue. |
2.3 Review method and evidence grading
This is a narrative review with structured evidence appraisal, not a systematic review in the PRISMA sense; the breadth of the method space precludes a single search protocol. Sources were drawn from operations management, quality engineering, industrial engineering and manufacturing systems literature, with priority given to (i) large-sample empirical studies, (ii) meta-analyses and systematic literature reviews, and (iii) methodological papers in the recognised journals of the field. Every source cited by author and year has been checked against a publisher, repository or indexing record; bibliographic details are given in Section 12.
Each method is assigned an evidence grade using the scheme below. The grade describes the strength of the causal evidence linking the method to operational outcomes — not the method’s usefulness, and not the confidence of its advocates. Several methods that are highly useful carry a low grade because the field has not yet produced controlled evidence for them.
Grade |
Criterion |
Interpretation |
A |
Large-sample empirical study with a control or matched comparison group, or a quantitative meta-analysis, published in a peer-reviewed operations or quality journal. |
Effect direction can be treated as established. Effect magnitude is contingent on context. |
B |
Multiple independent survey-based or multi-case studies with consistent findings, or a systematic literature review synthesising a substantial article corpus. |
Effect direction is well supported; confounding by co-adopted practices is generally not excluded. |
C |
Single case study, action research, or design-science artefact with limited evaluation; or a coherent body of practitioner-reported results without independent control. |
Plausible and often valuable, but susceptible to publication and survivorship bias. Treat reported magnitudes as upper bounds. |
D |
Conceptual, normative or emerging; empirical validation confined to laboratory or simulation settings. |
Adopt as a designed experiment with explicit success criteria, not as a committed programme. |
Table 2.2 Evidence grading scheme used throughout Section 4.
A caution applies throughout. Mabin and Balderstone (2003), assembling the largest quantitative synthesis of Theory of Constraints applications, reported that despite extensive searching they found no published reports of failure. That result is not credible as a statement about the method; it is a direct measurement of publication bias in the improvement literature. Every effect magnitude quoted from case-based sources in this document should be read with that bias in mind, and the reader should assume that the population mean lies below the published mean.
3 A Layered Taxonomy of Operational Excellence Methods
3.1 Why a layer model rather than a list
Methods are usually catalogued alphabetically or by historical origin. Neither ordering helps an engineer decide what to do. A more useful organisation follows the physics of the system being controlled: what is being regulated, at what time constant, and by what feedback mechanism. Under that criterion the method space resolves into five layers with markedly different characteristic times, from strategic cycles measured in quarters down to closed-loop process control measured in milliseconds.
Layer |
Name |
Controlled variable |
Time constant |
Methods (Section 4 reference) |
L0 |
Daily management and problem solving |
Time from abnormality occurrence to detection, and from detection to resolution; retention of gains |
Minutes to days |
Standardised and leader standard work (4.17); skills matrix (4.18); gemba (4.19); kaizen (4.20); PDCA and A3 (4.21); 8D, Shainin and the problem-solving family (4.22) |
L1 |
Strategic direction and capability |
Alignment between corporate objectives and shop-floor improvement effort; rate of capability acquisition |
Quarters to years |
Hoshin Kanri and company-specific production systems (4.1); Toyota Kata (4.2) |
L2 |
Flow and capacity |
Throughput, work-in-process, lead time, due-date performance |
Days to weeks |
Lean production systems (4.3); Theory of Constraints (4.6); Quick Response Manufacturing and POLCA (4.7); reinforcement-learning scheduling (4.13) |
L3 |
Variation and quality engineering |
Mean and variance of critical-to-quality characteristics; defect rate |
Hours to weeks |
Six Sigma and DMAIC (4.4); statistical engineering (4.4.6); design for Six Sigma and design space (4.5); Zero Defect Manufacturing (4.9); machine-learning-augmented SPC (4.10) |
L4 |
Asset reliability and availability |
Availability, performance rate, mean time between failures, mean time to repair |
Hours to months |
Total Productive Maintenance and OEE (4.8); predictive and prescriptive maintenance (4.11) |
L5 |
Digital substrate and inference |
Fidelity, latency and completeness of the process representation on which all other layers depend |
Milliseconds to hours |
Digital twin (4.12); process mining (4.14); Quality 4.0 and LSS4.0 integration frameworks (4.10, 4.15) |
Table 3.1 Six-layer control taxonomy. Layer L0 is the substrate on which every other layer depends; layer L5 is an enabler that changes the observability and controllability of layers L2 to L4 but produces no operational benefit on its own.
3.2 Coupling between layers
The layers are not independent, and the coupling is directional. Three couplings dominate practice.
L0 to everything (detection coupling). No improvement method can act on a problem the organisation has not noticed. Standardised work supplies the baseline against which abnormality is visible; tiered accountability supplies the cadence at which it is escalated; leader standard work supplies the managerial attention to observe it. Where these are absent, the detection interval is measured in shifts or weeks, and the improvement rate is bounded by that interval no matter how sophisticated the analytical layer above it. This is developed in Section 5.6.
L1 to L2–L4 (authority coupling). Improvement at the flow, quality and reliability layers consumes capacity that is otherwise allocated to production. Without a strategy-deployment mechanism that formally reserves that capacity and adjudicates between competing improvement claims, the improvement system is starved by the production system on any day that production is behind. This is the mechanism by which most improvement programmes die, and it is a resource-allocation failure rather than a technical one. Hoshin Kanri and its Western analogues exist to close this loop.
L4 to L2 (availability coupling). Flow methods presuppose predictable capacity. Drum-buffer-rope sizing, POLCA loop parameters and kanban quantities are all functions of process variability, of which unplanned downtime is normally the largest component. Deploying a pull system on an asset base with unstable availability yields a system that starves and floods alternately, and the failure is usually misattributed to the pull system. Shah and Ward (2003) capture the same relationship empirically: their four practice bundles — just-in-time, total quality management, total productive maintenance and human resource management — contribute jointly, and the joint contribution exceeds what any bundle achieves alone.
L5 to L3 (observability coupling). Every advanced quality method — predictive quality, Zero Defect Manufacturing, adaptive process control — is a function of the measurement system that feeds it. In ophthalmic manufacture the measurement system is largely automated optical inspection, which means that the effective resolution of the entire L3 layer is set by the inspection algorithm’s sensitivity and specificity. An automated inspection system with an elevated false-discovery rate does not merely waste yield; it corrupts every downstream statistical model that consumes its output. Measurement system analysis is therefore not a preliminary formality in this sector but the load-bearing element of the whole architecture.
Sequencing rule derived from the coupling structure Establish L0 (standardised work, visual controls, tiered accountability, leader standard work, one working problem-solving method) first. Without it, problems are not detected and improvements are not retained. Establish L1 (direction, protected capacity, capability-building routine) before committing to any L2–L4 programme; otherwise improvement effort is preempted by production pressure. Establish measurement adequacy and basic asset stability (L4) before deploying flow control (L2) or advanced statistical control (L3). Deploy L5 last as an accelerant, never first as a substitute. Instrumenting an unstable, non-standardised process produces a high-fidelity model of chaos. |
3.3 What “advanced” means in this document
Advanced is used here in a specific sense: a method is advanced if it satisfies at least two of four criteria. (i) Mechanistic depth — it acts on the generating structure of the problem rather than on its symptoms; design space definition qualifies, defect sorting does not. (ii) Predictive rather than reactive control — it acts before the deviation reaches the product; predictive maintenance and prediction-strategy ZDM qualify, end-of-line containment does not. (iii) Capability residue — it leaves the organisation more able to solve the next problem; Toyota Kata is the paradigm case. (iv) Formal evidence base — its effect has been measured against a comparison group rather than asserted.
Under this definition several widely marketed practices are excluded. Kaizen events in isolation are excluded: they produce a step change with negligible capability residue and, in a validated environment, an outsized change-control burden per unit of benefit. Generic 5S deployment is excluded on the same grounds, notwithstanding its role as a precondition for visual control. Certification-driven quality management is excluded because compliance and capability are only weakly correlated; a plant can be fully conformant to ISO 13485 and operationally mediocre, and the reverse is also observed.
4 Method Monographs
Each monograph follows a common structure: theoretical basis, mechanism of action, evidence with grade, implementation protocol, a worked example situated in ophthalmic device manufacture, limitations and failure modes, and primary metrics. The examples are constructed illustrations built on documented process characteristics of soft contact lens manufacture — cast moulding, hydration, automated optical inspection, blister packaging and sterilisation — and are intended to be dimensionally realistic rather than to describe any specific facility.
4.1 Hoshin Kanri and Company-Specific Production Systems
At a glance |
|
Layer |
L1 — Strategic direction and capability |
Evidence grade |
B for company-specific production systems as a class; C for Hoshin Kanri specifically |
Primary lever |
Alignment of improvement capacity with strategic objectives; formal reservation of improvement resource |
Time to measurable effect |
2–4 quarters to alignment effects; 4–8 quarters to operational effects |
Regulatory friction (device plant) |
Low. Strategy deployment is not a validated process; it does, however, require documented management review linkage under ISO 13485 clause 5.6. |
4.1.1 Theoretical basis
Hoshin Kanri (方針管理, literally direction management) is a policy-deployment system developed within Japanese total quality control practice. Formally it is a nested plan-do-check-act structure in which a small number of breakthrough objectives are decomposed through the organisational hierarchy by a negotiation protocol — catchball — that requires each level to propose the means by which it will contribute to the level above, and requires the level above to accept or modify those means. Two properties distinguish it from conventional cascaded objective-setting: the decomposition is of means as well as ends, and the review cadence is monthly rather than annual, which makes it a control loop rather than a planning artefact.
The related construct of the company-specific production system (XPS) — the Volvo Production System, the Bosch Production System, and their many analogues — is the institutional carrier of the same idea. Netland (2013) analysed thirty such systems and framed the central question as “one best way or own best way”: whether firms should adopt a canonical model or construct a bespoke one. The empirical answer is that the content of these systems converges strongly on lean principles while the packaging — vocabulary, structure, sequencing, governance — is deliberately idiosyncratic, because the packaging is what carries organisational identity and therefore commitment.
4.1.2 Mechanism
The mechanism is resource protection under competing claims. Improvement work and production work draw on the same engineers, technicians and equipment time. Absent a formal allocation mechanism, the claim with the shorter feedback loop wins, and production always has the shorter loop. Hoshin Kanri converts improvement capacity from a residual into a committed quantity by making it an explicit object of the annual plan, and the monthly review makes underdelivery visible before the year is lost.
A second mechanism is selectivity. The discipline of limiting breakthrough objectives to a small number — conventionally three to five at corporate level — is a constraint on managerial appetite, and it is the constraint that does most of the work. Organisations running twenty simultaneous improvement priorities are, operationally, running none.
4.1.3 Evidence
Direct controlled evidence for Hoshin Kanri as an isolated intervention is thin; the construct is difficult to separate from general management quality. The stronger evidence is indirect and comes from two directions. Netland (2016), surveying 432 practitioners across 83 factories in two multinational corporations, found that management support and commitment ranked among the highest-rated critical success factors for lean implementation, and — importantly for generalisation — that the ranking was largely invariant across corporation, factory size, implementation stage and national culture. Anand and colleagues (2009), studying continuous improvement initiatives across five companies, identified project selection and governance, resource allocation and organisational structure for improvement as infrastructure decision areas that determine whether continuous improvement functions as a dynamic capability. Both findings identify the deployment layer, not the tool layer, as the locus of success.
Recent work extends Hoshin Kanri to non-financial objective classes. Roche and colleagues (2025) present a case study of corporate sustainability targets deployed through Hoshin Kanri in a medium-sized firm, demonstrating that the mechanism generalises beyond cost and quality objectives — relevant to device firms now deploying post-market surveillance and environmental objectives through the same channel.
4.1.4 Implementation protocol
Define three to five breakthrough objectives with quantified targets and explicit baselines. Reject any objective that cannot be stated as a measurable delta from a known starting value.
For each objective, construct the means-ends tree down to the level at which an owner can commit to a specific countermeasure with a specific completion date. Stop the decomposition when the leaf is an action, not an aspiration.
Run catchball vertically and horizontally. Each level proposes its means; the level above tests sufficiency (do the proposed means, if fully delivered, close the gap?) and feasibility (is the required capacity available?). Insufficiency at this stage is information, not failure; it is the point at which targets or resources must move.
Publish the X-matrix or equivalent artefact linking objectives, strategies, actions, metrics and owners. In a device plant, cross-reference each improvement action to the affected quality-system process and to its change-control classification at this stage, not later.
Reserve improvement capacity explicitly: named people, named hours per week, protected from production escalation except by a documented exception decision at plant-manager level.
Review monthly against leading indicators with a standard problem-solving format for gaps. Reviewing outcome metrics alone converts the system into reporting.
Conduct an annual reflection (hansei) that assesses the process of deployment as well as the results, and carries forward the diagnosis into the next cycle.
4.1.5 Worked example — ophthalmic
A lens manufacturer sets a breakthrough objective of reducing total cost of poor quality from 6.1 percent to 3.5 percent of cost of goods sold within eighteen months. Decomposition identifies three principal contributors from the quality cost ledger: optical-zone cosmetic rejects at automated optical inspection (roughly 48 percent of scrap value), hydration-stage dimensional drift causing base-curve non-conformance (roughly 22 percent), and blister seal-integrity failures at package inspection (roughly 14 percent). Each becomes a second-level objective with a distinct owner and a distinct method: the cosmetic reject stream is routed to a design-of-experiments programme on mould release and monomer dosing; the dimensional drift stream to a statistical engineering study of hydration bath thermal and ionic profiles; the seal stream to a Total Productive Maintenance intervention on the sealing head, because failure analysis attributes the majority of events to platen wear and temperature drift rather than to material.
The catchball round produces the decisive finding. The engineering capacity required by the three programmes exceeds the available process-engineering establishment by an estimated 1.7 full-time equivalents once validation documentation load is included. Under a conventional cascade this would surface at month nine as slippage. Under catchball it surfaces before commitment, and the organisation makes a real choice: defer the seal-integrity programme by two quarters, or fund contract validation support. Making that choice consciously, in month zero, is the entire value of the mechanism.
4.1.6 Limitations and failure modes
Ritualisation. The X-matrix becomes an annual artefact rather than a live control document. Diagnostic: ask when the matrix was last modified mid-year. If the answer is never, the loop is open.
Target inflation. Objectives set by negotiation-from-above rather than by capability analysis produce commitments that are known to be undeliverable at the moment they are signed, which trains the organisation to treat all targets as notional.
Metric substitution. Where the breakthrough metric is hard to move, activity metrics (projects launched, people trained) migrate into the review. This is the most common quiet failure.
Weak coupling to the quality system. In device plants, an improvement plan that is not reconciled with change-control capacity is a plan to generate a validation backlog.
4.1.7 Primary metrics
Percentage of breakthrough objectives with a green leading indicator at each monthly review; improvement capacity delivered against improvement capacity committed (hours); count of active breakthrough objectives (a control metric — increases signal loss of selectivity); mean age of open countermeasures; percentage of improvement actions with a completed change-control classification at plan approval.
4.2 Toyota Kata: Improvement Kata and Coaching Kata
At a glance |
|
Layer |
L1 — Strategic direction and capability |
Evidence grade |
C — action research and multi-case evidence; no large-sample controlled study |
Primary lever |
Deliberate practice of scientific problem solving; conversion of improvement from event to routine |
Time to measurable effect |
2–3 quarters to behavioural change; 4–6 quarters to operational effect |
Regulatory friction (device plant) |
Low to moderate. The experimental cycle must be conducted within change control; trivial experiments on validated processes are not trivial to execute. |
4.2.1 Theoretical basis
Rother (2010) advanced the proposition that the durable element of the Toyota Production System is not its tools but a pair of behavioural routines: the Improvement Kata, a four-step pattern by which an individual works from a current condition toward a target condition through short experimental cycles, and the Coaching Kata, a structured dialogue by which a supervisor develops that pattern in a subordinate. The theoretical claim is that scientific problem-solving behaviour is a skill acquired through deliberate practice under correction, in the same sense as a musical or athletic skill, and that it therefore cannot be transmitted by training courses.
The four steps are: understand the direction or challenge; grasp the current condition through direct observation and measurement; establish the next target condition, defined as a description of a desired process state with a date; and iterate toward it through single-factor experiments, each with an explicit prediction. The distinguishing feature relative to conventional plan-do-check-act is the target condition — an intermediate, process-level state rather than an outcome goal — and the insistence on stating the expected result before running the experiment, which makes learning detectable.
4.2.2 Mechanism
Two mechanisms operate. The first is epistemic: requiring a prediction before each experiment converts the threshold-of-knowledge boundary into an observable. When the prediction is right, the mental model is confirmed; when it is wrong, the discrepancy is the information. Conventional improvement work, which reports only outcomes, discards this signal entirely.
The second is organisational. The Coaching Kata places the manager in the role of developing the subordinate’s reasoning rather than supplying the answer. Where it is practised consistently, it changes what managers do with their time, and it is that reallocation — rather than any individual improvement — that produces the compounding effect.
4.2.3 Evidence
The evidence base is weaker than the method’s prominence suggests, and this should be stated plainly. Published work is dominated by action research and bibliographic analysis rather than controlled comparison. Tillema and van der Steen (2015) and subsequent bibliographic analyses identify the availability of experienced coaches as the dominant barrier, and note that organisations without prior exposure to systematic problem-solving routines struggle to bootstrap the practice. Action research in construction and other non-automotive settings reports successful transfer of the routine but rarely isolates its operational contribution from co-occurring interventions. Grade C is appropriate: the mechanism is theoretically well motivated and consistently reported to work by those who use it, but the counterfactual has not been established.
4.2.4 Implementation protocol
Select one value stream and one challenge with a two-to-three-year horizon, stated in process terms (for example, ability to produce any lens power in any batch without changeover penalty).
Train a small number of coaches — typically two to four for a first wave — through practice with an experienced second coach, not through classroom instruction. This is the rate-limiting step and it cannot be compressed.
For each learner, establish the current condition quantitatively: process capability, cycle time distribution, defect Pareto by mode, not by opinion.
Set the first target condition with a horizon of two to four weeks. Target conditions of longer horizon degrade into projects and lose the experimental character.
Run daily or near-daily coaching cycles of ten to twenty minutes using the five coaching questions. Frequency matters more than duration; weekly coaching does not build the routine.
Record every experiment with its prediction, its result and what was learned. The experiment record is the primary artefact; in a regulated plant it also constitutes useful development evidence when the resulting change enters change control.
Expand only when the first coaches can coach reliably. Premature scaling produces a vocabulary without a practice, which inoculates the organisation against the real thing.
4.2.5 Worked example — ophthalmic
The challenge for a lens moulding cell is stated as: produce every scheduled stock-keeping unit every day within a two-shift window (an every-part-every-interval challenge). The current condition, measured over six weeks, shows an EPEI of 4.2 days, driven by a mean mould changeover of 96 minutes with a standard deviation of 31 minutes, and a post-changeover qualification run averaging 41 minutes before first-article approval.
The first target condition is set at: mould changeover completed in under 60 minutes with a standard deviation under 10 minutes, by a date four weeks out. The learner runs a sequence of single-factor experiments — external preparation of the mould set, a dedicated tooling cart, a modified clamping sequence, a pre-heated tool. The fourth experiment fails against prediction: pre-heating reduces changeover time but increases first-article rejects, because the thermal profile at start-up now differs from the validated steady state. That failure is the most valuable result in the sequence, because it identifies a coupling between the changeover process and the validated process window that a conventional setup-reduction workshop would have discovered only after the change had been implemented and a deviation raised.
The experiment record then becomes an input to the change-control package: the organisation can demonstrate, with data, why the pre-heat parameter was bounded where it was bounded.
4.2.6 Limitations and failure modes
Coach scarcity. The practice is transmitted person to person. An organisation cannot buy its way past this constraint, and attempts to do so through large training contracts reliably produce vocabulary adoption without behavioural adoption.
Incompatibility with command culture. Where managers are rewarded for having answers, the coaching dialogue is experienced as an evasion and reverts to instruction within weeks.
Experiment friction in validated processes. On a validated process, even a small parameter experiment may require a protocol. The workable response is to conduct the experimental phase on a qualification line or on non-product material, and to enter change control only with the converged answer — but this must be designed in from the start.
Metric ambiguity. Because the output is capability, the method resists conventional return-on-investment justification, which makes it politically fragile during cost pressure.
4.2.7 Primary metrics
Experiments per learner per week; proportion of experiments with a documented prior prediction; proportion of predictions falsified (a health metric — a rate near zero indicates the learner is not working at the knowledge threshold); target conditions achieved on date; coaching cycles conducted against coaching cycles scheduled; number of qualified coaches.
4.3 Lean Production Systems: Practice Bundles and the Hard–Soft Distinction
At a glance |
|
Layer |
L2 — Flow and capacity (with L3/L4 components) |
Evidence grade |
A — large-sample empirical evidence with quantified variance explained |
Primary lever |
Waste elimination, flow, pull, levelling; supported by problem-solving capability and human-resource practices |
Time to measurable effect |
2–4 quarters for local effects; 6–12 quarters for system effects |
Regulatory friction (device plant) |
Moderate. Layout, flow and pull changes are generally not validated processes; changes to process parameters and inspection points are. |
4.3.1 Theoretical basis
Lean production is best treated not as a philosophy but as a measurable multidimensional construct. Shah and Ward (2003) operationalised it as twenty-two practices grouped into four internally consistent bundles — just-in-time, total quality management, total preventive maintenance and human resource management — and tested both the contextual determinants of adoption and the performance consequences. Shah and Ward (2007) subsequently developed and validated a ten-factor measurement model spanning supplier, customer and internal dimensions, which remains the reference instrument for empirical lean research.
The theoretical content is a queueing-theoretic proposition: variability, in demand, in process time and in availability, must be buffered by inventory, capacity or time. Lean methods reduce the variability so that the buffer requirement falls, rather than merely removing buffers, which is the characteristic failure of superficial implementations.
4.3.2 Mechanism
Setup reduction lowers the economic batch quantity, which reduces cycle inventory and shortens the every-part-every-interval; pull control caps work-in-process, which by Little’s law caps lead time for a given throughput; levelling reduces demand variability seen by upstream stages; jidoka and source inspection reduce the time between defect creation and defect detection, which reduces the quantity at risk per event. Each mechanism is independently derivable, which is why the bundle behaves better than any element alone.
4.3.3 Evidence
Grade A. Shah and Ward (2003) report that lean bundles account for approximately 23 percent of the variance in operational performance after controlling for industry and contextual factors — a substantial figure for a cross-sectional operations study, and one that should also be read as a statement about the 77 percent that lean practices do not explain. The same study finds plant size to be a robust predictor of lean adoption, while plant age and unionisation have weaker effects than conventional accounts assume.
The most consequential refinement comes from Bortolotti, Boscari and Danese (2015). Comparing plants classified as successful and unsuccessful lean implementers, they find that hard lean practices function as order qualifiers — adopted at broadly similar levels by both groups — whereas soft practices (small-group problem solving, training, supplier and customer involvement, continuous-improvement leadership) discriminate the successful group, and that successful plants exhibit a distinct organisational-culture profile. The managerial implication is uncomfortable and worth stating directly: a plant that has implemented the technical apparatus of lean and achieved nothing has probably implemented the qualifier and omitted the winner.
Netland (2016) supplies the deployment-side complement, with management commitment and capability development ranking highest among critical success factors across a 432-respondent, 83-factory sample.
4.3.4 Implementation protocol
Establish the measurement baseline: value stream map with process time, changeover time, uptime, first-pass yield, batch size and inventory at each stage, measured rather than estimated.
Stabilise before flowing. Where availability at any stage is below roughly 85 percent or first-pass yield is materially below the design intent, address those first; flow control on an unstable base produces oscillation.
Reduce changeover on the constraint and on any stage whose batch size is set by changeover economics. Apply the internal-to-external conversion sequence, then improve the external elements.
Cap work-in-process with an explicit pull mechanism sized from measured variability, not from a textbook formula. Recompute the sizing quarterly.
Level the schedule to the extent demand and changeover economics allow, and make the levelling pattern visible.
Install source inspection and stop-the-line authority where the detection lag is long and the quantity at risk per event is large.
Deploy the soft bundle deliberately and concurrently: structured small-group problem solving with protected time, competence development against a skills matrix, and supervisory behaviour standards. This is the element most often deferred and it is the element the evidence identifies as decisive.
4.3.5 Worked example — ophthalmic
A cast-moulding line runs sixteen lens powers in campaign batches of 40,000 units. Changeover between powers takes 96 minutes; the resulting EPEI of 4.2 days requires roughly 11 days of finished-goods cover to maintain a 98 percent service level, and the finished-goods pool for the family carries approximately 1.9 million units.
Sequenced intervention: (i) measurement system analysis on the automated optical inspection station, which establishes that a portion of the apparent cosmetic reject rate is measurement variation rather than product variation — a result that changes the improvement target before any process work begins; (ii) availability recovery on the moulding presses from 81 to 92 percent through the Total Productive Maintenance intervention described in 4.8; (iii) changeover reduction from 96 to 54 minutes, which halves the economic batch and brings EPEI to 2.1 days; (iv) a two-bin pull signal between moulding and hydration with work-in-process capped at 1.4 days; (v) small-group problem solving on the top three cosmetic defect modes with protected weekly time.
The lead-time and inventory effects follow from Little’s law and from the batch-size reduction. The yield effect follows from step (i) and step (v). The instructive point is the ordering: had the pull cap been installed first, at 81 percent availability, the line would have starved hydration repeatedly, and the pull system — rather than the maintenance deficit — would have been blamed.
4.3.6 Limitations and failure modes
Buffer removal without variability reduction. The most common and most damaging error. Inventory is a symptom; removing it without removing its cause converts an inventory problem into a service problem.
Tool-first deployment. Implementing the hard bundle alone reproduces the pattern Bortolotti and colleagues associate with unsuccessful implementers.
Contextual over-extrapolation. The canonical lean model derives from high-volume repetitive discrete assembly. High-mix low-volume environments are better served by the Quick Response Manufacturing variant described in 4.7.
Regulated-environment friction. Reductions in inspection or in in-process controls that appear as waste elimination may be validated controls. Every such change must be assessed against the device risk file before it is proposed, not after.
4.3.7 Primary metrics
Dock-to-dock lead time; work-in-process turns; every-part-every-interval; changeover time (mean and standard deviation); first-pass yield by stage; schedule adherence; suggestions implemented per employee per year; hours of protected problem-solving time delivered against planned.
Talk to us →