Planned Maintenance System and fleet reliability
Module objectiveDistinguish the four maintenance strategies and choose the one suited to each component by criticality, condition data and cost.
Every onboard component can be managed with a different maintenance strategy: there is no universally best approach, but a choice that depends on the component's criticality, the availability of condition data and the relative cost of the different options.

| Strategy | Principle | When to use it |
|---|---|---|
| Reactive (run-to-failure) | Intervention only on failure | Non-critical, redundant components with low replacement cost |
| Preventive (time-based) | Fixed intervals from the PMS schedule | Onboard standard, widely accepted by class |
| Condition-based (CBM) | Intervention triggered by indicators (vibration, oil, thermography) | Critical machinery monitorable with non-invasive techniques |
| Predictive (analytics) | Models and trends estimating remaining life | Fleets with widespread sensors and a data platform |
Table 1.1 — Comparison of the four maintenance strategies.
Most real fleets use a mix of the four strategies, applied component by component. The most common mistake is not choosing the wrong strategy for a single component, but applying the same strategy indiscriminately to the whole system, without differentiating by criticality.
Understanding how a component's failure rate varies over time is essential for planning genuinely effective maintenance, instead of applying fixed intervals that ignore the life stage the component is in.
| Phase | Characteristics and maintenance implications |
|---|---|
| Infant mortality | High but decreasing failure rate, typical of new or recently overhauled components; requires careful commissioning and close monitoring in the first hours |
| Useful life | Low, relatively constant failure rate, dominated by random events; time-based preventive maintenance has diminishing returns in this phase |
| Wear-out / end of life | Increasing failure rate due to wear phenomena; here preventive or predictive maintenance has the greatest value, anticipating replacement before failure |
Table 2.1 — The three phases of the bathtub curve and their implications.
Very few. It is the finding that changed how maintenance is reasoned about, and it comes from the study Nowlan and Heap carried out for United Airlines in 1978, analysing failure data from a real fleet. Six patterns came out of it, not one.

| Pattern | Failure rate behaviour | Share |
|---|---|---|
| A — bathtub | Infant mortality, then a constant rate, then rising wear-out | 4% |
| B — wear-out | Constant or slowly rising, with a clear final wear-out zone | 2% |
| C — fatigue | Gradual, continuous increase with no identifiable wear-out zone | 5% |
| D — break-in | Low at first, then rapidly constant | 7% |
| E — random | Constant throughout the component’s life | 14% |
| F — infant mortality | High at first, then constant or slightly decreasing | 68% |
Table 2.2 — The six failure patterns of Nowlan and Heap (1978) and the share of components in each.
Only patterns A, B and C are age-related, and together they account for 11% of components. The other three — 89% — show no relationship between age and probability of failure. Two consequences follow, and they carry the whole course. First: for the large majority of components a fixed replacement interval does not reduce the probability of failure, because there is no age at which failure becomes more likely. Second, less intuitively: the dominant pattern is infant mortality, which means every opening-up and every replacement reintroduces risk. Stripping down a healthy component to meet the calendar is not neutral: it puts it back at the start of the curve.
Electronic and electrical components tend towards a constant failure rate for most of their useful life (patterns E and F); mechanical components subject to wear — bearings, seals, liners — are among the few that genuinely follow A or B. That is why the fixed interval remains right exactly where the Module 01 table places it, and becomes waste elsewhere: the choice is not made system by system, it is made failure mode by failure mode.
Reliability Centered Maintenance is a structured methodology for deciding, component by component, which maintenance strategy to apply, starting from the analysis of functions, failure modes and their consequences.

«RCM» is a label used loosely. There is, however, a standard setting the minimum criteria for a process to be called that: SAE JA1011 — Evaluation Criteria for Reliability-Centered Maintenance Processes, first issued in 1999 and revised in August 2009, with the companion guide SAE JA1012. A process is RCM if, and only if, it answers seven questions in order.
| # | Question | What it produces |
|---|---|---|
| 1 | What is the operational context, and what are the asset's functions and associated performance standards? | A definition of what the component must do, and how well |
| 2 | In what ways can it fail to fulfil those functions? | Functional failures |
| 3 | What causes each functional failure? | Failure modes |
| 4 | What happens when each failure occurs? | Failure effects |
| 5 | What are the consequences of the failure? | The classification: safety, environment, operations, cost |
| 6 | What should be done to predict or prevent each failure? | Maintenance tasks and their intervals |
| 7 | If no task is effective, what other strategy is preferable? | One-time changes: redesign, redundancy, accepting the risk |
Table 3.1 — The seven questions of SAE JA1011.
The structure itself says something. Only the sixth question reaches maintenance tasks: the five before it establish what the component must do, how it can stop doing it and what it costs when it stops. An analysis that starts from the list of existing jobs and tries to rationalise them is doing something else — possibly useful, but not RCM under JA1011. And the seventh question is the one most often skipped: when no maintenance task is effective, the right answer may be to redesign, to add redundancy, or to knowingly accept the failure, not to invent an interval.
Applying RCM to an entire system is a significant initial investment of time and expertise. The return shows over time: a maintenance programme built on RCM tends to reduce both unforeseen failures and unnecessary maintenance interventions, optimising resources instead of simply applying them uniformly.
The PMS is the structured record of scheduled maintenance work and its related records. It is also the tool through which many Administrations and classes accept maintenance schemes as an alternative to certain traditional surveys (approved Planned Maintenance Scheme).
| Component | Function |
|---|---|
| Machinery register | Structured list of all systems and components subject to maintenance |
| Job list and intervals | Definition of maintenance activities and their frequency |
| Recording of interventions | History of every intervention carried out, with date, performer, outcome |
| Deviation (overdue) management | Monitoring of jobs not carried out within the scheduled deadline |
| Integration with spares management | Link between maintenance jobs and availability of required spares |
Table 4.1 — Essential components of a well-governed PMS.
A well-designed PMS structures the work, but does not eliminate the need for technical judgement: a Chief Engineer who blindly follows PMS intervals without integrating direct observation of the system's actual condition loses valuable information no programme can capture on its own.
Module objectiveRecognise what ISM Code section 10 actually requires for critical equipment, and decide which spares to keep on board.
The critical equipment list is not good practice: it is a requirement of ISM Code section 10, and its text is worth reading, because what it asks is more precise — and different — from how it is usually summarised.
«The Company should identify equipment and technical systems the sudden operational failure of which may result in hazardous situations. The SMS should provide for specific measures aimed at promoting the reliability of such equipment or systems. These measures should include the regular testing of stand-by arrangements and equipment or technical systems that are not in continuous use.» (10.3)
«The inspections mentioned in 10.2 as well as the measures referred to in 10.3 should be integrated into the ship's operational maintenance routine.» (10.4)ISM Code, Part A, section 10 — Maintenance of the ship and equipment
This is the nuance summaries lose. Section 10.3 does not require installing a second generator or a second pump: it requires that, where they exist, stand-by arrangements and equipment that does not run continuously be regularly tested. And it bites on exactly the category of equipment where reliability is most easily imagined: the kind that looks fine because it is never started. Section 10.4 closes the loop by preventing this from becoming a separate exercise — the stand-by test belongs inside the PMS, not in a register of its own.

An effective spare parts stocking policy does not simply follow the manufacturer's generic recommendations, but cross-references operational criticality, real procurement times (lead time) and the availability of alternative suppliers along the ship's typical routes.
The true cost of a missing critical spare is machinery downtime, any resulting off-hire, and in the most serious cases a safety risk. Assessing stock only on the purchase cost of the part, ignoring the cost of unavailability, almost always leads to insufficient stock for the most critical components.
ISM Code 10.3 says that the list must exist. It does not say how to build it, and this is where most companies proceed by custom: the list is inherited from the sister ship, or from the previous manager, or from the PMS software template. A method does exist, it comes from failure mode analysis, and it is the same one underpinning the RCM of Module 03 and — as we shall see — the class-approved schemes.
FMEA (Failure Mode and Effects Analysis) walks through a system component by component and asks, for each, in what ways it can fail and what happens when it does. FMECA adds the letter that matters for our purpose — the C for Criticality — the assessment of how much that failure weighs. It is the same logical chain as questions 2, 3, 4 and 5 of SAE JA1011, applied in a tabular format.
| Column | What goes in it |
|---|---|
| Component and function | What it does, and to what performance standard |
| Failure mode | How it stops doing it: fails to start, fails to stop, leaks, sticks in position |
| Cause | What produces that failure mode |
| Local and system effect | What happens to the component, and what happens to the ship |
| Severity (S) | How serious the worst reasonable consequence is |
| Occurrence (O) | How likely that cause is to arise |
| Detection (D) | How likely it is to be noticed before the effect appears |
| Existing control and action | What already prevents or detects it, and what must be added |
Table 6.1 — The columns of a FMECA.
For decades the practice was to compute the Risk Priority Number — RPN = S × O × D — and rank actions by descending RPN. The problem is arithmetical, and it has safety consequences. A failure mode with S 9, O 2, D 3 gives an RPN of 54; one with S 3, O 6, D 6 gives 108. Ranking by RPN means acting first on the second, which is an inconvenience, and later on the first, which is a hazard to people. The defect is structural: ordinal scales are being multiplied, and the product does not preserve which factor was high. That is why the 2019 AIAG-VDA FMEA handbook replaced RPN with Action Priority, a table mapping each S-O-D combination to high, medium or low priority — in which a severity of 9 or 10 yields high priority regardless of how rare or how detectable the failure is.
The coincidence is worth noticing, because it simplifies the work. The Code's criterion — equipment «the sudden operational failure of which may result in hazardous situations» — is a severity-only criterion: it does not ask how likely, nor how detectable. It is exactly the Action Priority logic applied to the high-severity row. Anyone building the critical list by ranking on RPN therefore risks producing a list that does not satisfy 10.3, because it excludes precisely the rare, severe failures the Code wants included.
«Critical» in a shipping company means at least three things, with three different lists and three different readers. Keeping them apart — and knowing where they overlap — avoids both inflated lists and omissions.
| Criticality | The criterion | The consequence |
|---|---|---|
| For safety | ISM 10.3: sudden failure may result in hazardous situations | Specific reliability measures and regular testing of stand-by arrangements, integrated into the PMS (10.4) |
| For operations | The failure stops the ship or the cargo: risk of off-hire, penalties, delay | Spares stocking policy, commercial redundancy, service contracts |
| For class | The item falls within the machinery survey cycle, or within the approved PMS or CBM scheme | Openings, records and audits per the applicable regime (Module 11) |
Table 6.2 — The three criticalities, their criteria and what each entails.
An emergency generator is critical for safety and for class, but its unavailability produces no off-hire. A deck crane on a general cargo ship is critical for operations and for class, but its failure does not in itself create a hazardous situation. A boiler alarm system may be safety-critical without appearing in any survey cycle. Building a single «critical equipment» list and using it for all three purposes always produces the same result: over-stocking items that do not deserve it, and under-testing items that 10.3 would require to be tested.

A list that generates no action is a formality. Each of the three criticalities produces a different output, and these are the outputs an internal audit should verify.
| Output | Where it ends up |
|---|---|
| Periodic testing of stand-by equipment and of equipment not in continuous use | A PMS job, with an interval and a record — not a separate reminder (10.4) |
| Opening and recording regime | The applicable survey cycle: ordinary, approved PMS or approved CBM |
| Stock level and acceptable lead time | The spares policy of the previous module, which at this point has a criterion instead of a custom |
| Monitoring parameters and baseline values | The data base needed to ask class for an approved CBM scheme |
Table 6.3 — What the list must produce, if it is not to remain a list.
The commonest defect is not having the wrong list: it is having written it once and never touched it again. Every plant modification, every conversion, every change of trade shifts what is critical — and that is precisely the management of change seen in the ISM Code course. A critical equipment list identical to the delivery one, on a fifteen-year-old ship that has changed trade twice, does not describe that ship.
Module objectiveDistinguish MTBF, MTTF, MTTR and availability and use them to compare performance over time and between different ships in the fleet.
Systematically measuring reliability makes it possible to compare performance over time and between different ships in the fleet, and to identify signs of deterioration early, before they turn into a serious failure.

| Metric | Definition | How it is calculated |
|---|---|---|
| MTBF Mean Time Between Failures | Average time between successive failures of a repairable component | Operating hours ÷ number of failures in the period |
| MTTF Mean Time To Failure | Average time to failure of a non-repairable component, one that is simply replaced | Operating hours ÷ number of units failed |
| MTTR Mean Time To Repair | Average time needed to return the component to service after a failure | Downtime for repair ÷ number of failures |
| Availability | Proportion of time the component is ready for use | A = MTBF ÷ (MTBF + MTTR) |
| Maintenance backlog | Maintenance jobs open beyond their scheduled deadline | A count, to be read by criticality and by the age of the overrun |
Table 7.1 — Main reliability metrics and how they are calculated.
The distinction between the first two rows is not academic: MTBF applies to what is repaired and returned to service, MTTF to what is replaced and does not come back. Using one for the other makes comparisons between ships meaningless. As for availability, the formula shows something dashboards hide: it improves either by lengthening the time between failures or by shortening the repair. On many shipboard systems the second lever is the faster one, and it depends almost entirely on something Module 05 treats separately — having the right spare on board.
An improving MTBF is a good sign, but should always be read together with the maintenance backlog and critical spares availability: an MTBF that improves because scheduled interventions are simply being postponed tells a very different story from a genuine improvement in reliability.
Condition-based maintenance techniques allow intervention based on the component's actual condition, rather than a predefined time interval, reducing both unforeseen failures and unnecessary interventions.
| Technique | What it detects | Typical application |
|---|---|---|
| Vibration analysis | Imbalances, misalignments, bearing wear | Engines, pumps, compressors, generators |
| Lubricating oil analysis | Metal particles, contamination, additive degradation | Main engine, reduction gears |
| Infrared thermography | Abnormal hot spots, faulty electrical connections | Switchboards, bearings, seals |
| Ultrasonics | Air/gas leaks, defects in slow-rotating bearings | Pneumatic systems, valves |
Table 8.1 — Main Condition Based Maintenance techniques.
A vibration analysis tool in the hands of an officer not trained to interpret its results produces data nobody uses. Investment in CBM technology only makes sense if accompanied by an equivalent investment in the crew's ability to correctly interpret the signals collected.
Module objectiveRecognise what sets predictive maintenance apart and which conditions are needed to adopt it without skipping the intermediate stages.
Predictive maintenance represents the most advanced evolution of maintenance strategies: rather than simply detecting an abnormal condition, it uses statistical or machine-learning models to estimate a component's remaining life and plan the intervention at the optimal moment.
Companies that try to jump directly from a paper PMS to predictive maintenance, without the intermediate stages of digitalisation and condition monitoring, tend to fail for lack of reliable historical data on which to build the models.
Progressing towards predictive maintenance realistically requires passing through several stages of digital maturity, with no shortcuts.

| Level | Characteristics |
|---|---|
| Paper records | Manual management, high risk of information loss, no systematic historical analysis |
| Basic digital PMS | Digitalised checklists and records, but analysis still mostly manual |
| Basic sensors (CBM) | Condition monitoring on selected components, data not yet centralised |
| Integrated fleet data platform | Data from multiple ships centralised and comparable, basis for comparative analysis |
| Predictive maintenance (AI/ML) | Predictive models fed by reliable, widespread historical data |
Table 10.1 — Levels of maintenance digital maturity.
The most common bottleneck is not the sophistication of the predictive algorithm, but the quality and consistency of the data collected over previous years. Investing in recording discipline today is the enabling condition for any future predictive ambition.
Module objectiveDistinguish the three machinery survey regimes — UR Z18, UR Z20, UR Z27 — and the conditions class sets for CBM.
Class recognition of maintenance is neither a discretionary concession nor a practice that varies from Society to Society: it is the subject of IACS Unified Requirements, rules the association's members apply uniformly. There are three regimes, and they sit one on top of the other: you do not choose between them, you climb.

| Ordinary regime | Approved PMS | Approved CBM | |
|---|---|---|---|
| Reference | IACS UR Z18 Survey of Machinery | IACS UR Z20 Planned Maintenance Scheme for Machinery, Req. 2001/Rev.2 2019 | IACS UR Z27 Condition Monitoring and Condition Based Maintenance, Req. 2018 |
| Who opens the machine | Opened in the surveyor's presence, per the survey cycle | The company opens it, on its own programme; the surveyor is not present | Opened when monitoring detects an abnormality, not on a calendar |
| What it replaces | — | Continuous Machinery Survey for the items covered | The periodic openings the PMS provides for, on the items covered |
| Prerequisite | None | Documentary approval by the Society | A ship already on an approved PMS |
Table 11.1 — The three machinery survey regimes and how they relate.
This is the most commonly misunderstood point, and it has a planning consequence. UR Z27 applies only to vessels already operating on an approved PMS survey scheme. There is no path leading from the ordinary regime straight to recognition of condition-based maintenance: first you demonstrate to the Society that you can manage and document planned maintenance, then you may ask for part of those openings to be replaced by monitoring. Items not covered by the CBM scheme remain under Z18 and Z20 — the three regimes coexist on the same ship, component by component.
UR Z27 is the most recent and the most demanding regime, and its conditions are worth seeing because they show what class considers condition-based maintenance proper, as distinct from having a few sensors fitted.
| Condition | Content |
|---|---|
| Approval of the scheme | The Society approves the CBM scheme and its extent, that is which components fall within it — it is not a generic authorisation |
| Approval of the equipment | The monitoring system and instruments must be approved too, not only the procedure |
| Responsible person on board | The chief engineer: responsibility for monitoring and condition-based maintenance cannot be delegated to an outside supplier |
| Documentation to be submitted | Seven items, including the equipment list, the acceptable parameters, the description of the scheme, the instrument specifications, the baseline data and personnel qualifications |
| Documentation to be kept on board | Maintenance instructions, monitoring data, calibration records, maintenance and repair history |
| Audit | Annual, by a Society surveyor, concurrently with the class annual survey |
Table 11.2 — UR Z27 conditions for an approved condition-based maintenance scheme.
Among the seven documents to be submitted there is one that decides whether the whole exercise is feasible: the acceptable parameters and the baseline data for each component. It is not enough to declare that vibration will be measured: you have to know in advance what value is normal on that machine, and what value triggers an opening. That is the technical reason the maturity ladder of Module 10 cannot be skipped — without a reliable history there is no baseline, and without a baseline there is no approvable scheme.
Neither recognition is a one-off acquired right. The annual audit does not check that the system exists, it checks that it works: consistent records, qualified personnel, overruns managed. A deterioration in management quality found at audit can return the ship to the previous regime — and the practical consequence is that the items concerned go back to being opened in the surveyor's presence, with the costs and downtime that follow.
No maintenance system, however well designed, works without the people who carry it out. The real quality of maintenance depends largely on the crew's competence, motivation and organisational culture.
The most effective companies invest in ongoing training, recognise and reward the quality of maintenance execution (not just its speed), and maintain open communication channels between crew and technical office to share observations the PMS alone does not capture.
A PMS log with no delays can reflect excellent management, or a formal completion that does not correspond to actual practice on board. Only direct verification, through inspections and audits, allows the two situations to be told apart.
Module objectiveRecognise the forces that will keep fleet reliability management evolving: digitalisation, decarbonisation, integration between ship and shore.
Fleet maintenance and reliability management will continue to evolve driven by digitalisation, decarbonisation and growing integration between onboard data and shore-based management platforms.
A well-maintained system is not just more reliable: it consumes less fuel for the same performance, with a direct impact on the ship's CII rating. Reliability management and decarbonisation management, seen separately in their respective courses, are in practice increasingly the same competence applied from two angles.
From the Mistake Library of SuperbaKnowledge, filtered to the subjects this course covers. This view selects and organises content published in SuperbaKnowledge; it does not modify or replace it. The linked Knowledge page remains the reference version, while official texts remain authoritative.
| Topic | Mistake | Typical consequence | Topic sheet |
|---|---|---|---|
| IoT Predictive Maintenance | Sensors installed but data collected without systematic trend analysis | Impending failure not anticipated despite the instrumentation being available | See the topic sheet |
| Planned Maintenance System | Standby equipment not tested because it is 'not in use' | A failure of the primary unit reveals that the standby unit doesn't work either | See the topic sheet |
| Engine Room Energy Efficiency | Engine performance deterioration attributed generically to 'wear' without checking hull/propeller fouling | Hull cleaning not scheduled in time, growing impact on consumption | See the topic sheet |
| Critical Spare Parts Management | Critical spares list not updated after changes to the identified critical equipment | Missing spares for equipment recently classified as critical | See the topic sheet |
| Fire Pumps and Fixed Systems in the Engine Room | Emergency pump test conducted using the same power supply as the main engine room | The test does not genuinely verify the emergency condition the pump is designed for | See the topic sheet |
| Bunkering Operations and Fuel Quality Control | Pre-bunkering safety checklist completed as a formality without genuine verification of conditions | Spill risk not adequately mitigated | See the topic sheet |
| Cyber Risk Management of Automation Systems | OT and IT networks not segmented, with shared access points | A compromise of the IT network (e.g. via email) can propagate to critical control systems | See the topic sheet |
From the PSC Knowledge Base of SuperbaKnowledge. This view selects and organises content published in SuperbaKnowledge; it does not modify or replace it. The linked Knowledge page remains the reference version, while official texts remain authoritative.
| Deficiency | Regulation | Indicative frequency | Possible consequence | Topic sheet |
|---|---|---|---|---|
| Non-functioning emergency fire pump or insufficient pressure | SOLAS Chapter II-2, Reg. 10 | High | Serious deficiency, possible detention | See the topic sheet |
| Acronym | Definition |
|---|---|
| AI/ML | Artificial Intelligence / Machine Learning |
| CBM | Condition-Based Maintenance |
| CMS | Continuous Machinery Survey |
| FMEA / FMECA | Failure Mode (and Effects) (and Criticality) Analysis |
| CII | Carbon Intensity Indicator |
| ISM | International Safety Management Code |
| MTBF | Mean Time Between Failures |
| MTTF | Mean Time To Failure — for non-repairable components |
| MTTR | Mean Time To Repair |
| PMS | Planned Maintenance System |
| RCM | Reliability Centered Maintenance — in the sense of SAE JA1011 |
| RUL | Remaining Useful Life |
Consolidated list of the sources cited. Updated as of August 2026; for application to a specific system, always consult the manufacturer's technical documentation and the classification society's rules.
| Source | Scope |
|---|---|
| ISM Code section 10 (Res. A.741(18) and amendments) | Maintenance of the ship and equipment; 10.3 critical equipment and testing of stand-by arrangements; 10.4 integration into the maintenance routine |
| IACS UR Z18 — Survey of Machinery | The ordinary machinery survey regime, including continuous surveys |
| IACS UR Z20 — Planned Maintenance Scheme (PMS) for Machinery, Req. 2001/Rev.2 2019 | The approved planned maintenance scheme, as an alternative to Continuous Machinery Survey |
| IACS UR Z27 — Condition Monitoring and Condition Based Maintenance, Req. 2018 | The Society-approved condition-based maintenance scheme, available to ships already on PMS |
| IACS Rec. No. 74 — Managing Maintenance (2001/Rev.2 2018) | IACS recommendation on maintenance management |
| SAE JA1011 (1999, rev. August 2009) and SAE JA1012 | Evaluation criteria for RCM processes: the seven questions, and the companion guide |
| Nowlan and Heap, Reliability-Centered Maintenance (United Airlines, 1978) | The six failure patterns and their shares: 11% age-related, 89% not |
| John Moubray, RCM II, and subsequent literature | Development and dissemination of the RCM methodology |
This course is educational material for training purposes and does not constitute a professional certification or qualifying credential. Read the full disclaimer.