Digital Maintenance for Refining: CMMS Meets RCM

Connect with Industry Experts, Share Solutions, and Grow Together!

Join Discussion Forum
digital-maintenance-for-refining-cmms-meets-rcm

Refinery reliability lives and dies on the turnaround cycle. Units run three to six years between planned outages, and a single unplanned trip of a crude column, an FCC wet gas compressor, or an HF alkylation train pulls millions of dollars per day out of margin before the maintenance team even gets to the asset. The gap most refineries carry isn't the CMMS or the RCM analysis in isolation — it's that the RCM binder sits on a reliability engineer's shelf while the CMMS runs calendar-based PMs that no failure mode ever asked for. This guide shows how top refineries close that gap: the RCM 7-question decision tree per failure mode, the routing from failure mode to CMMS task type, the unit-criticality tiers that focus effort on the TAR-critical 20%, and the KPI set that plant managers actually open. Start free on OxMaint to configure the RCM-to-CMMS loop for your unit, or book a demo to see the failure-mode library in action.

Where CMMS Ends and RCM Begins — Refineries Bridge the Two
Failure-mode-driven maintenance for refining reliability programs.
3–6 yr
Typical run length between planned refinery turnarounds — TAR windows are the hard deadline every RCM decision maps to
7.7 yr
Pump MTBF benchmark for well-managed refineries (Bloch: 1,200 pumps, 156 repairs/yr)
$50–200M
Cost of a major crude/FCC/hydrocracker turnaround — 4–8 weeks per event
30–45%
Documented maintenance-cost reduction range when RCM findings actually reach CMMS execution

The Gap Most Refineries Carry · RCM on the Shelf, PMs on the Calendar

The pattern is remarkably consistent across refineries of every size. A reliability team completes an RCM study on the crude unit or the FCC, generates a thousand-page FMEA workbook, and hands it to the maintenance team. Six months later the CMMS still runs the same OEM-recommended calendar PMs it did before the study, because there was no live routing from a documented failure mode to a specific PM trigger, condition-monitoring alert, or spare-part reorder point. The RCM analysis dies on the shelf, and the crude unit still trips on the same seal failure the study identified. The rest of this guide is the operational bridge — where every RCM decision becomes a CMMS configuration, and every completed work order feeds back to validate or update the analysis.

The RCM 7-Question Decision Tree · Walked Per Failure Mode

John Moubray's seven questions (later codified as SAE JA1011) are the discipline behind any credible RCM program. Each failure mode on a critical refinery asset gets walked through the same seven questions — and the answer to question seven determines the CMMS task type that gets configured. Below is the ladder as reliability engineers actually use it on the plant floor.

Q1
What are the functions and desired performance standards of this asset in its operating context?
Example — an FCC main air blower delivers X kSCFM at Y psig at the regenerator, sustained for the TAR interval. Function defined before failure is defined.
Q2
In what ways can it fail to fulfill those functions? (Functional failures)
Air blower falls below required kSCFM, loses discharge pressure, trips on vibration, fails hidden protective function on surge control.
Q3
What causes each functional failure? (Failure modes)
Rotor imbalance, bearing wear, seal degradation, fouling on impeller, IGV linkage sticking, lube oil contamination.
Q4
What happens when each failure mode occurs? (Failure effects)
The physical, operational, and safety story per mode — what the operator sees, what the DCS records, what the downstream process does.
Q5
In what way does each failure matter? (Failure consequences)
Safety / environmental / operational / non-operational — a hidden protective failure gets treated differently from a redundant seal on a spared pump.
Q6
What can be done to predict or prevent each failure? (Proactive tasks)
Vibration route, oil analysis, thermography, IR scans, IGV stroke test, seal weep-hole inspection — whatever is technically feasible and worth doing.
Q7
What should be done if no suitable proactive task can be found? (Default actions)
Failure-finding for hidden functions, redesign, or run-to-failure if consequences allow. The answer to Q7 becomes the CMMS task type.

Failure Mode → CMMS Task Type · The Routing That Bridges RCM to Execution

This is the routing that separates a live RCM program from a shelved one. Every failure mode analyzed under the seven questions maps to exactly one of five CMMS task types — and the CMMS configuration (trigger, procedure, spare-part reorder, KPI feed) follows from the task type. The matrix below is the operational lookup reliability engineers use to configure OxMaint per failure mode.

RCM Task TypeWhen It AppliesCMMS ConfigurationRefinery Example
On-ConditionP-F interval measurable by a monitoring technique; failure gives warningCondition-monitoring PM route, threshold-triggered auto-WO on excursionCentrifugal pump vibration route on 1,200-pump population, oil analysis on turbines
Scheduled RestorationWear-out age exists, restoration returns capability, interval < wear-out ageTime-based or run-hour PM triggered at RCM-derived intervalOverhaul of FCC WGC coupling at 24-month interval; column tray restoration at TAR
Scheduled DiscardWear-out age exists, item replaced regardless of conditionTime-based or run-hour PM with mandatory-replace step and parts reorderFirewater pump batteries replaced every 3 yr; heater tube retubes at TAR interval
Failure-FindingHidden function; failure has no warning to operator until demand eventFunction-test PM at RCM-derived interval; auto-WO on any test failureRelief valve pop-testing, ESD test on FCC WGC surge trip, PSV overhaul cycle
Run-to-FailureNo safety or operational consequence; corrective cost < proactive costNo PM. Spare stocked, corrective procedure documented, MTBF trackedNon-critical utility motors, spared instrument air compressors on tertiary train
Load Your RCM Findings Into a CMMS That Runs the Routing — Free Forever
Sign up on OxMaint's free forever plan and import your failure-mode library. Each mode maps to one of the five task types, and the corresponding PM trigger, condition threshold, or function test configures itself. No spreadsheets, no dead binders. No card, no time limit.

Refinery Unit Criticality Tiers · Where the 20% of Assets Live

Every RCM implementation guide says the same thing: focus on the critical 10–20% of assets that drive 80% of unplanned downtime. In a refinery, that 20% is well-defined — the units and equipment that sit on the TAR critical path or that constrain crude throughput directly. The four-tier grid below is the criticality frame most refineries use to sequence RCM analysis and CMMS configuration.

Tier 1
TAR-Critical Path
FCC wet gas compressor (WGC) — consistently on the TAR critical path
FCC main air blower + reactor / regenerator internals
Crude distillation column + fired heater
Hydrocracker + hydrotreater reactors
HF alkylation reactor + acid handling systems
Coker drums + heater tubes
Full RCM analysis. Every failure mode mapped, every task type configured, condition monitoring on every rotating asset.
Tier 2
Unit-Critical
Reformer + reboiler + product pumps
Sulfur recovery unit (SRU) + amine treating
Cooling tower + main utilities
Charge pumps to Tier 1 units
API 653 field-erected storage tanks (product / crude)
Blending line rundown pumps + custody transfer meters
RCM with priority on high-consequence modes. Time-based + condition-based mix; RBI drives inspection intervals.
Tier 3
Supporting
Spared / parallel pumps and blowers
Instrument air compressors (with spare)
Non-critical exchangers
Ancillary utility motors
Laboratory support equipment
Warehouse and shop equipment
Standard PM package. Failure-finding on protective devices. MTBF tracked, not necessarily improved.
Tier 4
RTF-Eligible
Fully redundant utility motors
Non-critical lighting circuits
Office HVAC equipment
Non-process building assets
Landscape and grounds equipment
Spared tertiary instrument air trains
Run-to-failure by design. Spare stocked, corrective procedure documented, no PM waste.

The API Standards Overlay · What Governs Inspection Intervals

Refinery maintenance is regulated maintenance. API standards define the design integrity, inspection scope, and interval framework for every critical asset class — and RCM decisions live inside those constraints, not around them. The card grid below is the standards stack every refinery CMMS should reference against the asset record.

API 653
Field-Erected Tanks
Inspection, repair, alteration, and reconstruction of aboveground storage tanks. Floor mapping, NDE, weld vacuum-box testing, deflection limits.
API 610
Centrifugal Pumps
Design and specification of centrifugal pumps for petroleum service. Anchor for pump MTBF benchmarks and seal-life targets.
API 570
Piping Inspection
In-service inspection of piping systems. Thickness monitoring, corrosion-loop analysis, CML surveys, remaining-life calculations.
API 510
Pressure Vessels
In-service inspection of pressure vessels. External / internal / on-stream inspection frequencies and NDE requirements.
API 580 / 581
Risk-Based Inspection
Probability of Failure × Consequence of Failure framework to optimize inspection scope and intervals across the fixed-equipment population.
API 682
Pump Shaft Seals
Seal design, materials, and testing. ESA 2025 guidance: 60 months as a good site-level seal MTBF benchmark, best-in-class 100 months.

Refinery KPIs Plant Managers Actually Open

A CMMS that captures work orders but never surfaces KPIs is a system of record, not a reliability engine. The set below is the tight KPI dashboard every refinery reliability program should feed from execution data — with target ranges anchored to industry benchmarks, not vendor marketing numbers.

MTBF (Rotating Equipment)
Pump target: 5–7+ years (7.7 yr Bloch benchmark for well-managed refineries)
Tracked per pump population — trends up as RCM-derived vibration routes catch bearing wear early.
MTTR (Mean Time to Repair)
Target: trending down; segmented by unit and failure mode
A downward MTTR trend is the fingerprint of a mature spare-parts strategy and mobile work-order execution.
PM Compliance
Target: ≥ 90% within the compliance window
Falls sharply if the PM library is over-scoped from calendar-based habit — RCM optimization typically drops the PM count and raises compliance.
Maintenance Backlog
Target: 2–4 weeks of ready backlog per craft
Under 2 weeks signals planning gap; over 6 weeks signals resource gap or scope inflation.
Schedule Compliance
Target: ≥ 85% of scheduled work completed in the assigned week
The single strongest predictor of TAR readiness and month-over-month reliability improvement.
Wrench Time / RBI-Driven Coverage
Target: rising wrench-time %, rising RBI coverage on Tier 1/2 fixed equipment
Mobile CMMS + RBI-informed inspection routing lifts productive time and cuts calendar-based inspection waste.

How OxMaint Runs the CMMS-Meets-RCM Program for Refining

The seven-question tree, the failure-mode-to-task-type routing, the four-tier criticality frame, the API standards overlay, and the KPI dashboard all live in the same platform — asset hierarchy that matches unit/train/loop, mobile work orders that operators and mechanics actually complete, condition-based triggers wired to the DCS and vibration route, and a KPI feed that plant managers open on Monday.

Structure
Asset Hierarchy: Unit → Train → Loop → Tag
Refinery-native hierarchy mirrors the P&ID — every asset tag traceable up to unit and down to failure-mode library.
Library
Failure Modes Attached to Assets
Modes libraried once per asset class (WGC coupling wear, seal failure, tube fouling) and reused across the population.
Route
Every Mode → One CMMS Task Type
Each failure mode maps to on-condition, scheduled restoration, scheduled discard, failure-finding, or RTF — no orphan modes.
Trigger
DCS + Vibration + Time in One Engine
Auto work orders fire from DCS threshold, vibration excursion, run-hours, or calendar interval — routed to the right craft.
Execute
Mobile Work Orders + Spares
Technicians complete on the phone with photo, parts drawn against reservation, actual time captured against the failure mode.
Prove
KPI Dashboard + RCM Feedback Loop
MTBF, MTTR, PM compliance, backlog, schedule compliance surfaced by unit. Actual failure events feed back to update the RCM analysis.
Bridge the RCM Binder and the CMMS Work Queue in Weeks, Not Quarters
Free forever plan — no card, no time limit. Load your unit, load your failure-mode library, and every mode maps to a live CMMS task type before the next TAR planning cycle. Or book 30 minutes and we'll walk one of your critical assets end-to-end on the platform.

Frequently Asked Questions

What actually integrates CMMS and RCM in a refinery — beyond having both tools?
The integration is a routing: every failure mode identified in the RCM analysis must map to exactly one of the five CMMS task types (on-condition, scheduled restoration, scheduled discard, failure-finding, run-to-failure), and the corresponding PM trigger, condition threshold, or function test must exist in the CMMS. Refineries that treat RCM as a separate documentation exercise end up with a study on the shelf and a CMMS still running OEM calendar PMs. Book a demo to see the routing live.
Which refinery assets should the first RCM iteration cover?
Start with the TAR-critical path — FCC wet gas compressor, FCC main air blower, crude column and fired heater, hydrocracker and hydrotreater reactors, HF alkylation train, coker drums. Any unplanned trip on these units pulls throughput out of the entire refinery. Once Tier 1 is covered, extend to Tier 2 unit-critical assets (reformer, SRU, cooling tower, API 653 storage tanks). Tier 3 and 4 assets get standard PM packages and RTF-by-design respectively.
What are realistic MTBF benchmarks for refinery rotating equipment?
Bloch's widely-referenced benchmark for well-managed U.S. refineries is a pump-population MTBF of 7.7 years (1,200 installed pumps with 156 repair incidents in a year). ESA's 2025 guidance suggests 60 months as a good site-level pump-seal MTBF benchmark, with best-in-class operators targeting 100 months. Poorly-managed populations fall to 3 years — the spread is largely specification, installation, and operating discipline, not the pumps themselves.
How does API 580/581 Risk-Based Inspection interact with the RCM program?
RBI applies primarily to fixed equipment — tanks, pressure vessels, piping — and uses Probability of Failure × Consequence of Failure to prioritize inspection scope and intervals. RCM applies primarily to rotating and control equipment, and uses failure-mode analysis to select maintenance task types. In a refinery both frameworks feed the same CMMS: RBI drives the API 653 / 570 / 510 inspection schedule against asset records; RCM drives the PM library, condition-monitoring routes, and function tests. Together they replace calendar-based waste with consequence-based coverage.
How fast can a refinery deploy OxMaint against an existing RCM study?
Weeks, not quarters. Existing failure-mode libraries import against the asset hierarchy, each mode maps to one of the five CMMS task types, and PM triggers, condition thresholds, and function tests configure from the RCM decision. Mobile work orders roll out to craft in hours per crew. The KPI dashboard begins reporting MTBF, MTTR, PM compliance, backlog, and schedule compliance from the first week of live work-order data. Sign up free to start on your unit.

By William Jerry

Experience
Oxmaint's
Power

Take a personalized tour with our product expert to see how OXmaint can help you streamline your maintenance operations and minimize downtime.

Book a Tour

Share This Story, Choose Your Platform!

Connect all your field staff and maintenance teams in real time.

Report, track and coordinate repairs. Awesome for asset, equipment & asset repair management.

Schedule a demo or start your free trial right away.

iphone

Get Oxmaint App
Most Affordable Maintenance Management Software

Download Our App