Chillers FMEA Reference Guide for Reliability Teams

Connect with Industry Experts, Share Solutions, and Grow Together!

Join Discussion Forum
chillers-fmea-reference-guide-for-reliability-teams

Chillers are the single most energy-intensive and failure-prone asset class in most commercial and industrial facilities. ASHRAE Project 1043-rp — the canonical experimental study on centrifugal chiller faults — identifies seven fault categories that account for the majority of downtime events: reduced condenser water flow, non-condensable gas in refrigerant, condenser fouling, reduced evaporator water flow, excess oil, refrigerant overcharge, and refrigerant leak. Each of those faults degrades kW/ton efficiency before it becomes an outright failure — and every 10% efficiency drop on a 500-ton chiller costs $20,000 to $30,000 per year in additional electricity, long before it shows up as a breakdown. Recent Fuzzy-FMEA research on 31 real chiller station failure incidents ranked chiller surging (F-RPN 62.5), abnormal sound (60.8), and oil overheating (59.8) as the highest-priority failure modes. This is the working FMEA reference — every significant chiller failure mode organised by subsystem, with severity/occurrence/detection ratings on the standard 1-10 scale, Risk Priority Number, and the RCM-selected task per mode. Use it to run FMEA workshops, tighten your chiller PM programme, prioritise spares, and shift the operation from reactive fire-fighting to predictable, condition-based maintenance. Start free and load this FMEA reference into your chiller asset register this week, or book a demo to see the FMEA and condition-monitoring workflow mapped to your chiller fleet.

HVAC · Reliability Engineering · ASHRAE 1043-rp · Fuzzy-FMEA 2026

Chillers FMEA Reference Guide for Reliability Teams

Every significant chiller failure mode — compressor, evaporator, condenser, refrigerant loop, oil system, controls — with severity/occurrence/detection scoring, RPN, root cause, and RCM-recommended task. Anchored to ASHRAE Project 1043-rp canonical fault set, Fuzzy-FMEA field research on 31 real chiller stations, and industry-standard S×O×D scoring.

Start Free Trial Book a Demo

  • 7

    canonical chiller fault categories (ASHRAE Project 1043-rp)

  • $20–30K

    annual electricity cost of 10% kW/ton efficiency loss on a 500-ton chiller

  • 1°F / 1–2%

    approach-temperature drift per efficiency loss — tube fouling signature

  • Weeks

    early warning from vibration spectrum trending before mechanical failure

The Scoring Framework

How RPN Is Calculated — the Working Method

Every failure mode in this reference is scored using the industry-standard Risk Priority Number formula: RPN = Severity × Occurrence × Detection. Each factor is rated on a 1–10 scale. The higher the RPN, the higher the maintenance priority. Below are the working rating anchors so scores are defensible across FMEA workshops and reviews.

S · Severity

Consequence of the Failure

  • 1–2 · No noticeable effect, minor inefficiency
  • 3–4 · Reduced capacity, higher energy cost
  • 5–6 · Partial loss of cooling, occupant discomfort
  • 7–8 · Full loss of cooling, regulated-space impact
  • 9–10 · Safety event, refrigerant release, catastrophic damage
O · Occurrence

Likelihood of the Failure Mode

  • 1–2 · Rare — no history in the fleet
  • 3–4 · Occasional — 1 event in 5+ years per asset
  • 5–6 · Moderate — annual event across the fleet
  • 7–8 · Frequent — multiple events per year
  • 9–10 · Certain — occurs on every asset routinely
D · Detection

Likelihood of Catching Before Failure

  • 1–2 · Certain — continuous condition monitoring in place
  • 3–4 · High — routine inspection catches it reliably
  • 5–6 · Moderate — visible on periodic PM checks
  • 7–8 · Low — only caught with specialised testing
  • 9–10 · None — failure is the first indication

The Complete Reference

Chiller FMEA Worksheet — 20 Priority Failure Modes

This is the working reference. Twenty priority failure modes, organised by subsystem, with severity, occurrence, and detection scores giving the RPN — plus the RCM-selected task per mode. Higher RPN sits at the top of the maintenance queue. Values reflect typical commercial chiller operation; calibrate to your specific fleet failure history.

Subsystem Failure Mode Root Cause S O D RPN RCM Task
CompressorChiller surgingLow load, high lift, IGV mis-position854160Continuous surge monitoring, IGV calibration
CompressorBearing wear / spallingFatigue, lubrication breakdown944144Vibration spectrum analysis quarterly
Oil SystemOil overheating (>120°C)Condenser fouling, low voltage, low oil844128Oil temperature trend, oil analysis quarterly
RefrigerantRefrigerant leakSeal failure, fitting corrosion, tube crack853120Annual leak test, EPA 608 compliance, purge log
CondenserCondenser tube foulingPoor water treatment, scale build-up673126Approach temperature trend, annual tube brush
CompressorAbnormal sound / vibrationBearing, impeller imbalance, foreign object753105Vibration FFT monthly, acoustic monitoring
RefrigerantNon-condensable gas in refrigerantAir ingress through leaks, improper service663108Purge unit runtime log, condensing pressure trend
EvaporatorReduced chilled water flowStrainer blockage, pump degradation, air-lock74384Flow monitoring, strainer inspection quarterly
CondenserReduced condenser water flowCooling tower fouling, pump fault64372Flow trend, tower fill inspection annually
ControlsTemperature sensor driftSensor aging, calibration loss564120Annual calibration against reference standard
RefrigerantRefrigerant overchargeImproper service, guess-charge practice554100Weighed-charge verification at service, superheat check
Oil SystemExcess oil in refrigerant loopOil separator carryover, high migration545100Oil sight-glass check, separator diff-pressure trend
EvaporatorEvaporator tube freezeLow flow, low refrigerant, EXV fault92590Freeze protection setpoint verify, flow interlock test
CompressorMotor winding insulation degradationThermal cycling, moisture ingress934108Annual meg-ohm test, insulation resistance trend
ControlsFlow switch failureContact wear, contamination, wiring83496Quarterly flow-switch functional test
RefrigerantExpansion valve mis-tuningBulb charge loss, orifice contamination54480Superheat verification each PM cycle
CompressorShort-cycling (1–3 min intervals)Low buffer volume, refrigerant leak, control drift73363Runtime log review, buffer tank sizing check
ControlsVFD parameter driftFirmware issue, parameter overwrite53460Annual VFD parameter backup and verify
StructureAnchor / vibration isolator wearCorrosion, elastomer degradation44348Annual visual inspection, isolator torque check
Oil SystemOil filter clogContamination, extended service interval44348Oil filter diff-pressure alarm, scheduled change-out

How to read this: Modes with RPN ≥ 100 (navy-highlighted) go into the priority queue for condition monitoring and predictive intervention. Modes 60–99 (amber-highlighted) get scheduled inspection with trend-triggered action. Modes below 60 (soft-navy) get standard periodic PM with failure-finding tasks on hidden functions.

Subsystem Breakdown

FMEA by Chiller Subsystem — Where the Risk Actually Sits

Aggregating RPN by subsystem reveals where the maintenance dollar earns its highest return. Compressor and refrigerant loop routinely dominate — together representing more than half of aggregate chiller risk on centrifugal machines — with oil system, controls, and heat exchangers rounding out the profile.

Compressor

Highest Aggregate RPN

Surging, bearing wear, motor winding, short-cycling, sound/vibration. Vibration spectrum trending and IGV calibration are the highest-return interventions. Weeks of P-F lead time available on bearing modes.

Refrigerant Loop

Regulatory + Environmental Exposure

Leaks, overcharge, non-condensable gas, EXV drift. EPA 608 compliance mandatory; purge unit runtime log and superheat verification are the operational levers.

Heat Exchangers

The Efficiency Money Pit

Condenser and evaporator tube fouling. 1°F approach drift equals 1–2% efficiency loss. Approach-temperature trending catches fouling weeks before it starts costing serious kW/ton.

Oil System

Compressor Life Determinant

Oil overheating, excess oil migration, filter clog. Quarterly oil analysis and separator differential pressure trending are the two-thirds coverage for compressor life extension.

Controls & Sensors

Silent Efficiency Drift

Temperature sensor drift, flow switch failure, VFD parameter loss. Annual calibration and functional testing prevent the "chiller is running fine but the numbers are all wrong" failure mode.

Structure

Low RPN, Non-Zero Impact

Anchor wear, vibration isolator degradation. Low occurrence but annual visual inspection retained because the failure-finding task is cheap and the consequence of cascade failure is not.

The Economic Case

10% kW/ton Efficiency Loss = $20–30K per Year on a 500-Ton Chiller

Chiller PM is the one maintenance programme in a facility where cutting corners always shows up in the energy bill before it shows up as a breakdown. Approach-temperature drift, IGV mis-calibration, refrigerant charge error, and sensor drift each cost more in electricity in one year than the FMEA-driven PM programme costs to run for three. Oxmaint operationalises the entire chiller FMEA reference — RPN scoring per asset, condition-monitoring intervals per mode, and mobile work orders with the right refrigerant, torque, and calibration specs per chiller family.

Start Free Trial Book a Demo

Task Selection Decision Path

How the RPN Score Translates Into a Maintenance Task

RPN is a prioritisation tool, not the maintenance strategy itself. The strategy is selected using RCM decision logic — safety consequence, technical feasibility, cost-effectiveness, and the P-F interval available. Below is the working task-selection path applied to chiller failure modes.

  1. 01

    Safety or Environmental Consequence?

    If yes (refrigerant release, catastrophic pressure event, fire risk), task must be feasible and effective regardless of cost. Failure-finding retained on all hidden functions.

  2. 02

    Is Condition Monitoring Technically Feasible?

    If a detectable P-F interval exists and monitoring technology is available (vibration FFT, approach temperature, oil analysis), on-condition task is the default choice.

  3. 03

    Is the Failure Age-Related?

    If wear-out pattern documented (Nowlan-Heap Patterns A/B/C — rare in complex chiller subsystems), fixed-interval time or usage-based PM at end of useful life.

  4. 04

    Hidden Function?

    Flow switches, freeze protection interlocks, refrigerant relief valves — failure only reveals itself when the primary function fails. Failure-finding task scheduled at less than half the mean-time-to-failure.

  5. 05

    No Cost-Effective Task Available?

    Redesign consideration (dual-pump, N+1 chiller, redundant sensor). Run-to-failure accepted only on low-RPN modes with adequate sparing and no cascade consequence.

Built for Reliability Teams

How Oxmaint Runs the Chiller FMEA End to End

  • Live FMEA per Asset

    RPN Scored and Updated on Every Event

    Every chiller in the register carries its own live FMEA. Occurrence scores update automatically as work-order history accumulates; RPN recalculates and reprioritises the maintenance queue.

  • Condition Monitoring

    Vibration, Approach Temp, Oil Analysis

    IoT sensors and manual sample results flow into the same asset record. Approach temperature, vibration spectrum, and oil analysis all trend against the P-F thresholds documented per mode.

  • Auto Work Orders

    Trigger on Threshold, Not Calendar

    Condition-based work orders fire when trend data crosses P-detection thresholds — 1°F approach drift, vibration signature shift, oil metal-particle count spike. Fewer PMs, better catches.

  • EPA 608 Refrigerant Tracking

    Leak Rate, Recovery, Purge Log

    Purge runtime, leak-test dates, refrigerant recovery quantities, and technician certification all tracked at the asset level — audit-ready EPA 608 compliance without spreadsheets.

  • Mobile Execution

    Right Charge Weight, Torque, PPE In-Hand

    Technicians receive work orders with refrigerant type and weighed-charge target per chiller model, torque specs, and PPE requirements. Photo history from the last three services attached.

  • Reliability Reporting

    RPN Trend, kW/ton Baseline Deviation

    Facility manager and reliability director dashboards show RPN aggregate trend per chiller, efficiency drift from commissioning baseline, and cost-of-drift converted to annual electricity impact.

Measured Outcomes

What Reliability Teams Report Using This FMEA Reference

  • 10%+

    kW/ton Efficiency Recovered

    Teams tracking approach temperature and acting on 1°F drift routinely recover 10% or more of commissioning kW/ton baseline — $20–30K per year on a 500-ton chiller.

  • Weeks

    Bearing Failure Prevented

    Vibration spectrum trending on centrifugal compressors detects bearing wear, impeller fouling, and refrigerant flooding weeks before mechanical failure.

  • EPA 608

    Compliance Audit-Ready

    Leak-test dates, refrigerant recovery quantities, purge runtime, and technician certification tracked per asset — full compliance record produced in seconds.

  • $0

    Free Forever Plan to Start

    Cloud-based, mobile-first. Load this FMEA reference into one chiller asset today, run RPN scoring on your fleet, and scale to the full facility when the reliability gain is proven.

Frequently Asked

Chiller FMEA Questions

Where does the seven-fault ASHRAE 1043-rp reference come from?

ASHRAE Project 1043-rp is the canonical experimental study on centrifugal chiller faults, conducted on a McQuay PEH048J 90-ton chiller with faults artificially induced at four severity levels each. The seven categories — reduced condenser water flow, non-condensable gas, condenser fouling, reduced evaporator water flow, excess oil, refrigerant overcharge, and refrigerant leak — are the reference set every serious chiller reliability programme benchmarks against. Start free and load the ASHRAE seven-fault set into your chiller register today.

Should we adjust the RPN scores for our specific facility?

Yes — always. The values in this reference reflect typical commercial chiller operation. Occurrence in particular should be calibrated to your actual fleet failure history over the last 3–5 years. Severity should reflect your specific consequence categories (data centre uptime is a different severity than a school gym). Detection reflects the monitoring technology you have deployed. Recalibrate quarterly as data accumulates.

Does this apply to screw and absorption chillers, or only centrifugal?

The core failure modes — refrigerant loop, heat exchanger fouling, controls, oil system — apply across all chiller types. Compressor-specific modes differ: centrifugal machines get surge and IGV modes; screw chillers get rotor profile wear and slide-valve modes; absorption units get concentrator and generator modes with lithium bromide chemistry. Use this reference as the base and layer type-specific compressor modes for your fleet mix. Book a demo to see chiller-type-specific FMEA templates.

How does approach temperature trending catch fouling?

Approach temperature is the difference between the refrigerant saturation temperature and the leaving-water temperature in the condenser or evaporator. As tube surfaces foul, heat-transfer efficiency drops, and the approach temperature drifts up. A 1°F drift corresponds to a 1–2% kW/ton efficiency loss. Trending this metric weekly catches fouling weeks before the plant operator notices increased power consumption.

Is there a free plan to run this FMEA on one chiller first?

Yes. Oxmaint offers a free forever plan — enough to load this FMEA reference into one chiller, connect condition-monitoring data, calibrate RPN to your facility's history, and prove the reliability and efficiency gains before rolling out to the full plant. Cloud-based, mobile-first — no server procurement or infrastructure commitment. Sign up for the free plan and stand up chiller FMEA on one asset today.

Score · Prioritise · Monitor · Act

FMEA Is Not a Workshop Deliverable. It Is a Live Operating Document.

Twenty failure modes, ranked by RPN, mapped to root cause and RCM-selected task — that is the working chiller reliability programme. Anchored to ASHRAE Project 1043-rp, calibrated to your fleet history, and updated every time a work order closes with a root cause tag. Oxmaint runs the entire cycle: live FMEA per asset, condition monitoring against P-F thresholds, mobile execution with the right specs, and reporting that converts efficiency drift into annual electricity cost.

Start Free Trial Book a Demo


By William Jerry

✨

Experience
Oxmaint's
Power

Take a personalized tour with our product expert to see how OXmaint can help you streamline your maintenance operations and minimize downtime.

Book a Tour

Share This Story, Choose Your Platform!

Connect all your field staff and maintenance teams in real time.

Report, track and coordinate repairs. Awesome for asset, equipment & asset repair management.

Schedule a demo or start your free trial right away.

iphone

Get Oxmaint App
Most Affordable Maintenance Management Software

Download Our App