- TAPPI OnDemand
- An Effective Roadmap for Failure Mode and Effects Analysis (FMEA)
An Effective Roadmap for Failure Mode and Effects Analysis (FMEA)
May 22, 2026
Failure Mode and Effects Analysis (FMEA) stands as one of the most effective risk assessment tools available to reliability engineers, quality professionals, and operations managers. While the methodology may initially appear complex, understanding its systematic approach reveals a logical framework for identifying, prioritizing, and mitigating potential failures before they impact operations, safety, or customer satisfaction. This guide provides a practical roadmap for conducting FMEA, from initial planning through implementation and continuous improvement.
UNDERSTANDING FMEA:
PURPOSE AND APPLICATIONS
FMEA is a structured, proactive methodology used to identify potential failure modes within a system, assess their risk, and develop appropriate mitigation strategies. Originally developed by the aerospace industry in the 1960s, FMEA has since become a standard practice across manufacturing, health care, automotive, energy, and numerous other sectors.
The fundamental objective of FMEA is to anticipate problems before they occur, enabling organizations to allocate resources effectively toward prevention rather than firefighting. By systematically examining how systems, processes, or products can fail, teams can make informed decisions about design improvements, maintenance strategies, and quality controls.
WHEN TO CONDUCT AN FMEA
FMEA delivers the greatest value when applied to situations where failure consequences are significant. Consider implementing FMEA in the following scenarios:
- New product or process development—Identifying potential issues during the design phase is significantly more cost-effective than addressing problems after production begins.
- Major modifications to existing systems —Changes to proven designs or processes can introduce unforeseen failure modes that warrant systematic analysis.
- Critical equipment or processes—Assets or operations where failure could result in safety incidents, environmental damage, or substantial financial losses merit comprehensive failure analysis.
- Chronic reliability issues—When recurring failures continue despite traditional troubleshooting efforts, FMEA can reveal root causes and systemic vulnerabilities.
THE FMEA PROCESS:
A STEP-BY-STEP APPROACH
Step 1: Establish the Analysis Team. Effective FMEA requires cross-functional collaboration. Assemble a team that includes individuals with diverse perspectives and expertise relevant to the system being analyzed. Typical participants include design engineers, process engineers, maintenance personnel, operators, quality specialists, and subject matter experts. A team size of four to eight members generally provides sufficient diversity without becoming unwieldy. Designate a facilitator to guide the analysis and ensure productive discussions.
Step 2: Define the Scope and Objectives. Clearly delineate what your team will analyze. Is the scope a complete production line, a single machine, a specific subsystem, or an entire manufacturing process? Establishing precise boundaries prevents scope creep and maintains focus.
Document the analysis objectives and gather relevant information including technical drawings, process flow diagrams, operating procedures, historical failure data, and maintenance records. This preparation streamlines the analysis process and ensures decisions are grounded in facts.
Step 3: Identify System Components or Process Steps. Break down the system or process into logical elements that can be analyzed individually. For equipment FMEA, this typically involves identifying major subsystems, assemblies, and components. For process FMEA, identify each discrete operation or manufacturing step.
The level of detail should be appropriate to the analysis objectives. Excessive granularity can make the analysis unmanageable, while insufficient detail may overlook critical failure modes.
Step 4: Document Functions and Requirements. For each element identified in Step 3, clearly define its intended function and performance requirements. Functions should be specific and measurable where possible.
For example, rather than stating “pump moves fluid,” specify “pump delivers 500 GPM of water at 150 PSI with efficiency greater than 75 percent.” Include both primary function (the main purpose) and secondary functions (supporting capabilities such as containing pressure or maintaining alignment).
Step 5: Identify Potential Failure Modes. A failure mode represents any way a component or process step could fail to perform its intended function. Brainstorm all conceivable failure modes for each function, drawing upon team experience, historical data, and theoretical possibilities.
Consider various failure categories including complete loss of function, partial or degraded performance, intermittent operation, and unintended operation. Even seemingly unlikely failures should be documented at this stage, as their risk will be quantified in subsequent steps.
Step 6: Determine Failure Causes. For each failure mode, identify the underlying root causes. What conditions, defects, or circumstances could trigger this particular failure? A single failure mode may have multiple potential causes, and each should be documented separately.
Understanding failure mechanisms provides the foundation for developing effective preventive measures. Causes might include design deficiencies, material degradation, manufacturing defects, improper maintenance, operating errors, or environmental factors.
Step 7: Analyze Failure Effects. Describe the consequences of each failure mode, considering impacts at multiple levels. What are the immediate, local effects on the component or process step itself? How does the failure propagate to affect the overall system? What are the ultimate consequences for operations, safety, environment, and the customer?
This analysis should trace the chain of events from initial failure through final impact, providing a comprehensive understanding of why each failure mode matters.
Step 8: Assign Severity Ratings. Quantify the seriousness of each failure mode’s effects using a standardized rating scale, typically ranging from 1-10:
- 1-3: Minor impact with negligible consequences.
- 4-6: Moderate impact causing measurable problems (but not critical).
- 7-8: Significant impact with serious operational, safety, or financial consequences.
- 9-10: Catastrophic impact potentially causing injury, major environmental damage, or severe business disruption.
Establish clear rating criteria specific to your site and apply them consistently throughout the analysis.
Step 9: Document Current Controls. Identify existing measures that either prevent the failure from occurring or detect it before significant effects materialize. Prevention controls include design features, protective devices, operating procedures, and maintenance practices. Detection controls include inspections, testing, monitoring systems, and alarms. This inventory of current controls informs the occurrence and detection ratings in the following steps.
Step 10: Rate Occurrence Probability. Assess how frequently each failure mode is likely to occur given existing prevention controls. Use a 1-10 scale where:
- 1-2: Remote probability; failure is unlikely during equipment or product lifetime.
- 3-4: Low probability; isolated failures may occur.
- 5-6: Moderate probability; occasional failures can be expected.
- 7-8: High probability; repeated failures are likely.
- 9-10: Very high probability; failure is almost inevitable.
Base your occurrence ratings on historical failure data, when available, supplemented by engineering judgment for new designs or insufficient data situations.
Step 11: Rate Detection Capability. Evaluate the likelihood that existing detection controls will identify the failure mode before the effects occur or reach the customer. Again, you can use a 1-10 scale like this:
- 1-2: Very high detection capability; multiple reliable methods ensure discovery.
- 3-4: High detection capability; likely to be discovered.
- 5-6: Moderate detection capability; may or may not be discovered.
- 7-8: Low detection capability; unlikely to be discovered in time.
- 9-10: Very low or no detection capability; failure will reach customer or cause effects.
Failures that are difficult to detect warrant particular attention, as they can cause significant problems before being discovered.
Step 12: Calculate Risk Priority Number (RPN). The RPN provides a quantitative measure for prioritizing failure modes:
RPN = Severity × Occurrence × Detection
RPN values range from 1 to 1,000, with higher numbers indicating greater risk. While RPN serves as a useful prioritization tool, it should not be applied mechanically. High-severity failures merit attention even when occurrence and detection ratings suggest lower overall RPN values. Many organizations establish thresholds for severity, occurrence, and detection that trigger action regardless of the calculated RPN.
Step 13: Develop and Implement Corrective Actions. For high-risk failure modes, develop specific actions to reduce risk. These actions should target one or more of the three risk factors:
- Reducing Severity: Modify the design or process to eliminate the failure mode or minimize its consequences. This is often the most effective but also the most challenging approach.
- Reducing Occurrence: Implement measures to prevent the failure from happening, such as improved maintenance procedures, upgraded components, enhanced process controls, or operator training.
- Improving Detection: Add or enhance monitoring, inspection, or testing capabilities to identify failures earlier, before significant effects occur.
For each recommended action, assign responsibility, establish target completion dates, and document the rationale. Track implementation progress and verify that actions achieve the intended risk reduction.
Step 14: Recalculate Risk and Maintain the FMEA. After implementing corrective actions, reassess the severity, occurrence, and detection ratings to determine the resulting RPN. This quantifies the improvement achieved and verifies that risk has been reduced to acceptable levels. The FMEA should be treated as a living document, regularly reviewed and updated as conditions change, new failure modes emerge, or lessons are learned from operational experience.
THE BUSINESS CASE FOR FMEA
Organizations that effectively implement FMEA realize substantial benefits, including reduced warranty costs, improved product quality, enhanced operational reliability, better resource allocation, and reduced liability exposure. By identifying and addressing potential failures proactively, companies avoid the significantly higher costs associated with failures in production or the field.
FMEA also provides defensible documentation of due diligence, demonstrating that systematic risk assessment informed design and operational decisions. This documentation can prove valuable for regulatory compliance, customer audits, and potential liability situations.
CONCLUSION
Failure Mode and Effects Analysis is a disciplined, systematic approach to anticipating and preventing problems before they impact operations or customers. While this path requires your mill to invest time and resources, the return on investment typically manifests through avoided failures, optimized maintenance strategies, and improved designs.
Success with FMEA comes not from rigid adherence to forms and procedures, but from fostering the right mindset. It is a mindset of proactive risk management and continuous improvement. Organizations that embed FMEA into their standard practices are better able to identify, assess, and mitigate risks across all aspects of their operations.
For those beginning their FMEA journey, start with a focused pilot project on a manageable system or process. Build competency through practice, learn from the experience, and gradually expand the application of FMEA to additional areas. The investment in developing this capability will pay dividends through improved reliability, reduced costs, and enhanced competitive advantage.
