In this episode of the Pipeliners Podcast, Russel Treat dives into key considerations for alarm management in pipeline operations, drawing on his years of personal experience.
The episode covers the challenges of defining alarm severity, rationalizing alarms, and navigating the technical complexities of SCADA systems. Additionally, the episode highlights best practices for ensuring alarms are meaningful, prioritizing critical alarms, and optimizing SCADA for effective alarm management.
Considerations in Alarm Management Show Notes, Links and Insider Terms
- Russel Treat is the CEO of EnerACT Energy Services, as well as the host of the Pipeliners Podcast and the founder of the Pipeline Podcast Network. Connect with Russel on LinkedIn.
- EnerACT Energy Services is a holding company of oil and gas technology and service companies. EnerSys Corporation is a frequent sponsor of the Pipeliners Podcast. Find out more about how EnerSys supports the pipeline control room through compliance, audit readiness, and control room management through the POEMS Control Room Management (CRM Suite) software suite.
- The CRM Rule (Control Room Management Rule as defined by 49 CFR Parts 192 and 195) introduced by PHMSA provides regulations and guidelines for control room managers to safely operate a pipeline. PHMSA’s pipeline safety regulations prescribe safety requirements for controllers, control rooms, and SCADA systems used to remotely monitor and control pipeline operations.
- Control Room Management is regulated by PHMSA under 49 CFR Parts 192 and 195 for the transport of gas and hazardous liquid pipelines, respectively. PHMSA’s pipeline safety regulations prescribe safety requirements for controllers, control rooms, and SCADA systems used to remotely monitor and control pipeline operations.
- Alarm management is the process of managing the alarming system in a pipeline operation by documenting the alarm rationalization process, assisting controller alarm response, and generating alarm reports that comply with the CRM Rule for control room management.
- AOC (Abnormal Operating Condition) is defined by the 49 CFR Subpart 195.503 and 192.803 as a condition identified by a pipeline operator that may indicate a malfunction of a component or deviation from normal operations that may indicate a condition exceeding design limits or result in a hazard(s) to persons, property, or the environment.
- ISA-18.2 defines an alarm as “an audible and/or visible means of indicating to the operator an equipment malfunction, process deviation, or abnormal condition requiring a response.
- Alarm rationalization is a component of the Alarm Management process of analyzing configured alarms to determine causes and consequences so that alarm priorities can be determined to adhere to API 1167. Additionally, this information is documented and made available to the controller to improve responses to uncommon alarm conditions.
- API 1167 provides operators with recommended industry practices in the development, implementation, and maintenance of an Alarm Management program. The implementation of API 1167 is required by reference in the CRM Rule.
- SCADA (Supervisory Control and Data Acquisition) is a system of software and technology that allows pipeliners to control processes locally or at remote location. SCADA breaks down into two key functions: supervisory control and data acquisition. Included is managing the field, communication, and control room technology components that send and receive valuable data, allowing users to respond to the data.
- HMI (Human Machine Interface) is the user interface that connects an operator to the controller in pipeline operations. High-performance HMI is the next level of taking available data and presenting it as information that is helpful to the controller in understanding the present and future activity in the pipeline.
- Alarm Response: Specified actions for pipeline controllers when alarms occur.
- Emergency Response: Actions taken when an emergency is identified, aiming to mitigate consequences.
- Operating Condition: Pipeline system state categorized as normal, abnormal (AOC), or emergency.
- ESD (Emergency Shut Down Systems) are high-powered control systems designed to protect people, pipelines, and the environment in the event of a pipeline operating beyond set control limits.
- Situational Awareness: Understanding of current operating conditions and potential risks.
- HCA (High-Consequence Areas) are defined by PHMSA as a potential impact zone that contains 20 or more structures intended for human occupancy or an identified site. PHMSA identifies how pipeline operators must identify, prioritize, assess, evaluate, repair, and validate the integrity of gas transmission pipelines that could, in the event of a leak or failure, affect HCAs.
- Roles and Responsibilities: Designation of individuals with authority to direct actions or supersede controller authority.
- GIS Mapping: Identifying geographic information, such as proximity to valves or HCAs, for emergency planning.
- Alarm Analysis: Reviewing alarm responses and emergency situations to identify areas for improvement.
- Continual Improvement: Ongoing efforts to enhance alarm management and emergency response programs.
- CRMP (Control Room Management Plan) captures the policies and procedures that are to be followed in the control room to ensure the safe operations of pipeline assets. Operators that are subject to the PHMSA CRM Rule are required to have a CRMP.
- KPI (key performance indicator) is a measurable value that is intended to show how well a business is adhering to its business model and strategies.
- PLCs (Programmable Logic Controllers) are programmable devices placed in the field that take action when certain conditions are met in a pipeline program.
Considerations in Alarm Management Full Episode Transcript
Russel Treat: Welcome to the “Pipeliners Podcast,” episode 356, sponsored by EnerSys Corporation, providers of POEMS, the Pipeline Operations Excellence Management System, operations and compliance software for the pipeline operator to address safety program management, control room management, and field operations. Find out more about POEMS at enersyscorp.com.
[background music]
Announcer: The Pipeliners Podcast, where professionals, bubba geeks, and industry insiders share their knowledge and experience about technology, projects, and pipeline operations.
Now your host, Russel Treat.
Russel: Thanks for listening to this podcast. I appreciate you taking the time. To show the appreciation, we give away a customized YETI tumbler to one listener every episode. This week, our winner is Scott Olsen with ONEOK. Congratulations, Scott. Your YETI is on its way. To learn how you can win this signature prize, stick around till the end of the episode.
This week, our guest is myself and myself only. I’m going to be talking about considerations in alarm management. Let’s just dive in. Normally, when I have a guest on, I ask the guest to do an introduction. I’ll do a little introduction of myself. I’ll talk a little bit about my background and where I come from in alarm management.
I started working in field measurement in about 1998. Prior to that, I’d been working running a company that was building software for the measurement back office. From that field work in measurement, we started doing some telecommunications. From doing telecommunications, we started doing some simple HMIs. Then we started doing some control rooms.
By 2007, we had done quite a number of control rooms in pipeline operations. We thought we were pretty good at what we were doing. In 2007, PHMSA published the control room management rule as a notice of preliminary rulemaking.
We heard about that, went and read it, and said, “What the heck is all this? What does this mean?” We went and found some industry experts to help educate and guide us. Some of those guys have been on the podcast over the years.
In particular, Doug Rothenberg was a gentleman I worked with in-depth for a number of years, went to school on him. Doug literally wrote the book on process alarm management. I read that thing and did a bunch of workshops with him and, through that process, built up expertise in alarm management.
Since about 2012, we started building our own alarm management software as part of our control room management software suite. We’ve implemented that with a number of pipeline operators. We’re actually at a place in our process where we’re doing a number of alarm management projects for several different operators.
I am the person on our team that runs point on all that. While we have other people that do other things, when we get to the really deep, nitty-gritty alarm management stuff, that ends up being in my domain. I work with our customers to do that.
Through working with this half a dozen customers, where we’re helping them implement alarm management, we’ve been learning some things. We’ve been learning quite a bit.
First thing I want to do is I want to talk about some of the key challenges in alarm management. I would first frame these as, how do you approach rationalization, particularly as it relates to alarm severity? The API 1167 definition of an alarm is a notification to a controller that requires action. I’m paraphrasing a little bit, but basically, that’s the API 1167 definition.
When we’re doing alarm management, we add to that definition and say a notification of the controller requiring action in a time frame or a bad outcome will occur, so being a little bit more specific. The time frame is important as you’re doing rationalization. Understanding the nature of the potential bad outcome is important as well.
One of the things, when we’re doing our alarm workshops that we’re trying to anchor with the customers, with the operators, is “What is the purpose of an alarm?” and “What is the purpose of alarm management?”
The purpose of alarm is to provide the operator with a notice that demands attention because if that attention doesn’t come to that notice, you’re going to have a bad outcome. You want to make those things evident. That’s one purpose, is to make sure every alarm is actually an alarm.
One of the challenges is that, in the oil and gas pipeline operations world, we have used the word “alarm” to mean a lot of different things. There’s alarms in the PLC. There’s alarms in the SCADA system. I’m talking about when you’re working in the background. You’re building things. All the notifications in SCADA are typically called an alarm when you go in to configure them.
There’s a distinction between a little “a” alarm, something I configure in the automation and a big “A” Alarm, something demanding action. Many of the things that come to the control room are really to provide information to the controller, not to request action of the controller. That’s a really important distinction to make.
Most operators are clear about that at this point. We’ve been working with the alarm management requirements of the control management rule now for about a dozen years. We’re beginning to have a clear understanding of that. Getting the entire organization to understand little “a” alarm, big “A” Alarm and get them all in the same language can be quite challenging.
The other thing about rationalization, the purpose is to put the highest severity alarm at the top of the list for the console operator to work on. If I’m working a console and I get, over the course of a typical day, 20 to 30 alarms, those things don’t come in spaced out. They tend to come in in groups.
What I’m trying to do with setting severity is make sure that the alarm that I want the controller to work first lands at the top of their list. I want to make it clear that these alarms demand immediate attention and these alarms demand attention but they’re not necessarily immediate.
You have to define, in every control room, what immediate looks like. Immediately is typically within five minutes. Not immediate is something between 5 minutes and 30. Then “I have lots of time to respond” would be 30 to an hour. That creates a frame for bounding what that means.
The purpose of alarm management is, first, to make sure every alarm is truly an alarm. It’s, second, to put the highest severity alarm at the top of the operator’s console. It’s top of their list to work.
Then thirdly, it is to make sure that every alarm can be understood. I understand the workflow to determine what’s the root cause behind the alarm and what I would do to address it. That’s the purpose.
Our mechanism for determining severity is to look at risk. If I get an alarm and it demands an action and it demands that action in a time frame or a bad outcome will occur, what is the consequence of that bad outcome determines the alarm severity. We’re not trying to say that, “This alarm will have this consequence.” It’s a way to understand and calculate and quantify risk.
There’s a legacy approach. The reality for many operators when the control room measurement rule came out is they had to get something in place relatively quickly to address alarm management. A lot of those operators made some simplifying assumptions. They said things like, “All high high pressures, we’re going to consider critical alarms.”
That might be useful, but in our experience, it’s not. There are two basic frames or approaches for doing this risk assessment to determine alarm severity. One is the frame that’s in API 1167 is I build a frame. I take the highest-score risk.
Then there’s another mechanism that’s defined more in ISA-18.2, which says, “I’m going to define a matrix. I’m going to define multiple risks. I’m going to quantify a total risk score.” Our belief is that the best practice is more the 18.2 approach.
There could be situations where the combined risk, while not severe in any one area, might overall lead to a calculation to cause you to rate it higher. That’s a subtlety.
What’s important is you understand your frame and you understand your mechanism for determining risk and to remember that the purpose of all this is just to put the most critical alarms to the top of the list. In other words, what do I want the controller to work on first if there’s more than one on the console?
One of the things that we’re discovering as we’re working with these customers is we started by taking a standard safety risk matrix and using it to do the alarm rationalization and tweaking it a little bit. What we’re beginning to find is that can be complicated. It can be challenging.
For example, in a gas utility, typically, the key things that determine risk are the pressure that I’m operating that particular facility at, the location of the facility, and how that facility affects deliverability.
Let me break this down a little bit. Pressure. In a typical gas utility system, I might have a feeder coming into my system that’s operating in the 600 to 850 PSI range. I might have a 325 PSI loop around the city that’s my main feeder.
I’ll have lines coming off of that where I’m regulating the gas down to 125 to 150 PSI. Then off of that, I might be regulating it down to 50 or 60 PSI before I regulate it down to the pressure I feed to the house. A 350 PSI line has more risk than a 60 PSI line, one way to think about risk.
The other part of the typical frame is location. In a utility, the site I’m looking at, is it in an open, green area, relatively unpopulated, rural? Is it in a neighborhood? Is it in an apartment complex? Is it inside of an apartment building or underground infrastructure in a highly populated city center? Those things would go to risk. More risk where there’s more population.
Lastly, there’s deliverability. That is, what’s the impact of not having flow through this site? That is a risk frame that makes sense for a utility operator because now I can use that as a way to calculate my risk and make sure that I get that highest-severity alarm to the top of the controller’s console.
A different frame might be a crude oil system. In a crude oil system, I’m probably more concerned about the size of the line and where the facility is located. There’s a difference between a crude oil system in West Texas and a crude oil system that’s running through a beautiful environmental area that a lot of the public use and enjoy.
Different kind of operations, different kind of risk frame, but, again, trying to do something that makes it easy to determine, as I do rationalization, where that goes, what I do in order to make something go to the top of the list.
Once you settle on what your frame’s going to be, it gets relatively straightforward. Typically, you’re looking at risk to people/health, risk to environment, risk to property. You can call that a financial risk. Some operators might look at legal risk or reputation risk. You can put that in the domain of license to operate.
We’ve worked with operators that have assets in cities or near schools, where something as simple as cracking a relief, where that might have very little impact if that were on a remote site, it has some significant impact if it’s inside of a city and located near a school, within a quarter- or half-mile of a school.
These are the considerations that you look at, as each operator, to get your risk frame together. Again, remembering the whole purpose of the exercise is keep it as simple as possible in order to put the most severe potential outcome at the top of the pipeline controller’s alarm list.
That’s one of the key challenges, is just getting clear about what my approach to risk management is and trying to frame that in a way that’s more tied to my operation and the nature of where and how I operate and with what fluid.
The next big challenge — this is a much bigger, much more complicated challenge — is the whole challenge around SCADA. Not all SCADA systems are equal. What I mean by that is the technical capability of a SCADA system, particularly in terms of its ability to do graphical representation and animations in the HMI, can be very different from vendor to vendor.
There are a number of operators that we are working with that have older SCADA systems, where those software platforms just don’t have some of the capabilities in it that you need in order to do proper alarm management.
I need to define a little bit about what I mean by proper alarm management. There’s several things you need to be able to do in the SCADA system as it relates to the alarming specifically. In the alarm list, I need to be able to segregate alarms from alerts.
Alarms are those things requiring action. Alerts are notifications, things I want the controller to be aware of, but they don’t require action, so things like confirming that a valve closed, that sort of thing, statuses, that sort of thing. Those are often called alarms in a SCADA system. In fact, they’re not generally an alarm in the context of it requires action.
One thing I got to do is segregate my alerts from my alarms. The other thing I got to do is I’ve got to be able to identify and sort my alarm list. Ideally, I’m going to be able to put the highest-severity unacknowledged alarm first, followed by the highest-severity acknowledged alarms.
Then I need symbology, animation, color, whatever, that tells me, “This alarm is active, unacknowledged. This alarm is active and acknowledged. This alarm is cleared and unacknowledged.”
One of the things that can’t happen is an alarm can go into alarm and then clear and come out. Then there’s no notice to the controller that that occurred. That’s a problem because it’s a loss of operating context for the controller. Those are features you need in the alarm list.
In addition and probably, in many systems that I’ve seen, not well-implemented, I need a mechanism for animating the alarm in the HMI. I need to know that, for that pressure, that pressure is an alarm. I need to know when that thing goes in an alarm. I need to know how severe is that alarm on that pressure.
Those are basic requirements. There are more advanced capabilities, like I need an overview. I need to have a site overview. I need to do alarm roll-up, which would mean that, in my overview, like my pipeline overview, I see a site.
I want to know, are there any alarms at that site? If there are alarms, I want to know, what is the highest-severity? I want to animate based on the same criteria I used for the list. If there is an active, unacknowledged alarm, that’s the animation I want to use. Once everything’s acknowledged, I want to use the highest-severity uncleared alarm.
The idea is I can get high-level operating context and drive into it. Now we start getting into more elegant features, more complex features to implement in the SCADA platform.
Not all SCADA is equal. One is technical capability. That’s what we’ve been talking about. The other is actual implementation. I can have the same product ArchestrA in location A and ArchestrA in location B. To the untrained eye, they can look identical. Yet they can be very, very different implementations.
Likewise, I could have different products and make them look and function relatively the same. What I’m talking about is the user experience, what appears to the console operator and how they interact with a tool.
That’s built by the integrator or the SCADA tech, the SCADA engineer. They’re building those features. They’re building the objects. They’re building the animations. All of that’s being built into the SCADA system.
I can have things that also look very different, even in the same platform, because they’re built by different engineers. Maybe I acquired a company that used ArchestrA. I’m using ArchestrA. I’m like, “They have ArchestrA. I have ArchestrA. That’ll be easy.” Maybe not, because it has to do with how the automation is implemented, how the automation is deployed.
That ends up becoming a really big challenge, particularly for larger operators where they have multiple operations of different types in different locations. What tends to happen is every one of those SCADA systems becomes unique, even if it’s from the same software vendor.
The other challenge is, once you define what you want to do with your alarming and you get your rationalization in place and now you need to implement these changes in SCADA, making changes in SCADA is hard.
My experience is this. Most operators have teams internally that allow them to maintain the SCADA system. That’s add points, drop points, add sites, drop sites, make small changes to the HMI, that sort of thing.
They do not typically have resources internally to make major modifications to the SCADA system. There’s exceptions to that, but making those SCADA changes can be very difficult.
If you start changing animation on screens, given the way the control room management rules work, I can end up driving lots of need for point-to-point. Point-to-point is very expensive and time-consuming. Making changes is not easy. It’s expensive. It’s time-consuming. It’s got to be well-planned.
What we’re finding is, most of the operators we go and work with, they do not have a mechanism to segregate alarms and alerts. They don’t have a mechanism to handle multiple alarm severities.
We’re also finding a lot of operators that have very limited animation implemented in the HMI. In some cases, that animation is not possible because it’s a limitation of the software they’re using. In other cases, it’s just never been implemented. The capability exists in the software, but the SCADA engineers have never built it fully out that way.
All of this conversation is driving to another point I want to make. That is that SCADA integration and building an operations management system are different skill sets. SCADA integration is highly technical.
It’s all about getting the automation in place, getting the telecommunications in place, wiring up the data to the actual engine in the SCADA system. All of that is typical hardcore technical SCADA integration.
SCADA integration is a very different skill set than operations management. Operations management is looking at how is the HMI built, what automations and animations are available in the HMI, how does it support accomplishment of very specific tasks, and then how do those tasks relate to managing operating risk.
That’s a very different conversation than all the technical things to wire up the data. One is a “How do I optimize tasks? How do I optimize safety?” That’s a human factors or an operations management conversation. That is distinct and different from the SCADA integration, which is more of a data-centric getting all the data from the field to some place so I can display it.
Most companies have some SCADA integration capability internally. Few companies have operations management or human factors capabilities internally. These are some of the things we’re finding that are causing some challenges.
All of that is really just laying out a problem statement, if you will, about current state of alarm management and things you need to be thinking about as you’re going through it. What we’ve come up with is an approach. I think this approach is very useful.
The first thing we focus on is philosophy and specification. You can consider the API standards, API 1167 for alarm management, API 1165 for HMI, guidelines or philosophy about how you ought to approach these things. Typically, on the HMI side, you have a philosophy and a style guide. The style guide says, “This is exactly what we build and how we build it in our implementation.”
I think you need to do the same thing on the alarm side. You need an alarm philosophy which says, “This is how we approach it.” Then you need an alarm management plan that specifies, “This is how we’re going to do rationalization. This is how we’re going to implement rationalization in the SCADA system.”
“This is our master alarm database where all of our rationalization and documentation is stored, maintained, and used to govern implementation. These are the specific capabilities we need in the SCADA system and in the HMI to enable our approach for alarm management.”
Then you got to specify the governance. Once you get all this in place, you need pretty strong governance, or what will happen over time is that the technical infrastructure will decay, things like I’ll pull a site offline. I’ll take it out of service, but I never remove all that information out of SCADA. It’s still there.
Or I make a change, but I don’t document that change. Or I make it at a site, and I don’t make it at all sites. Those things tend to happen if you don’t have really good governance. Our approach is philosophy, which is generic, and specification. Specification needs to define governance. Those are key elements to an effective program.
What’s the opportunity? Couple things. One, there’s an opportunity to take what’s done in process safety management around facilities and what’s done around alarm management for control room management with pipelines and unify those processes. There’s a lot of alignment. There are distinctions, but there’s a lot of alignment.
One of the biggest challenges of doing effective alarm management and effective, high-performance HMI is it requires a disciplined, thoughtful design and then a maintenance of that design over a long period of time by a team.
If you do PSM, that context building that you do for PSM is very useful. It’s very much the same thing you do for alarm management. As you’re doing those efforts, if you can unify those processes, there’s a lot of cost savings to be had. There’s a lot of time savings to be had. There’s a lot of opportunity for improving how you’re able to operate these facilities.
One of the things with alarm management is we build things and we do alarm management as an afterthought. It really needs to be done up front when we’re doing the engineering.
The other thing I would say is that building a capability or a mechanism inside the company where you have a focus on engineering for human performance in the control room, that has a really large opportunity to transform and improve operations.
Couple takeaways. What we’re seeing is probably, of all the operators we work with, maybe 10 or 20 percent are doing a reasonably good job and are continuing to make progress. Everybody else is struggling with some of these issues that I’ve just laid out.
The industry’s made a huge amount of progress in the last 12 years, since the publishing and requirement for control room management.
We really don’t have a proper understanding yet about the distinction between what are the capabilities I need to do a good job of SCADA and managing the data and governing that versus what are the capabilities I need for effective human performance as implemented through alarm management and HMI management. We have a lot of opportunity and a long way to go.
Hope that’s helpful for you as listeners. I hope you enjoyed this…Not conversation this time because it’s just me talking. I hope you find this useful and helpful.
Just a reminder before you go. You should register to win our customized Pipeliners Podcast YETI tumbler. Simply visit pipelinepodcastnetwork.com/win and enter yourself in the drawing.
If you’d like to support this podcast, please leave us a review. You can do that on Apple Podcast, Google Play, Spotify, wherever you happen to listen. You can find instructions at pipelinepodcastnetwork.com.
[background music]
Russel: If you have ideas, questions, or topics you’d be interested in, please let me know, either on the Contact Us page on pipelinepodcastnetwork.com or reach out to me on LinkedIn. Thanks for listening. I’ll talk to you next week.
[music]


