-
Step 1: Identify Your Critical Systems and Their Redundancies (Or Lack Thereof)
-
Step 2: Pre-Quote and Budget for Emergency Services
-
Step 3: Create a 'Rapid Response' Vendor Sheet
-
Step 4: Define Your Own 'Emergency Threshold'
-
Step 5: Train Your Team (Not Just for Basics, but for Emergencies)
-
Common Mistakes to Avoid
In my role coordinating urgent repairs for commercial buildings, I've seen a lot of panic. A chiller goes down in July. A data center cooling system starts throwing alarms on a Friday afternoon. A pool heater fails right before a big event. And suddenly, the person in charge is scrambling, calling vendors they've never worked with, and paying a premium for a solution that may or may not arrive in time.
I've handled over 200 of these emergency situations in the last 8 years. If I remember correctly, about 60% of them could have been avoided or made significantly less painful with a simple, actionable plan. This checklist is that plan. It's designed for anyone responsible for a critical system—a Johnson Controls chiller, a building automation system, or even a single freezer—who cannot afford a prolonged breakdown.
Here are the 5 steps you need to take, ideally now, before the alarm goes off.
Step 1: Identify Your Critical Systems and Their Redundancies (Or Lack Thereof)
This sounds obvious, but you'd be surprised how many facility managers don't have a clear, written list. The first thing I ask when I get a panicked call is: 'What happened, and what's the backup?'
You need to walk through your facility and answer these questions for every piece of equipment that would cause a major business disruption if it failed:
- What is its specific role? (e.g., 'This Johnson Controls YCIVL chiller provides chilled water to air handlers in the west wing.')
- What is the current, documented failure mode? (e.g., 'No backup chiller. Single point of failure.')
- Can the load be shifted or shed? (e.g., 'Can we shut down the west wing and move staff to the east?')
- What is the maximum tolerable downtime? (e.g., 'Data center: 30 minutes. Office space: 4 hours. Cold storage: 2 hours.')
Don't just think this through—write it down. I once had a client who 'knew' their critical freezer had a backup generator. When the power went out, they discovered the generator hadn't been tested in 3 years and the fuel had been siphoned. That's a $20,000 mistake in lost product that a simple checklist would have caught.
Step 2: Pre-Quote and Budget for Emergency Services
This is the step that separates the pros from the panic-stricken. In normal procurement, you get three quotes and pick the cheapest. In an emergency, you don't have that luxury. In my experience managing over 200 rush orders, the lowest quote has cost us more in 60% of cases.
Let me tell you about a time I ignored my own rule. A client called at 4 PM on a Friday needing a replacement heat pump for a school. The normal lead time was two weeks. I found a vendor who offered a 'rush job' for only $200 more than the base cost. I went with them. The unit arrived on Tuesday, but it was the wrong voltage. We paid $800 in overtime for a contractor to install it anyway and had to eat the cost of replacing a fried board in a different part of the system. That $200 savings turned into a $1,500 problem.
Here's what to do before an emergency:
- Call your go-to vendors (like a Johnson Controls authorized service provider) and ask for their emergency service rates and response time SLAs. For example, is there a 24/7 number? What's the premium for a weekend call-out? Is there a guaranteed response time for a critical chiller?
- Identify a backup vendor. I always recommend having one or two alternatives. Even the best vendors can be overwhelmed.
- Create a simple budget line. This doesn't have to be a huge number. I worked with a manufacturing plant that kept a $5,000 'emergency repair fund' each quarter. They almost never used it, but when they did—like the time their main pool heater failed during a product test—it saved them from a production delay that would have cost $12,000.
View this budget as an insurance premium, not an expense. The best part of having this pre-approved: no frantic calls to accounting at 3 AM for approval on a $3,000 repair.
Step 3: Create a 'Rapid Response' Vendor Sheet
Don't make your team search for phone numbers when the s**t hits the fan. I learned this the hard way. After three failed rush orders where the contact person was 'on vacation,' we now use a simple, always-accessible sheet.
For each of your top two vendors, have this information ready:
- Primary & Secondary Contacts: Not just a general line, but names and cell phone numbers. 'We only use the sales engineer for our account, not the main switchboard.'
- After-Hours Procedure: Is there a specific email or text line? Do they have a technician on call?
- Parts Availability: Do they stock common parts for your specific equipment (e.g., a control board for a Metasys system)? The most expensive part of an emergency is often expedited shipping.
- Preferred Communication: 'When I'm triaging a rush order, I need to send photos and specs via text, not email.'
Keep a physical copy of this sheet near the equipment and a digital copy accessible from a phone. In March 2024, I got a call from a junior technician at a data center. The main cooling fan had shut down. He couldn't find the vendor's number. By the time he called me, and I found it, three servers had already throttled down. A simple, printed sheet next to the variable frequency drive would have saved 20 minutes of panic and a lot of explaining to management.
Step 4: Define Your Own 'Emergency Threshold'
Here's a nuance that most people miss. Not every alarm requires a full emergency response. You need to define what constitutes a 'can wait until Monday' problem versus a 'wake up the CEO at 2 AM' problem.
To some extent, this depends on your facility. But here's a framework I've used:
- Level 1 (Minor issue): A single, non-critical zone is uncomfortable. No business impact. Action: Log it and call for service within 48 hours.
- Level 2 (Significant issue, can wait): A secondary pump fails, but the primary is working. Business impact is possible but not immediate. Action: Call for service within 24 hours; start monitoring.
- Level 3 (Moderate emergency): A critical fan for a server room fails, but the backup is working. Business interruption is imminent if not resolved within hours. Action: Call the on-call technician; escalate internally.
- Level 4 (Full emergency): The main chiller is down in a July heatwave. The backup for a data center is offline. Business is shutting down. Action: Activate the full emergency plan; call the vendor's 24/7 line; authorize overtime.
Take this with a grain of salt: your thresholds will be different. But having them pre-defined is key. It prevents your team from overreacting to a simple alarm (and wasting money) or under-reacting to a critical failure.
Step 5: Train Your Team (Not Just for Basics, but for Emergencies)
Finally, and this is the most overlooked step: drill the plan. I don't mean a full-scale simulation that shuts down the building. I mean a simple 30-minute dry run.
Gather the team. Hand them the vendor sheet. Give them a scenario. 'It's 5:30 PM on Friday. The main freezer's alarm is ringing. What do you do?'
Don't hold me to this, but based on our internal data from 200+ rush jobs, about 70% of first-time responses are wrong. People call the wrong person, wait too long, or authorize the wrong fix. A quick training session, repeated once a year, changes that.
After 8 years of cleaning up other people's emergencies, I've come to believe that the biggest risk isn't the equipment failing. It's the plan (or lack of one) that's failing. Spend a couple of hours building this checklist and reviewing it with your team. You'll save yourself a lot of stress, a lot of money, and maybe even a couple of gray hairs.
Common Mistakes to Avoid
- Only having one vendor on speed dial. What happens if they're swamped? Always have a backup.
- Not checking parts availability. A skilled technician is useless without the right motor or control board.
- Relying on memory. The best plan is worthless if it's not written down and accessible.
- Forgetting about the 'human factor.' Your on-call person might be on their commute or have a bad cell signal. Make sure multiple people know the plan.
Pricing is for general reference only. Actual costs and specific equipment capabilities vary by vendor, location, and system specifications. Always verify current contact information and SLAs directly with your service providers.