Need help selecting the right controls? Talk to our specialists — response within 24 hours.

What To Do When Your Commercial HVAC System Fails: 8 Questions to Ask

So your HVAC system just failed. Here are the questions you need to ask.

I’ve been on the other end of those panicked calls for over a decade. A chiller goes offline at 2 PM on a Friday. A data center rooftop unit starts throwing error codes during a heatwave. The building manager is asking how fast we can respond.

In my role coordinating emergency service for commercial facilities (this was back in 2022, when supply chains were still a mess), I learned that the first questions you ask determine whether you’re back online in 12 hours or stranded until Monday.

Below are the most common questions I get from facility managers when things go sideways. Plus one they should be asking but usually don’t.

1. My rooftop unit just stopped cooling. What do I do first?

Don’t panic, but act fast. The very first thing: check the thermostat setpoint. You'd be surprised how often someone accidentally bumped it into heat mode. If that’s fine, look for obvious signs: a tripped breaker, a frozen coil (ice on the outdoor unit), or a blinking error code on the controller.

The next step is to call a qualified service provider. If you’re under a maintenance contract—which, if you’re a commercial building manager and you don’t have one, that’s its own conversation—call your provider immediately. They have dispatch prioritization. If you’re on a time-and-materials basis, expect a longer wait.

Critical: If the issue is a compressor failure on a rooftop unit or a chiller, do NOT attempt to restart it repeatedly. You can cause secondary damage to the system.

2. Our chiller is down and we have a temperature-critical process (like a data center). What's the play?

This is about triage, not repair. The immediate goal is temperature stability, not fixing the chiller.

First, is there redundant capacity? Most data centers have N+1 cooling. If so, the failed chiller is isolated, and you have time. If you’re running at full capacity (N), you have a much smaller window—typically measured in minutes to hours before server inlet temps spike.

We had a situation in July 2024 where a data center's primary chiller failed at 3 PM on a Thursday. Normal turnaround for a major repair is 3-5 days. The client needed cooling that night. We paid a premium for an emergency rental chiller—about $3,500 for a 48-hour rental plus expedited freight—but the alternative was a potential $50,000-per-hour revenue loss from server downtime. The upside was massive. The risk was doing nothing.

Bottom line: For mission-critical cooling, the time to plan for a chiller failure is before it happens. Pre-negotiate a rental agreement. Have a load-shedding plan. That said, if you’re in it now, the first call is to a chiller rental company (Carrier, Trane, Johnson Controls all have programs).

3. How do I know if I should repair my old HVAC system or replace it?

This is the classic Hobson's choice during an emergency. The unit is dead, parts are on backorder, and you’re tempted to just replace the whole thing. Or, the repair is expensive, and you wonder if you're throwing good money after bad.

Most buyers focus on the immediate repair cost and completely miss the total cost of ownership over the next 3-5 years. The question everyone asks is, 'How much is the repair?' The question they should ask is, 'What's the remaining useful life of the equipment, and what's the efficiency delta compared to a modern replacement?'

Here’s my rough rule of thumb (based on internal data from about 200 replacement decisions):

  • System is under 10 years old: Repair unless the repair cost exceeds 50% of a new unit's price.
  • System is 10-15 years old: Repair only if the fix is simple (like a contactor or capacitor) and the unit is otherwise healthy.
  • System is over 15 years old: Strongly consider replacement. The R-22 refrigerant phase-out alone is a deal-breaker for many older systems.

4. Is emergency HVAC service really worth the extra cost?

This is where my perspective is pretty clear. Yes, but not because the hourly rate is justified. You’re not paying for the technician’s time. You’re paying for the certainty that someone will show up.

Standard service might have a 3-5 day wait. Emergency service guarantees a response within 4 hours—sometimes same-day. That premium (often 1.5x to 2x the standard rate) buys you a slot on the dispatcher’s priority board. It buys you a technician who might be called away from another job.

An example: In March 2024, we had a client whose main chiller went down on a Saturday morning. Normal service wasn’t available until Tuesday. Their warehouse needed cooling for stored goods (perishable inventory worth about $15,000). We paid $400 extra for an emergency callback fee. The alternative was a complete loss of inventory. The math was a no-brainer.

So, when you call and ask about emergency rates, don't just think of the hourly rate. Think of the cost of not having cooling for another three days.

5. I have a Johnson Controls Metasys system. Can any HVAC contractor work on it?

Technically, any licensed contractor can physically work on the mechanical components (compressors, fans, coils). But for the controls—the brain of the system—it’s different.

Metasys is a proprietary building automation system. To properly diagnose and repair issues at the controller level, a technician needs specific training and access to software tools. A general HVAC tech can change a failed actuator, but they can’t re-program the sequence of operation or troubleshoot a network communication fault.

In my experience, this is a point of friction. Many facility managers call a local shop that handles generic Trane or Carrier units, and are surprised when the tech can’t figure out why the VAV box isn’t responding to the central system. Then they have to call us, and we’re dealing with a 2-hour diagnostic fee on top of the original trip charge.

Recommendation: If you have a Metasys or any JCI controls platform, make sure your service provider has a certified controls technician on staff. It’s a non-negotiable.

6. How long does it take to get a replacement compressor for a chiller?

This has been the source of a lot of gray hair since 2020. As of January 2025, lead times for major chiller components (compressors, microchannel condensers, ECM motors) are still extended compared to pre-pandemic levels.

For a common chiller model (like a Carrier 30RB or a Trane CGAM), a new compressor might be 2-3 days if it’s in a regional warehouse. For older models or less common brands, you’re looking at 1-2 weeks—or more if it’s a custom-built unit for a data center.

This is where the 'time certainty' argument gets real. If you have an older chiller, I strongly recommend having a critical spares agreement with your service provider (note to self: I really should formalize this for our own clients). Paying a small annual fee to have a compressor or a control board on consignment at the local branch can turn a 2-week outage into a 6-hour repair.

7. What are the most common mistakes building managers make during an HVAC emergency?

Three things, in order.

First: Not having a plan. Most buildings don’t have a documented emergency shutdown or isolation procedure. When a refrigerant leak is detected, people don’t know which valve to close first. That costs time and money.

Second: Underestimating the lead time for parts. They assume a compressor is on the shelf at the local supply house. It’s not. Especially for larger tonnage units (above 50 tons), components are typically stocked regionally, not locally.

Third: Not asking the 'what if' question during routine maintenance. 'The unit has been running fine for years' is not a maintenance strategy. It’s a recipe for an unplanned replacement. This was true 15 years ago when controls were simpler. Today, with complex systems like VRF (Variable Refrigerant Flow) and building analytics platforms, the 'set it and forget it' approach is a deal-breaker.

8. (The question you probably haven't asked) What's the 'time to failure' for my emergency response?

Most people ask 'how fast can you get here?' The better question is: 'What’s the average time to restore cooling once your technician is on-site?'

I’ve seen a 2-hour dispatch time turn into a 12-hour repair because the first technician was a basic install guy, not a diagnostics specialist. The real metric isn't arrival time. It's Mean Time To Repair (MTTR).

When you’re evaluating a service provider, ask them for their MTTR for rooftop units and chillers. A good provider should be able to give you a ballpark: 'For a standard RTU with a simple component failure, 4-6 hours. For a chiller with a compressor failure, 2-3 days including parts procurement.' If they can't answer, that’s a red flag.

Leave a Reply