Servers need continuous power and controlled temperatures. Data-centre resilience therefore begins with electrical and mechanical systems long before software architecture becomes relevant.

1) Utility power and site feeds

A facility normally receives electrical service from a utility through one or more feeds. Two feeds are useful only when they are genuinely independent; separate cables that share a substation or upstream path may still fail together.

The site design should identify the actual source, switching arrangement, maintenance procedures and expected restoration process.

2) Uninterruptible power supplies

A UPS bridges short interruptions and conditions power while generators start or utility service is restored. Battery runtime is limited, so the UPS is not usually the long-duration backup by itself.

UPS capacity, battery age, maintenance bypasses and load distribution matter. A redundant design can still be vulnerable if both paths depend on one switchboard or if maintenance places the system on a single path.

3) Generators and fuel

Generators provide longer-duration backup. Their practical resilience depends on start reliability, fuel quantity, refuelling contracts, testing under load and the ability to operate during a wider regional emergency.

A generator that is tested only without load may not reveal problems that appear when the full electrical system transfers.

4) Power distribution to the rack

Power moves through switchgear, distribution panels, power distribution units and rack-level outlets before reaching equipment. Dual power supplies help only when each supply is connected to a separate and healthy path.

Capacity planning must include both normal load and failure conditions. If one side fails, the remaining path must be able to carry the transferred load.

5) Cooling and airflow

Cooling removes the heat created by servers and network equipment. Common designs use computer-room air handlers, chilled water, direct expansion systems or more specialized liquid cooling.

Airflow management is as important as raw cooling capacity. Hot-aisle and cold-aisle separation, blanking panels and sensor placement help prevent localized hot spots.

6) What to monitor

  • Utility, UPS and generator status
  • Battery condition and estimated runtime
  • Temperature and humidity at multiple rack locations
  • Cooling unit performance and alarms
  • Fuel levels and refuelling readiness
  • Power draw by circuit and rack

7) Resilience questions

Facility reviews should ask what happens during maintenance, not only during failure. Planned work often disables redundant paths and creates temporary single points of failure.

Document who receives alarms, who can access the site, how long backup systems can operate and which external suppliers are required during an extended event.

Operational review questions

Use these questions to connect the concept to a real service or environment:

  • Are power, cooling and connectivity paths genuinely independent?
  • What maintenance state creates a temporary single point of failure?
  • How long can the site operate without utility service?
  • Who receives alarms and can access the facility?
  • What external suppliers are required during an extended disruption?