Servers need continuous power and controlled temperatures. Data-centre resilience therefore begins with electrical and mechanical systems long before software architecture becomes relevant.
1) Utility power and site feeds
A facility normally receives electrical service from a utility through one or more feeds. Two feeds are useful only when they are genuinely independent; separate cables that share a substation or upstream path may still fail together.
The site design should identify the actual source, switching arrangement, maintenance procedures and expected restoration process.
2) Uninterruptible power supplies
A UPS bridges short interruptions and conditions power while generators start or utility service is restored. Battery runtime is limited, so the UPS is not usually the long-duration backup by itself.
UPS capacity, battery age, maintenance bypasses and load distribution matter. A redundant design can still be vulnerable if both paths depend on one switchboard or if maintenance places the system on a single path.
3) Generators and fuel
Generators provide longer-duration backup. Their practical resilience depends on start reliability, fuel quantity, refuelling contracts, testing under load and the ability to operate during a wider regional emergency.
A generator that is tested only without load may not reveal problems that appear when the full electrical system transfers.
4) Power distribution to the rack
Power moves through switchgear, distribution panels, power distribution units and rack-level outlets before reaching equipment. Dual power supplies help only when each supply is connected to a separate and healthy path.
Capacity planning must include both normal load and failure conditions. If one side fails, the remaining path must be able to carry the transferred load.
5) Cooling and airflow
Cooling removes the heat created by servers and network equipment. Common designs use computer-room air handlers, chilled water, direct expansion systems or more specialized liquid cooling.
Airflow management is as important as raw cooling capacity. Hot-aisle and cold-aisle separation, blanking panels and sensor placement help prevent localized hot spots.
6) What to monitor
- Utility, UPS and generator status
- Battery condition and estimated runtime
- Temperature and humidity at multiple rack locations
- Cooling unit performance and alarms
- Fuel levels and refuelling readiness
- Power draw by circuit and rack
7) Resilience questions
Facility reviews should ask what happens during maintenance, not only during failure. Planned work often disables redundant paths and creates temporary single points of failure.
Document who receives alarms, who can access the site, how long backup systems can operate and which external suppliers are required during an extended event.
Operational review questions
Use these questions to connect the concept to a real service or environment:
- Are power, cooling and connectivity paths genuinely independent?
- What maintenance state creates a temporary single point of failure?
- How long can the site operate without utility service?
- Who receives alarms and can access the facility?
- What external suppliers are required during an extended disruption?
Related guides
How Data Centers Connect to the Internet
A plain-language explanation of how data centers connect to the internet through carriers, meet-me rooms, cross-connects, upstream networks, and exchange points.
Network & DeliveryAnycast Routing Explained — Why CDNs and DNS Work So Fast
A plain-language explanation of anycast routing and why it allows DNS providers, CDNs, and global platforms to deliver traffic from the nearest location.
Compute & StorageConsistency vs Availability Explained
A detailed, plain-language explanation of consistency vs availability in distributed systems, including trade-offs, real-world examples, and why systems cannot maximize both.
Cloud ArchitectureHow Cloud Regions and Availability Zones Actually Work
A clear, architecture-first explanation of how cloud regions and availability zones are designed, connected, and operated — and why they matter for resilience and latency.