Glossary

Digital Infrastructure Glossary

Definitions of common terms used in networking, cloud, data centres, resilience and infrastructure operations.

Anycast
A routing method in which the same network address is announced from multiple locations and traffic is directed according to routing conditions.
Autonomous system
A network or group of networks operated under a common routing policy and identified by an autonomous system number.
Availability
The proportion of a defined period during which a service is usable according to stated measurement rules.
Availability zone
A cloud-provider failure boundary within a region, usually designed with separate facility infrastructure.
Backup
A recoverable historical copy of data or system state kept for restoration.
BGP
Border Gateway Protocol, the routing protocol used by autonomous systems to exchange internet reachability information.
Blast radius
The scope of services, users or components affected by a failure or change.
Capacity headroom
Unused capacity reserved for demand growth, bursts, maintenance and failure states.
Cloud region
A geographic area in which a cloud provider operates one or more availability zones and related services.
Colocation
A model in which customer-controlled equipment is installed in a provider-operated data centre.
Common-mode failure
One event or dependency that affects components believed to be redundant.
Content delivery network
A distributed system that serves cached or optimized content from locations closer to users.
Data centre
A facility designed to house computing, storage and network equipment with supporting power, cooling and security.
Disaster recovery
The people, procedures, infrastructure and data capabilities used to restore services after a serious disruption.
Edge computing
Processing placed closer to users, equipment or data sources rather than only in a central facility or cloud region.
Failover
The movement of service from a failed or unavailable component to an alternate component or location.
Failure domain
A set of resources that can be affected by the same failure event.
Fault tolerance
The ability to continue operating through a defined fault with little or no visible interruption.
High availability
An architecture and operating approach intended to limit service interruption through redundancy, monitoring, failover and testing.
Hypervisor
Software or firmware that creates and manages virtual machines on physical hardware.
Internet exchange point
A shared facility or switching platform where networks exchange traffic.
Latency
The time required for data or a request to travel through a system and produce a response.
Observability
The ability to understand a system’s internal behavior from the evidence it produces, such as metrics, logs and traces.
Peering
A direct traffic-exchange relationship between networks.
Replication
The process of maintaining copies of data in more than one location or system.
Recovery point objective
The maximum acceptable amount of recent data exposure, expressed as a time interval.
Recovery time objective
The target time to restore a service after disruption.
Route diversity
The use of network paths that avoid relevant common physical and upstream failure points.
Service level objective
An internal target for a measurable service outcome such as availability or latency.
Technical debt
Future cost or risk created by short-term technical choices, deferred maintenance or unsupported systems.
Transit
A paid service through which one network obtains reachability to the broader internet through another network.
UPS
An uninterruptible power supply that conditions power and bridges interruptions while longer-duration power is established.
Virtual machine
A software-defined computer environment that runs on a physical host through a hypervisor.