Glossary
Digital Infrastructure Glossary
Definitions of common terms used in networking, cloud, data centres, resilience and infrastructure operations.
- Anycast
- A routing method in which the same network address is announced from multiple locations and traffic is directed according to routing conditions.
- Autonomous system
- A network or group of networks operated under a common routing policy and identified by an autonomous system number.
- Availability
- The proportion of a defined period during which a service is usable according to stated measurement rules.
- Availability zone
- A cloud-provider failure boundary within a region, usually designed with separate facility infrastructure.
- Backup
- A recoverable historical copy of data or system state kept for restoration.
- BGP
- Border Gateway Protocol, the routing protocol used by autonomous systems to exchange internet reachability information.
- Blast radius
- The scope of services, users or components affected by a failure or change.
- Capacity headroom
- Unused capacity reserved for demand growth, bursts, maintenance and failure states.
- Cloud region
- A geographic area in which a cloud provider operates one or more availability zones and related services.
- Colocation
- A model in which customer-controlled equipment is installed in a provider-operated data centre.
- Common-mode failure
- One event or dependency that affects components believed to be redundant.
- Content delivery network
- A distributed system that serves cached or optimized content from locations closer to users.
- Data centre
- A facility designed to house computing, storage and network equipment with supporting power, cooling and security.
- Disaster recovery
- The people, procedures, infrastructure and data capabilities used to restore services after a serious disruption.
- Edge computing
- Processing placed closer to users, equipment or data sources rather than only in a central facility or cloud region.
- Failover
- The movement of service from a failed or unavailable component to an alternate component or location.
- Failure domain
- A set of resources that can be affected by the same failure event.
- Fault tolerance
- The ability to continue operating through a defined fault with little or no visible interruption.
- High availability
- An architecture and operating approach intended to limit service interruption through redundancy, monitoring, failover and testing.
- Hypervisor
- Software or firmware that creates and manages virtual machines on physical hardware.
- Internet exchange point
- A shared facility or switching platform where networks exchange traffic.
- Latency
- The time required for data or a request to travel through a system and produce a response.
- Observability
- The ability to understand a system’s internal behavior from the evidence it produces, such as metrics, logs and traces.
- Peering
- A direct traffic-exchange relationship between networks.
- Replication
- The process of maintaining copies of data in more than one location or system.
- Recovery point objective
- The maximum acceptable amount of recent data exposure, expressed as a time interval.
- Recovery time objective
- The target time to restore a service after disruption.
- Route diversity
- The use of network paths that avoid relevant common physical and upstream failure points.
- Service level objective
- An internal target for a measurable service outcome such as availability or latency.
- Technical debt
- Future cost or risk created by short-term technical choices, deferred maintenance or unsupported systems.
- Transit
- A paid service through which one network obtains reachability to the broader internet through another network.
- UPS
- An uninterruptible power supply that conditions power and bridges interruptions while longer-duration power is established.
- Virtual machine
- A software-defined computer environment that runs on a physical host through a hypervisor.