When you stream a film, save a photo to the cloud or buy something online, your request lands in a data centre: a building packed with servers, network equipment and the power and cooling that keep them running day and night. This lesson walks through one, from the servers in a rack to the internet connection, and explains why nearly everything in it is built twice.
What a data centre is
A data centre (US spelling: data center) is a building designed to house computer servers and the network that connects them, safely and without interruption. It can be a small room run by a hospital, a whole building owned by a bank or a warehouse-sized hall run by a cloud provider. Some companies rent space in a colocation ("colo") data centre, where many customers put their own racks in a shared building with shared power and cooling.
💡 In simple terms: a data centre is to servers what a hospital is to patients. It provides a safe place, constant power, the right temperature and a team watching around the clock, with a backup for everything that could go wrong.
Why data centres exist
A business could keep its servers in a cupboard at the office, and many small businesses do. But that cupboard has problems:
- Power cuts: one outage and every service stops.
- Heat: servers produce a lot of heat, and a failed air conditioner can let them overheat within an hour.
- One internet line: if it fails, customers can't reach anything.
- Security: anyone with the cupboard key can unplug or steal a server.
- Growth: there is no room for the next hundred servers.
A data centre solves all of these together, for hundreds or thousands of servers at once. The goal is availability: the services stay up even when parts fail, which they eventually always do.
Where it sits
Users reach a data centre over the internet or over private links from company offices. Inside, traffic is described in two directions:
| Direction | What it is | Example |
|---|---|---|
| North-south | Traffic entering or leaving the data centre | Your browser loading a shop's home page |
| East-west | Traffic between servers inside the data centre | The web server asking the database server for prices |
One click from a user can cause dozens of east-west messages between web, application and database servers. In modern data centres, east-west traffic is usually far larger than north-south traffic, and that shapes how the network is built.
The building blocks
The computers that run websites, applications and databases. Each physical server usually runs many virtual machines.
Connect the servers. Each rack has its own top-of-rack switches, which are joined together by larger switches.
Connect the data centre to the internet, to company offices and to other data centres.
Decide which traffic may enter or leave the data centre, and which traffic may pass between groups of servers.
Share incoming requests across several identical servers, and stop sending requests to a server that fails.
Standard steel frames or cabinets that hold the equipment, its cables and its power strips.
Two independent power paths, batteries (UPS) and generators, so that the power never stops.
Cooling units and a hot and cold aisle layout that remove the heat all this equipment produces.
A small data centre, end to end
Here is a small but complete design, the kind a mid-sized company might run. Follow a request from the internet down to a server, then a conversation between servers.
- 1. North-south in: a customer's request for 203.0.113.80 arrives over ISP B, passes through the firewall and reaches the load balancer's virtual IP address.
- 2. Load balancing: the load balancer chooses a healthy web server (10.10.1.12) and forwards the request to it.
- 3. East-west: the web server sends a request to an app server, which queries a database. Every path between servers on different leaves is leaf → spine → leaf.
- 4. Redundancy: if ISP B, edge router 2, firewall 2 or spine 2 fails, traffic uses the other half of the design.
Servers and racks
A data centre server is a powerful computer with no screen or keyboard, built to slide into a rack. Most servers run virtualisation: software that splits one physical server into many virtual machines (VMs), each acting like a separate computer with its own IP address. One physical server might run thirty VMs.
Servers are built for redundancy too. A typical server has two network interface cards (NICs), each cabled to a different switch, and two power supplies (PSUs), each plugged into a different power strip.
A rack is a standard steel frame that holds equipment 19 inches (48.3 cm) wide. Its height is measured in rack units (U). One U is 1.75 inches (4.45 cm), and a common rack is 42U or 48U tall. Racks stand in long rows, and cables run above them in trays or under a raised floor.
Switches: top-of-rack and spine-leaf
Top-of-rack (ToR) switches
Every rack has its own switch, usually two, mounted at the top: the top-of-rack (ToR) switches. Each server connects to both of them with short copper or fibre cables at 10, 25 or 100 Gbps. Only a few long fibre links leave the rack, which keeps the cabling tidy.
Spine-leaf: the concept
The ToR switches must be connected to each other. Older data centres used a tree of three layers (access, aggregation and core), which was designed for north-south traffic. Today, most use a spine-leaf design, built for heavy east-west traffic:
- The leaf switches are the ToR switches, and servers plug into them. In larger designs, the firewalls, load balancers and edge routers also connect through leaves (often called border leaves). The diagram above draws them next to the spines to keep it simple.
- The spine switches are a few large, fast switches that connect only to leaves.
- Every leaf connects to every spine. Leaves never connect to each other, and spines never connect to each other.
- 1. Path via spine 1: server A to server B goes leaf, spine, leaf: always the same number of hops.
- 2. Path via spine 2: a second conversation can use the other spine, so the load is spread across all spines.
- 3. Spine 1 fails: every flow moves to spine 2. Capacity is halved, but nothing is cut off.
| Why spine-leaf? | What it means in practice |
|---|---|
| Predictable | Any two servers on different leaves are always exactly leaf → spine → leaf apart, so delay is low and consistent |
| All links used | Traffic is shared across all spines at once, instead of leaving backup links idle |
| Easy to grow | Need more server ports? Add a leaf. Need more bandwidth between racks? Add a spine |
| Fails gracefully | Losing one spine of four removes a quarter of the capacity, not the connection |
🎓 Going further (CCNA and beyond): how spines and leaves share traffic (routing protocols and ECMP, equal-cost multipath) and how servers in different racks can still share a subnet (overlays such as VXLAN) are advanced topics. For this course, remember the shape and the reasons for it. The CCNA LAN switching with redundant links lesson shows how older designs handled redundancy.
The edge: routers, firewalls and load balancers
Edge routers
The routers at the edge connect the data centre to the outside world. A professional data centre buys internet service from two or more ISPs over physically separate cables, so a digger cutting one cable doesn't cut the whole site off.
Firewalls
Firewalls sit behind the routers and allow only the traffic the services need: for a web shop, TCP 443 to the load balancer and nothing else from the internet. Inside, more firewall rules separate groups of servers, so the web tier can reach the app tier but not the database directly. Firewalls are deployed as a pair: one is active, and the other is ready to take over within seconds. The two share their list of open connections, so existing sessions survive a failover.
Learn more: Network Segmentation
Load balancers
A load balancer stands in front of a group of identical servers. Users connect to one address, the virtual IP address (VIP), and the load balancer passes each new connection to one of the real servers behind it.
- Health checks: every few seconds, the load balancer tests each server (for example, by requesting a web page). A server that fails the test gets no new connections until it recovers.
- Maintenance: engineers can take a server out of the pool, update it and put it back with no downtime for users.
- Scaling: during a busy season, engineers add more servers to the pool.
What's in the packets
Follow the request from the diagram above. The customer's public address (after NAT on their home router) is 198.51.100.77. Servers inside the data centre use private addresses.
| Leg | Source IP:port | Destination IP:port | Note |
|---|---|---|---|
| Internet → firewall | 198.51.100.77:51544 | 203.0.113.80:443 | Public addresses only; the firewall allows TCP 443 to the VIP |
| Load balancer → web server | 10.10.0.5:40112 | 10.10.1.12:443 | In a common setup, the load balancer uses its own inside address as the source |
| Web → app server | 10.10.1.12:38920 | 10.10.2.31:8080 | East-west, leaf → spine → leaf |
| App → database | 10.10.2.31:44410 | 10.10.3.15:5432 | East-west; internal firewall rules allow only this port |
At every routed (Layer 3) hop, the Ethernet frame gets new MAC addresses, while the IP addresses inside stay the same on each leg. (Each leg in the table is a separate connection, which is why the addresses differ between legs.)
Learn more: Different-Subnet Communication
Power: it must never stop
Servers can't survive even a short power cut: they reboot, and data that is being written or sent can be lost. Data centres build the power supply in layers:
| Part | What it does |
|---|---|
| Utility feeds | Mains power, often from two separate grid connections |
| UPS (uninterruptible power supply) | Large battery banks that take over instantly when mains power drops. They cover the time until the generators are running |
| Generators | Diesel (or gas) generators that start automatically and can run for hours or days on stored fuel |
| Transfer switch | Automatically moves the load from mains power to the generators and back |
| PDUs (power distribution units) | Power strips in the rack, usually two per rack (A and B), often with remote monitoring per socket |
| Dual PSUs | Each server has two power supplies, one on PDU A and one on PDU B |
Cooling: hot and cold aisles
Nearly every watt a server uses ends up as heat. A single rack can produce as much heat as several household ovens. Cooling systems remove it: CRAC (computer room air conditioner) units, or liquid cooling for very dense racks. The layout that makes air cooling work is the hot aisle / cold aisle design:
- Servers pull cool air in at the front and push hot air out of the back.
- Rows of racks are placed front to front, so that cool air is delivered into a shared cold aisle.
- The backs face each other across a hot aisle, where the hot air is collected and returned to the cooling units.
- Many sites add doors and roofs to the aisles (containment) so hot and cold air can't mix at all.
💡 Small detail, big effect: empty spaces in a rack are closed with blanking panels. Without them, hot air from the back leaks round to the front, and servers draw in their own exhaust air.
Redundancy everywhere
The rule in a data centre is: no single point of failure. A single point of failure is any one part whose failure stops the service. Designers go through every part and ask, "What if this breaks?"
| Part | How it is made redundant |
|---|---|
| Internet connection | Two or more ISPs, over separate cable routes into the building |
| Edge routers and firewalls | Pairs; one takes over if the other fails |
| Spine switches | Two or more, all in use at once |
| Top-of-rack switches | Two per rack; each server cabled to both |
| Server network cards | Two NICs bonded together (NIC teaming / bonding) |
| Servers | Several behind a load balancer; VMs restart on another physical host |
| Power | A and B feeds (2N), UPS, generators, dual PSUs |
| Cooling | N+1: at least one more cooling unit than needed |
| The whole site | A second data centre in another region, for disasters |
A real-world example
An online shop runs in a rented colocation data centre. On Black Friday morning, three things happen:
- 07:10: a road crew cuts the ISP A fibre outside. Traffic moves to ISP B within seconds, and customers notice nothing.
- 09:30: web server 2 crashes under load. It fails the load balancer's health check, so the load balancer stops sending users to it. Web servers 1 and 3 carry on.
- 14:00: the grid has a 40-second power cut. The UPS takes over instantly, the generators start, and the transfer switch moves the load. No server reboots.
Three failures, zero outages. That is what all the duplicated equipment is for. The hidden risk is that each failure leaves the shop with no spare in that area until it is repaired. That is why monitoring and fast repairs matter.
What happens when things fail
| Failure | With redundancy | Without it |
|---|---|---|
| One ToR switch dies | Servers keep working over their second NIC to the other ToR | The whole rack goes offline |
| One spine dies | Less capacity between racks, but everything stays connected | Racks can't talk to each other |
| Mains power fails | The UPS, then the generators, carry the load | Every server drops at once |
| A cooling unit fails | The spare unit (N+1) keeps temperatures normal | Servers overheat and shut down to protect themselves |
| Both paths fail together | An outage. This is why the A and B paths must be truly separate: different cables, routes and power circuits | |
Troubleshooting and useful commands
Data centre engineers use the same basic network commands as everyone else, plus one habit: check that the redundancy is still there. Redundancy hides failures, so a server can run for months on one NIC without anyone noticing, until the second one fails too.
$ cat /proc/net/bonding/bond0 Ethernet Channel Bonding Driver: v6.8.0-45-generic Bonding Mode: IEEE 802.3ad Dynamic link aggregation Transmit Hash Policy: layer3+4 (1) MII Status: up MII Polling Interval (ms): 100 Slave Interface: ens1f0 MII Status: up Speed: 25000 Mbps Duplex: full Link Failure Count: 0 Permanent HW addr: 02:00:00:00:01:10 Slave Interface: ens1f1 MII Status: down Speed: Unknown Duplex: Unknown Link Failure Count: 3 Permanent HW addr: 02:00:00:00:01:11
bond0. The top-level MII Status: up means the bond is working, so the server is still online. But ens1f1 (the link to the second ToR switch) shows MII Status: down and a Link Failure Count of 3. The server has no redundancy left. Check that cable, its optic and the port on ToR switch B.$ ip -br link lo UNKNOWN 00:00:00:00:00:00 <LOOPBACK,UP,LOWER_UP> ens1f0 UP 02:00:00:00:01:10 <BROADCAST,MULTICAST,SLAVE,UP,LOWER_UP> ens1f1 DOWN 02:00:00:00:01:10 <NO-CARRIER,BROADCAST,MULTICAST,SLAVE,UP> bond0 UP 02:00:00:00:01:10 <BROADCAST,MULTICAST,MASTER,UP,LOWER_UP>
ens1f1 is DOWN with NO-CARRIER (no signal on the link), while bond0 is still UP. Both members show the bond's MAC address, because a bond presents one MAC address to the network.| Symptom | First checks |
|---|---|
| One server unreachable | Are both NIC links up? Check ip -br link, the bond status and the ToR port status |
| A whole rack unreachable | Are both ToR switches up? Are their uplinks to the spines up? Are the rack PDUs on? |
| Website slow, servers look fine | Load balancer health checks: is the pool down to one server? |
| Web tier reachable, database not | Internal firewall rules between tiers |
| Random reboots in one row | Check power (PDU and UPS alarms) and temperature (cooling) before the network |
Common mistakes
The server has two power supplies but only one power path, so one tripped breaker takes it down.
This protects against a bad cable, but not against the switch failing or being upgraded.
Two ISP cables that share a route can be cut by the same digger.
A failed backup that nobody has noticed is no backup at all.
Hot exhaust air leaks to the front of the rack, and the servers overheat.
Generators that have never run under load, or firewall pairs that have never switched over, often fail on the day they are needed.
- A data centre houses servers and gives them constant power, cooling, security and connectivity.
- Servers sit in racks, connected to two top-of-rack (leaf) switches; every leaf connects to every spine.
- North-south traffic enters and leaves the data centre; east-west traffic flows between servers and is usually larger.
- The edge has redundant routers (with two ISPs), firewall pairs and load balancers with a virtual IP address.
- Power: A/B feeds, UPS, generators, two PDUs and dual PSUs. Cooling: hot and cold aisles, N+1 units.
- Build in redundancy everywhere, avoid any single point of failure and monitor the spares.
Check yourself
In a spine-leaf network, server A is on leaf 1 and server B is on leaf 4. What path does traffic between them take?
The grid power fails. What keeps the servers running in the seconds before the generators start?
A web server behind a load balancer crashes. What do users notice?
A web server asks a database server in another rack for data. What kind of traffic is this?
A server's bond shows one NIC up and one down, but the server is working normally. What should you do?
Related lessons
Next, see how the same ideas look in software in Cloud networking basics. The other lessons below revisit the devices, the shapes networks take and how server groups are separated.
Learn more: SwitchesRoutersFirewallsNetwork TopologiesNetwork Segmentation