Routelearn.net
Course menu

Unit 15: Modern Networking IntroductionLesson 15.1 (1 of 2 in this unit)83 of 84 in the Network Fundamentals course

Data Centre Networking Basics

A data centre is a building full of servers, designed so that its services never stop. Learn the parts that make it work (racks, top-of-rack and spine-leaf switches, routers, firewalls, load balancers, power and cooling) and why almost everything comes in pairs.

Beginner · 16 min read · Before this: Switches, Routers, Firewalls, Public and private IP addresses

A data centre network is the network inside a facility that houses servers and storage. It connects racks through top-of-rack and spine-leaf switches, links them to the outside world through routers, firewalls and load balancers, and duplicates links, devices and power so that a single failure does not stop service.

In simple terms: A data centre is a building full of servers. Its network connects them to each other and to the internet, with a spare for almost everything, so nothing stops when one part fails.

When you stream a film, save a photo to the cloud or buy something online, your request lands in a data centre: a building packed with servers, network equipment and the power and cooling that keep them running day and night. This lesson walks through one, from the servers in a rack to the internet connection, and explains why nearly everything in it is built twice.

What a data centre is

A data centre (US spelling: data center) is a building designed to house computer servers and the network that connects them, safely and without interruption. It can be a small room run by a hospital, a whole building owned by a bank or a warehouse-sized hall run by a cloud provider. Some companies rent space in a colocation ("colo") data centre, where many customers put their own racks in a shared building with shared power and cooling.

💡 In simple terms: a data centre is to servers what a hospital is to patients. It provides a safe place, constant power, the right temperature and a team watching around the clock, with a backup for everything that could go wrong.

Why data centres exist

A business could keep its servers in a cupboard at the office, and many small businesses do. But that cupboard has problems:

  • Power cuts: one outage and every service stops.
  • Heat: servers produce a lot of heat, and a failed air conditioner can let them overheat within an hour.
  • One internet line: if it fails, customers can't reach anything.
  • Security: anyone with the cupboard key can unplug or steal a server.
  • Growth: there is no room for the next hundred servers.

A data centre solves all of these together, for hundreds or thousands of servers at once. The goal is availability: the services stay up even when parts fail, which they eventually always do.

Where it sits

Users reach a data centre over the internet or over private links from company offices. Inside, traffic is described in two directions:

DirectionWhat it isExample
North-southTraffic entering or leaving the data centreYour browser loading a shop's home page
East-westTraffic between servers inside the data centreThe web server asking the database server for prices

One click from a user can cause dozens of east-west messages between web, application and database servers. In modern data centres, east-west traffic is usually far larger than north-south traffic, and that shapes how the network is built.

The building blocks

Servers

The computers that run websites, applications and databases. Each physical server usually runs many virtual machines.

Switches

Connect the servers. Each rack has its own top-of-rack switches, which are joined together by larger switches.

Routers

Connect the data centre to the internet, to company offices and to other data centres.

Firewalls

Decide which traffic may enter or leave the data centre, and which traffic may pass between groups of servers.

Load balancers

Share incoming requests across several identical servers, and stop sending requests to a server that fails.

Racks

Standard steel frames or cabinets that hold the equipment, its cables and its power strips.

Power

Two independent power paths, batteries (UPS) and generators, so that the power never stops.

Cooling

Cooling units and a hot and cold aisle layout that remove the heat all this equipment produces.

A small data centre, end to end

Here is a small but complete design, the kind a mid-sized company might run. Follow a request from the internet down to a server, then a conversation between servers.

InternetEdge router 1ISP AEdge router 2ISP BFirewall 1Firewall 2Load balancerVIP 203.0.113.80Spine 1Spine 2Leaf 1 (ToR)Rack 1Leaf 2 (ToR)Rack 2Leaf 3 (ToR)Rack 3Web servers10.10.1.11–13App servers10.10.2.0/24Databases10.10.3.0/24
  1. 1. North-south in: a customer's request for 203.0.113.80 arrives over ISP B, passes through the firewall and reaches the load balancer's virtual IP address.
  2. 2. Load balancing: the load balancer chooses a healthy web server (10.10.1.12) and forwards the request to it.
  3. 3. East-west: the web server sends a request to an app server, which queries a database. Every path between servers on different leaves is leaf → spine → leaf.
  4. 4. Redundancy: if ISP B, edge router 2, firewall 2 or spine 2 fails, traffic uses the other half of the design.

Servers and racks

A data centre server is a powerful computer with no screen or keyboard, built to slide into a rack. Most servers run virtualisation: software that splits one physical server into many virtual machines (VMs), each acting like a separate computer with its own IP address. One physical server might run thirty VMs.

Servers are built for redundancy too. A typical server has two network interface cards (NICs), each cabled to a different switch, and two power supplies (PSUs), each plugged into a different power strip.

A rack is a standard steel frame that holds equipment 19 inches (48.3 cm) wide. Its height is measured in rack units (U). One U is 1.75 inches (4.45 cm), and a common rack is 42U or 48U tall. Racks stand in long rows, and cables run above them in trays or under a raised floor.

PDU A
Rack · front view
1UPatch panel (cables to other racks)
1UTop-of-rack switch A
1UTop-of-rack switch B
1UBlanking panel
2UServer 01 · 2 NICs · 2 PSUs
2UServer 02 · 2 NICs · 2 PSUs
2UServer 03 · 2 NICs · 2 PSUs
2UServer 04 · 2 NICs · 2 PSUs
4UStorage array
PDU B
A typical rack: two top-of-rack switches at the top, servers below, two power strips (PDUs) fed from different power paths. 1U = 1.75 inches (4.45 cm) of height.

Switches: top-of-rack and spine-leaf

Top-of-rack (ToR) switches

Every rack has its own switch, usually two, mounted at the top: the top-of-rack (ToR) switches. Each server connects to both of them with short copper or fibre cables at 10, 25 or 100 Gbps. Only a few long fibre links leave the rack, which keeps the cabling tidy.

Spine-leaf: the concept

The ToR switches must be connected to each other. Older data centres used a tree of three layers (access, aggregation and core), which was designed for north-south traffic. Today, most use a spine-leaf design, built for heavy east-west traffic:

  • The leaf switches are the ToR switches, and servers plug into them. In larger designs, the firewalls, load balancers and edge routers also connect through leaves (often called border leaves). The diagram above draws them next to the spines to keep it simple.
  • The spine switches are a few large, fast switches that connect only to leaves.
  • Every leaf connects to every spine. Leaves never connect to each other, and spines never connect to each other.
Spine 1Spine 2Leaf 1Leaf 2Leaf 3Leaf 4Server A10.10.1.11Server B10.10.4.21
  1. 1. Path via spine 1: server A to server B goes leaf, spine, leaf: always the same number of hops.
  2. 2. Path via spine 2: a second conversation can use the other spine, so the load is spread across all spines.
  3. 3. Spine 1 fails: every flow moves to spine 2. Capacity is halved, but nothing is cut off.
Why spine-leaf?What it means in practice
PredictableAny two servers on different leaves are always exactly leaf → spine → leaf apart, so delay is low and consistent
All links usedTraffic is shared across all spines at once, instead of leaving backup links idle
Easy to growNeed more server ports? Add a leaf. Need more bandwidth between racks? Add a spine
Fails gracefullyLosing one spine of four removes a quarter of the capacity, not the connection

🎓 Going further (CCNA and beyond): how spines and leaves share traffic (routing protocols and ECMP, equal-cost multipath) and how servers in different racks can still share a subnet (overlays such as VXLAN) are advanced topics. For this course, remember the shape and the reasons for it. The CCNA LAN switching with redundant links lesson shows how older designs handled redundancy.

The edge: routers, firewalls and load balancers

Edge routers

The routers at the edge connect the data centre to the outside world. A professional data centre buys internet service from two or more ISPs over physically separate cables, so a digger cutting one cable doesn't cut the whole site off.

Firewalls

Firewalls sit behind the routers and allow only the traffic the services need: for a web shop, TCP 443 to the load balancer and nothing else from the internet. Inside, more firewall rules separate groups of servers, so the web tier can reach the app tier but not the database directly. Firewalls are deployed as a pair: one is active, and the other is ready to take over within seconds. The two share their list of open connections, so existing sessions survive a failover.

Learn more: Network Segmentation

Load balancers

A load balancer stands in front of a group of identical servers. Users connect to one address, the virtual IP address (VIP), and the load balancer passes each new connection to one of the real servers behind it.

Client connects to the VIP
198.51.100.77 → 203.0.113.80, TCP 443
Load balancer chooses a server
Round robin or least connections; skips unhealthy servers
Forwards to the real server
→ 10.10.1.12, TCP 443
Reply goes back through the load balancer
The client still sees 203.0.113.80 as the server
One request through a load balancer
  • Health checks: every few seconds, the load balancer tests each server (for example, by requesting a web page). A server that fails the test gets no new connections until it recovers.
  • Maintenance: engineers can take a server out of the pool, update it and put it back with no downtime for users.
  • Scaling: during a busy season, engineers add more servers to the pool.

What's in the packets

Follow the request from the diagram above. The customer's public address (after NAT on their home router) is 198.51.100.77. Servers inside the data centre use private addresses.

LegSource IP:portDestination IP:portNote
Internet → firewall198.51.100.77:51544203.0.113.80:443Public addresses only; the firewall allows TCP 443 to the VIP
Load balancer → web server10.10.0.5:4011210.10.1.12:443In a common setup, the load balancer uses its own inside address as the source
Web → app server10.10.1.12:3892010.10.2.31:8080East-west, leaf → spine → leaf
App → database10.10.2.31:4441010.10.3.15:5432East-west; internal firewall rules allow only this port

At every routed (Layer 3) hop, the Ethernet frame gets new MAC addresses, while the IP addresses inside stay the same on each leg. (Each leg in the table is a separate connection, which is why the addresses differ between legs.)

Learn more: Different-Subnet Communication

Power: it must never stop

Servers can't survive even a short power cut: they reboot, and data that is being written or sent can be lost. Data centres build the power supply in layers:

PartWhat it does
Utility feedsMains power, often from two separate grid connections
UPS (uninterruptible power supply)Large battery banks that take over instantly when mains power drops. They cover the time until the generators are running
GeneratorsDiesel (or gas) generators that start automatically and can run for hours or days on stored fuel
Transfer switchAutomatically moves the load from mains power to the generators and back
PDUs (power distribution units)Power strips in the rack, usually two per rack (A and B), often with remote monitoring per socket
Dual PSUsEach server has two power supplies, one on PDU A and one on PDU B
⛽ Diesel generators
Start within seconds; run for hours or days on stored fuel
Feed A
Utility feed A
Mains power from the grid
Transfer switch
Switches to the generator if mains fails
UPS A
Batteries bridge the gap: seconds to minutes
PDU A
Power strip in the rack
Feed B
Utility feed B
Ideally a second, separate supply
Transfer switch
Switches to the generator if mains fails
UPS B
A separate battery system
PDU B
Second power strip in the rack
Server
PSU 1 ← APSU 2 ← B
Two independent power paths (A and B). Either one can fail completely and the server keeps running on the other power supply (PSU).

Cooling: hot and cold aisles

Nearly every watt a server uses ends up as heat. A single rack can produce as much heat as several household ovens. Cooling systems remove it: CRAC (computer room air conditioner) units, or liquid cooling for very dense racks. The layout that makes air cooling work is the hot aisle / cold aisle design:

  1. Servers pull cool air in at the front and push hot air out of the back.
  2. Rows of racks are placed front to front, so that cool air is delivered into a shared cold aisle.
  3. The backs face each other across a hot aisle, where the hot air is collected and returned to the cooling units.
  4. Many sites add doors and roofs to the aisles (containment) so hot and cold air can't mix at all.
Seen from above
Hot aisle
↑ Hot exhaust air (about 35–45 °C) rises and returns to the cooling units
Rack row
front ↓
front ↓
front ↓
front ↓
front ↓
Cold aisle
❄ Cool air (about 18–27 °C) supplied here; both rows breathe it in
Rack row
front ↑
front ↑
front ↑
front ↑
front ↑
Hot aisle
↑ Hot exhaust air returns to the cooling units
Cooling unit (CRAC)hot in, cold out
Rows of racks face each other front to front and back to back. Servers pull cool air in at the front and blow hot air out of the back, so hot and cold air never mix.

💡 Small detail, big effect: empty spaces in a rack are closed with blanking panels. Without them, hot air from the back leaks round to the front, and servers draw in their own exhaust air.

Redundancy everywhere

The rule in a data centre is: no single point of failure. A single point of failure is any one part whose failure stops the service. Designers go through every part and ask, "What if this breaks?"

PartHow it is made redundant
Internet connectionTwo or more ISPs, over separate cable routes into the building
Edge routers and firewallsPairs; one takes over if the other fails
Spine switchesTwo or more, all in use at once
Top-of-rack switchesTwo per rack; each server cabled to both
Server network cardsTwo NICs bonded together (NIC teaming / bonding)
ServersSeveral behind a load balancer; VMs restart on another physical host
PowerA and B feeds (2N), UPS, generators, dual PSUs
CoolingN+1: at least one more cooling unit than needed
The whole siteA second data centre in another region, for disasters

A real-world example

An online shop runs in a rented colocation data centre. On Black Friday morning, three things happen:

  1. 07:10: a road crew cuts the ISP A fibre outside. Traffic moves to ISP B within seconds, and customers notice nothing.
  2. 09:30: web server 2 crashes under load. It fails the load balancer's health check, so the load balancer stops sending users to it. Web servers 1 and 3 carry on.
  3. 14:00: the grid has a 40-second power cut. The UPS takes over instantly, the generators start, and the transfer switch moves the load. No server reboots.

Three failures, zero outages. That is what all the duplicated equipment is for. The hidden risk is that each failure leaves the shop with no spare in that area until it is repaired. That is why monitoring and fast repairs matter.

What happens when things fail

FailureWith redundancyWithout it
One ToR switch diesServers keep working over their second NIC to the other ToRThe whole rack goes offline
One spine diesLess capacity between racks, but everything stays connectedRacks can't talk to each other
Mains power failsThe UPS, then the generators, carry the loadEvery server drops at once
A cooling unit failsThe spare unit (N+1) keeps temperatures normalServers overheat and shut down to protect themselves
Both paths fail togetherAn outage. This is why the A and B paths must be truly separate: different cables, routes and power circuits

Troubleshooting and useful commands

Data centre engineers use the same basic network commands as everyone else, plus one habit: check that the redundancy is still there. Redundancy hides failures, so a server can run for months on one NIC without anyone noticing, until the second one fails too.

Example output from a Linux server with two bonded NICs, written for this lesson
$ cat /proc/net/bonding/bond0
Ethernet Channel Bonding Driver: v6.8.0-45-generic

Bonding Mode: IEEE 802.3ad Dynamic link aggregation
Transmit Hash Policy: layer3+4 (1)
MII Status: up
MII Polling Interval (ms): 100

Slave Interface: ens1f0
MII Status: up
Speed: 25000 Mbps
Duplex: full
Link Failure Count: 0
Permanent HW addr: 02:00:00:00:01:10

Slave Interface: ens1f1
MII Status: down
Speed: Unknown
Duplex: Unknown
Link Failure Count: 3
Permanent HW addr: 02:00:00:00:01:11
What to look for: the server's two NICs are combined into one logical interface, bond0. The top-level MII Status: up means the bond is working, so the server is still online. But ens1f1 (the link to the second ToR switch) shows MII Status: down and a Link Failure Count of 3. The server has no redundancy left. Check that cable, its optic and the port on ToR switch B.
Example output from a Linux server, written for this lesson
$ ip -br link
lo               UNKNOWN        00:00:00:00:00:00 <LOOPBACK,UP,LOWER_UP>
ens1f0           UP             02:00:00:00:01:10 <BROADCAST,MULTICAST,SLAVE,UP,LOWER_UP>
ens1f1           DOWN           02:00:00:00:01:10 <NO-CARRIER,BROADCAST,MULTICAST,SLAVE,UP>
bond0            UP             02:00:00:00:01:10 <BROADCAST,MULTICAST,MASTER,UP,LOWER_UP>
What to look for: the same problem, in brief. ens1f1 is DOWN with NO-CARRIER (no signal on the link), while bond0 is still UP. Both members show the bond's MAC address, because a bond presents one MAC address to the network.
SymptomFirst checks
One server unreachableAre both NIC links up? Check ip -br link, the bond status and the ToR port status
A whole rack unreachableAre both ToR switches up? Are their uplinks to the spines up? Are the rack PDUs on?
Website slow, servers look fineLoad balancer health checks: is the pool down to one server?
Web tier reachable, database notInternal firewall rules between tiers
Random reboots in one rowCheck power (PDU and UPS alarms) and temperature (cooling) before the network

Common mistakes

Plugging both PSUs into one PDU.

The server has two power supplies but only one power path, so one tripped breaker takes it down.

Cabling both NICs to the same switch.

This protects against a bad cable, but not against the switch failing or being upgraded.

Running "redundant" links in one trench.

Two ISP cables that share a route can be cut by the same digger.

Not monitoring the spare.

A failed backup that nobody has noticed is no backup at all.

Leaving out blanking panels.

Hot exhaust air leaks to the front of the rack, and the servers overheat.

Never testing failover.

Generators that have never run under load, or firewall pairs that have never switched over, often fail on the day they are needed.

Key takeaways
  • A data centre houses servers and gives them constant power, cooling, security and connectivity.
  • Servers sit in racks, connected to two top-of-rack (leaf) switches; every leaf connects to every spine.
  • North-south traffic enters and leaves the data centre; east-west traffic flows between servers and is usually larger.
  • The edge has redundant routers (with two ISPs), firewall pairs and load balancers with a virtual IP address.
  • Power: A/B feeds, UPS, generators, two PDUs and dual PSUs. Cooling: hot and cold aisles, N+1 units.
  • Build in redundancy everywhere, avoid any single point of failure and monitor the spares.

Check yourself

Predict · scenario 1

In a spine-leaf network, server A is on leaf 1 and server B is on leaf 4. What path does traffic between them take?

Predict · scenario 2

The grid power fails. What keeps the servers running in the seconds before the generators start?

Predict · scenario 3

A web server behind a load balancer crashes. What do users notice?

Predict · scenario 4

A web server asks a database server in another rack for data. What kind of traffic is this?

Predict · scenario 5

A server's bond shows one NIC up and one down, but the server is working normally. What should you do?

Related lessons

Next, see how the same ideas look in software in Cloud networking basics. The other lessons below revisit the devices, the shapes networks take and how server groups are separated.

Learn more: SwitchesRoutersFirewallsNetwork TopologiesNetwork Segmentation

FAQ

What is the difference between a data centre and a server room?
A server room is a single room in an office with a few racks, often with one power feed and one air-conditioning unit. A data centre is a building (or a large part of one) designed only for IT equipment, with redundant power, generators, industrial cooling, physical security and several network connections. The ideas are the same, but the scale and the level of redundancy are very different.
Is the cloud just someone else's data centre?
Physically, yes: cloud providers run huge data centres full of servers, switches and the same power and cooling systems described here. What makes it 'cloud' is the software on top: you rent virtual servers and networks on demand, pay for what you use and never touch the hardware. The Cloud Networking Basics lesson explains how that works.
Why is it called spine-leaf?
The name comes from the drawing: a few large switches across the top form the backbone (the spine), and many smaller switches below are the leaves, with every leaf connected to every spine. Servers plug into the leaves. A server can reach a server on any other leaf through exactly one spine, so the path length is always the same.
What does N+1 mean?
N is the number of units you need to carry the load, and N+1 means you have one spare. If a room needs 4 cooling units, N+1 is 5, so any one unit can fail or be serviced. 2N means a complete second copy of everything, such as two separate power paths that can each carry the whole load.