A real-life situation
HQ users call the server site over IP phones. Every evening at 18:00 a backup job copies files from HQ to SRV1 over the 10.0.12.0/30 link between R1 and R2, a 100 Mbps WAN circuit fed by a 1 Gbps LAN. During the backup, calls break up: words go missing and voices sound robotic. The link is full, and the router sends packets in the order they arrived. A voice packet stuck behind hundreds of backup packets arrives too late to be played. Quality of Service (QoS) lets R1 send the voice packets first.
- 1. Two kinds of traffic, one link. The backup is TCP and wants as much bandwidth as it can get. The call is small UDP packets every 20 ms.
- 2. Without QoS: one FIFO queue. R1's Gi0/1 queue fills with backup packets. Voice waits behind them and some voice packets are dropped.
- 3. With QoS: voice jumps the queue. R1 recognises voice by its DSCP marking and sends it from a priority queue. The backup slows slightly; the call is clear.
What QoS is
QoS is a set of tools that treat some traffic better than other traffic when there is not enough bandwidth for everything. It manages four things:
- Bandwidth: how many bits per second a type of traffic can use.
- Delay (latency): how long a packet takes from sender to receiver.
- Jitter: how much the delay changes from one packet to the next.
- Loss: the share of packets that never arrive, usually dropped from a full queue.
| Traffic | Needs | Typical targets (one way) |
|---|---|---|
| Voice | Low delay, low jitter, low loss; small bandwidth | Delay ≤ 150 ms, jitter ≤ 30 ms, loss ≤ 1% |
| Interactive video | Like voice, but much more bandwidth | Delay ≤ 150-200 ms, jitter ≤ 30-50 ms, loss ≤ 0.1-1% |
| Data (web, files, backup) | Tolerates delay; TCP resends lost data | Best effort |
Why it works that way
Congestion happens wherever traffic arrives faster than it can leave: a 1 Gbps LAN feeding a 100 Mbps WAN link, or many ports feeding one uplink. The extra packets wait in an output queue. Waiting adds delay and jitter; when the queue is full, new packets are dropped (tail drop).
Data traffic copes: TCP slows down and resends what was lost (see Windowing and flow control). Voice and video use UDP and play in real time. A late packet is as useless as a lost one. So QoS doesn't make the link faster; it decides who waits and who is dropped, so the traffic that can't wait doesn't have to.
How it works step by step
1. Classification and marking
Classification means sorting packets into classes, for example by ACL (addresses and ports), by the interface they came in on, or by looking deeper with NBAR (Network Based Application Recognition). Doing this on every router would be slow, so the first device writes the result into the packet: that is marking. Later devices just read the mark.
There are two common places for a mark:
Cisco and the RFCs use a set of standard DSCP names:
- EF (Expedited Forwarding, 46): for voice.
- AF (Assured Forwarding) AFxy: four classes (x = 1-4) with three drop precedences (y = 1-3). Value = 8x + 2y, so AF41 = 34 and AF21 = 18. A higher y is dropped first inside the same class.
- CS (Class Selector) CS0-CS7: values 0, 8, 16 … 56, matching the old 3-bit IP Precedence.
- DF (Default, 0): best effort.
| Traffic | DSCP name | DSCP value | Typical CoS |
|---|---|---|---|
| Voice (the call itself) | EF | 46 | 5 |
| Interactive video | AF41 | 34 | 4 |
| Call signalling | CS3 | 24 | 3 |
| Important business data | AF21 | 18 | 2 |
| Routing protocols | CS6 | 48 | 6 |
| Everything else (best effort) | DF / CS0 | 0 | 0 |
Trust boundary. Anyone can mark their own traffic EF and get priority. So the network only trusts marks from devices it controls. A Cisco IP phone marks its voice EF / CoS 5, and the access switch trusts the phone but re-marks the PC behind it to 0. The point where marks start being trusted is the trust boundary; it should be as close to the user as possible.
2. Queuing and scheduling
Each class gets its own queue on the output interface. The scheduler decides which queue sends next:
- FIFO: one queue, first in first out. The default on fast interfaces, and the cause of the problem above.
- CBWFQ (Class-Based Weighted Fair Queuing): each class is guaranteed a share of bandwidth during congestion, served in a round-robin way.
- LLQ (Low Latency Queuing): CBWFQ plus one priority queue that is always served first. Voice and video go there. The priority queue is policed to its rate, so it can't starve the others.
3. Congestion avoidance
When a queue fills and tail drop starts, every TCP flow loses packets at the same moment, all slow down together, then all speed up together. This wave is called TCP global synchronisation, and it wastes the link. WRED (Weighted Random Early Detection) drops a few packets at randombefore the queue is full, choosing low-priority marks (for example AF23 before AF21) first. A few flows slow down early instead of all of them at once.
4. Policing and shaping
| Policing | Shaping | |
|---|---|---|
| Excess traffic is | Dropped or re-marked | Held in a buffer and sent later |
| Adds delay? | No | Yes |
| Direction | Inbound or outbound | Outbound only |
| Typical use | An ISP enforcing the rate a customer pays for | A customer sending at the contracted rate so the ISP doesn't drop it |
Example: R1's Gi0/2 runs at 1 Gbps but the ISP contract is 100 Mbps, and the ISP polices at 100 Mbps. If R1 sends bursts at full speed, the ISP drops the excess. If R1 shapes to 100 Mbps, it queues the bursts itself, where its own QoS can still put voice first.
How to configure it on Cisco IOS
The CCNA asks you to explain QoS, not to configure it. Still, seeing a small example makes the ideas concrete. Cisco uses the MQC (Modular QoS CLI) in three steps: a class-map says what to match, a policy-map says what to do with each class, and service-policy applies it to an interface.
⚠️ Based on Cisco IOS / IOS XE documentation, not run on a lab device. Real QoS designs have more classes; this is a minimal sketch.
class-map match-any VOICE
match dscp ef
class-map match-any VIDEO
match dscp af41
!
policy-map WAN-EDGE
class VOICE
priority percent 10
class VIDEO
bandwidth percent 30
class class-default
fair-queue
!
interface GigabitEthernet0/1
service-policy output WAN-EDGEOn R1, toward R2. Voice gets a priority queue (LLQ) of up to 10% of the link, video a guaranteed 30% (CBWFQ), and everything else shares the rest fairly.
policy-map SHAPE-100M
class class-default
shape average 100000000
service-policy WAN-EDGE
!
interface GigabitEthernet0/2
service-policy output SHAPE-100MShaping on the ISP link: send at most 100 Mbps, and apply the WAN-EDGE queuing inside the shaped rate (a nested, or hierarchical, policy).
How to verify it
R1#show policy-map interface GigabitEthernet0/1 GigabitEthernet0/1 Service-policy output: WAN-EDGE Class-map: VOICE (match-any) 18432 packets, 3907584 bytes 5 minute offered rate 64000 bps, drop rate 0000 bps Match: dscp ef (46) Priority: 10% (10000 kbps), burst bytes 250000, b/w exceed drops: 0 ... Class-map: class-default (match-any) 924117 packets, 1287365112 bytes 5 minute offered rate 97000000 bps, drop rate 1210000 bps Match: any ...
show class-mapLists the class-maps and what each one matches.
What goes wrong and how to troubleshoot it
- Voice still breaks up. Check the marks really arrive. If the switch doesn't trust the phone, it re-marks voice to 0 and R1's VOICE class never matches. Check the class counters in
show policy-map interface. - The policy is on the wrong interface or direction. Queuing only helps on the congested outbound interface.
- Drops in the priority class. More voice and video than the priority rate. Increase the rate or limit the number of calls.
- The ISP drops traffic in bursts. It polices your rate; shape outbound to just below the contracted speed.
Common mistakes
- Thinking QoS adds bandwidth. It only shares it during congestion.
- Mixing up policing (drops, no delay) and shaping (buffers, adds delay, outbound only).
- Trusting markings from user PCs. Users could mark everything EF.
- Expecting CoS to survive a router. Only DSCP is in the IP header.
- Putting too much traffic in the priority queue, so it is no longer special.
💡 Exam tip: the blueprint says explain the forwarding per-hop behavior (PHB) for QoS, such as classification, marking, queuing, congestion, policing and shaping. Expect concept questions: EF = 46 for voice, the AF formula, DSCP (6 bits, IP header) vs CoS (3 bits, 802.1Q tag), what a trust boundary is, LLQ for voice, WRED against TCP global synchronisation, and policing vs shaping. Voice targets: 150 ms delay, 30 ms jitter, 1% loss.
Key takeaways
- QoS manages bandwidth, delay, jitter and loss when an interface is congested.
- Classify and mark at the edge; trust marks only from devices you control.
- DSCP (6 bits, IP header, end to end): EF 46 voice, AF41 video, DF 0 default. CoS is 3 bits in the 802.1Q tag.
- LLQ = a policed priority queue for voice plus CBWFQ shares for the rest. WRED drops early to avoid global synchronisation.
- Policing drops or re-marks excess; shaping buffers it and only works outbound.
Check yourself
A voice packet leaves PC1's IP phone marked CoS 5 and DSCP EF. After passing R1, which mark is still in the packet?
What is the DSCP value of AF31?
An ISP limits a customer to 50 Mbps by dropping anything above it. Which tool is the ISP using?
Which queuing method gives voice a strict priority queue while guaranteeing bandwidth to other classes?
During heavy congestion, all TCP flows slow down and speed up at the same time. Which tool helps?