Routelearn.net
Course menu

Course 12: IP Services on IOSLesson 3.1 (5 of 5 in this course)78 of 91 in the CCNA series

QoS concepts

Bandwidth, delay, jitter and loss; classification and marking (DSCP, CoS), queuing, policing and shaping, at CCNA depth.

Intermediate · 12 min read

QoS (Quality of Service) is a set of tools that classify, mark, queue, police and shape traffic so that chosen traffic types get better treatment when links are congested. It manages bandwidth, delay, jitter and loss, for example by giving voice packets priority over bulk downloads.

In simple terms: When a link is too busy for everything, QoS decides what goes first. A phone call gets through smoothly while a big download waits a little longer.

A real-life situation

HQ users call the server site over IP phones. Every evening at 18:00 a backup job copies files from HQ to SRV1 over the 10.0.12.0/30 link between R1 and R2, a 100 Mbps WAN circuit fed by a 1 Gbps LAN. During the backup, calls break up: words go missing and voices sound robotic. The link is full, and the router sends packets in the order they arrived. A voice packet stuck behind hundreds of backup packets arrives too late to be played. Quality of Service (QoS) lets R1 send the voice packets first.

Gi0/2 .2.1Gi0/0 .1192.168.10.0/24Gi0/1 .1.2 Gi0/010.0.12.0/30Gi0/1 .1192.168.20.0/24InternetNTP 198.51.100.10R1HQ edgeSW1mgmt 192.168.10.2PC1VLAN 10, DHCP clientR2server siteSRV1192.168.20.10
  1. 1. Two kinds of traffic, one link. The backup is TCP and wants as much bandwidth as it can get. The call is small UDP packets every 20 ms.
  2. 2. Without QoS: one FIFO queue. R1's Gi0/1 queue fills with backup packets. Voice waits behind them and some voice packets are dropped.
  3. 3. With QoS: voice jumps the queue. R1 recognises voice by its DSCP marking and sends it from a priority queue. The backup slows slightly; the call is clear.

What QoS is

QoS is a set of tools that treat some traffic better than other traffic when there is not enough bandwidth for everything. It manages four things:

  • Bandwidth: how many bits per second a type of traffic can use.
  • Delay (latency): how long a packet takes from sender to receiver.
  • Jitter: how much the delay changes from one packet to the next.
  • Loss: the share of packets that never arrive, usually dropped from a full queue.
TrafficNeedsTypical targets (one way)
VoiceLow delay, low jitter, low loss; small bandwidthDelay ≤ 150 ms, jitter ≤ 30 ms, loss ≤ 1%
Interactive videoLike voice, but much more bandwidthDelay ≤ 150-200 ms, jitter ≤ 30-50 ms, loss ≤ 0.1-1%
Data (web, files, backup)Tolerates delay; TCP resends lost dataBest effort

Why it works that way

Congestion happens wherever traffic arrives faster than it can leave: a 1 Gbps LAN feeding a 100 Mbps WAN link, or many ports feeding one uplink. The extra packets wait in an output queue. Waiting adds delay and jitter; when the queue is full, new packets are dropped (tail drop).

Data traffic copes: TCP slows down and resends what was lost (see Windowing and flow control). Voice and video use UDP and play in real time. A late packet is as useless as a lost one. So QoS doesn't make the link faster; it decides who waits and who is dropped, so the traffic that can't wait doesn't have to.

How it works step by step

Classify
Which kind of traffic is this?
Mark
Write DSCP / CoS
Police / shape
Limit the rate
Queue
One queue per class
Schedule
Priority first, then shares
The QoS tool chain. Classification and marking happen once, near the edge; queuing happens on every congested interface.

1. Classification and marking

Classification means sorting packets into classes, for example by ACL (addresses and ports), by the interface they came in on, or by looking deeper with NBAR (Network Based Application Recognition). Doing this on every router would be slow, so the first device writes the result into the packet: that is marking. Later devices just read the mark.

There are two common places for a mark:

Version4 bits
IHL4 bits
DSCP6 bitsQoS mark
ECN2 bits
Total length16 bits
First 32 bits of the IPv4 header. The old ToS byte is now DSCP (6 bits, values 0-63) plus ECN. IPv6 has the same 8 bits in its Traffic Class field. DSCP stays with the packet across routers.
TPID 0x810016 bits
PCP (CoS)3 bits0-7
DEI1 bit
VLAN ID12 bits
The 802.1Q tag. Its 3-bit Priority Code Point is the CoS value. It exists only on tagged (trunk) links and disappears at the first router.

Cisco and the RFCs use a set of standard DSCP names:

  • EF (Expedited Forwarding, 46): for voice.
  • AF (Assured Forwarding) AFxy: four classes (x = 1-4) with three drop precedences (y = 1-3). Value = 8x + 2y, so AF41 = 34 and AF21 = 18. A higher y is dropped first inside the same class.
  • CS (Class Selector) CS0-CS7: values 0, 8, 16 … 56, matching the old 3-bit IP Precedence.
  • DF (Default, 0): best effort.
TrafficDSCP nameDSCP valueTypical CoS
Voice (the call itself)EF465
Interactive videoAF41344
Call signallingCS3243
Important business dataAF21182
Routing protocolsCS6486
Everything else (best effort)DF / CS000

Trust boundary. Anyone can mark their own traffic EF and get priority. So the network only trusts marks from devices it controls. A Cisco IP phone marks its voice EF / CoS 5, and the access switch trusts the phone but re-marks the PC behind it to 0. The point where marks start being trusted is the trust boundary; it should be as close to the user as possible.

2. Queuing and scheduling

Each class gets its own queue on the output interface. The scheduler decides which queue sends next:

  • FIFO: one queue, first in first out. The default on fast interfaces, and the cause of the problem above.
  • CBWFQ (Class-Based Weighted Fair Queuing): each class is guaranteed a share of bandwidth during congestion, served in a round-robin way.
  • LLQ (Low Latency Queuing): CBWFQ plus one priority queue that is always served first. Voice and video go there. The priority queue is policed to its rate, so it can't starve the others.

3. Congestion avoidance

When a queue fills and tail drop starts, every TCP flow loses packets at the same moment, all slow down together, then all speed up together. This wave is called TCP global synchronisation, and it wastes the link. WRED (Weighted Random Early Detection) drops a few packets at randombefore the queue is full, choosing low-priority marks (for example AF23 before AF21) first. A few flows slow down early instead of all of them at once.

4. Policing and shaping

PolicingShaping
Excess traffic isDropped or re-markedHeld in a buffer and sent later
Adds delay?NoYes
DirectionInbound or outboundOutbound only
Typical useAn ISP enforcing the rate a customer pays forA customer sending at the contracted rate so the ISP doesn't drop it

Example: R1's Gi0/2 runs at 1 Gbps but the ISP contract is 100 Mbps, and the ISP polices at 100 Mbps. If R1 sends bursts at full speed, the ISP drops the excess. If R1 shapes to 100 Mbps, it queues the bursts itself, where its own QoS can still put voice first.

How to configure it on Cisco IOS

The CCNA asks you to explain QoS, not to configure it. Still, seeing a small example makes the ideas concrete. Cisco uses the MQC (Modular QoS CLI) in three steps: a class-map says what to match, a policy-map says what to do with each class, and service-policy applies it to an interface.

⚠️ Based on Cisco IOS / IOS XE documentation, not run on a lab device. Real QoS designs have more classes; this is a minimal sketch.

class-map match-any VOICE match dscp ef class-map match-any VIDEO match dscp af41 ! policy-map WAN-EDGE class VOICE priority percent 10 class VIDEO bandwidth percent 30 class class-default fair-queue ! interface GigabitEthernet0/1 service-policy output WAN-EDGE

On R1, toward R2. Voice gets a priority queue (LLQ) of up to 10% of the link, video a guaranteed 30% (CBWFQ), and everything else shares the rest fairly.

policy-map SHAPE-100M class class-default shape average 100000000 service-policy WAN-EDGE ! interface GigabitEthernet0/2 service-policy output SHAPE-100M

Shaping on the ISP link: send at most 100 Mbps, and apply the WAN-EDGE queuing inside the shaped rate (a nested, or hierarchical, policy).

How to verify it

Example output · based on Cisco documentation; exact format varies by platform and software version
R1#show policy-map interface GigabitEthernet0/1
 GigabitEthernet0/1

  Service-policy output: WAN-EDGE

    Class-map: VOICE (match-any)
      18432 packets, 3907584 bytes
      5 minute offered rate 64000 bps, drop rate 0000 bps
      Match: dscp ef (46)
      Priority: 10% (10000 kbps), burst bytes 250000, b/w exceed drops: 0
...
    Class-map: class-default (match-any)
      924117 packets, 1287365112 bytes
      5 minute offered rate 97000000 bps, drop rate 1210000 bps
      Match: any
...
Packets are matching the VOICE class, and none are dropped for exceeding the priority rate. The drops are in class-default, which is where you want them during a backup.
show class-map

Lists the class-maps and what each one matches.

What goes wrong and how to troubleshoot it

  • Voice still breaks up. Check the marks really arrive. If the switch doesn't trust the phone, it re-marks voice to 0 and R1's VOICE class never matches. Check the class counters in show policy-map interface.
  • The policy is on the wrong interface or direction. Queuing only helps on the congested outbound interface.
  • Drops in the priority class. More voice and video than the priority rate. Increase the rate or limit the number of calls.
  • The ISP drops traffic in bursts. It polices your rate; shape outbound to just below the contracted speed.

Common mistakes

  • Thinking QoS adds bandwidth. It only shares it during congestion.
  • Mixing up policing (drops, no delay) and shaping (buffers, adds delay, outbound only).
  • Trusting markings from user PCs. Users could mark everything EF.
  • Expecting CoS to survive a router. Only DSCP is in the IP header.
  • Putting too much traffic in the priority queue, so it is no longer special.

💡 Exam tip: the blueprint says explain the forwarding per-hop behavior (PHB) for QoS, such as classification, marking, queuing, congestion, policing and shaping. Expect concept questions: EF = 46 for voice, the AF formula, DSCP (6 bits, IP header) vs CoS (3 bits, 802.1Q tag), what a trust boundary is, LLQ for voice, WRED against TCP global synchronisation, and policing vs shaping. Voice targets: 150 ms delay, 30 ms jitter, 1% loss.

Key takeaways

  • QoS manages bandwidth, delay, jitter and loss when an interface is congested.
  • Classify and mark at the edge; trust marks only from devices you control.
  • DSCP (6 bits, IP header, end to end): EF 46 voice, AF41 video, DF 0 default. CoS is 3 bits in the 802.1Q tag.
  • LLQ = a policed priority queue for voice plus CBWFQ shares for the rest. WRED drops early to avoid global synchronisation.
  • Policing drops or re-marks excess; shaping buffers it and only works outbound.

Check yourself

Predict · scenario 1

A voice packet leaves PC1's IP phone marked CoS 5 and DSCP EF. After passing R1, which mark is still in the packet?

Predict · scenario 2

What is the DSCP value of AF31?

Predict · scenario 3

An ISP limits a customer to 50 Mbps by dropping anything above it. Which tool is the ISP using?

Predict · scenario 4

Which queuing method gives voice a strict priority queue while guaranteeing bandwidth to other classes?

Predict · scenario 5

During heavy congestion, all TCP flows slow down and speed up at the same time. Which tool helps?

FAQ

Does QoS create more bandwidth?
No. QoS only decides who goes first, and who waits or is dropped, when a link is congested. If the link is never full, QoS does nothing visible. If it is always full, more bandwidth is the real fix; QoS just chooses which traffic suffers.
What is the difference between DSCP and CoS?
CoS is a 3-bit value in the 802.1Q VLAN tag of an Ethernet frame, so it only exists on trunk links and is lost when a router removes the Layer 2 header. DSCP is a 6-bit value in the IP header, so it stays with the packet from end to end across routers. Networks usually classify on DSCP and map CoS to DSCP at the edge.
Why is voice given a priority queue but limited in size?
Voice needs low delay and jitter, so it is always sent first. But if the priority queue could take any amount of bandwidth, a flood of traffic marked EF could starve every other queue. LLQ therefore polices the priority queue to its configured rate during congestion.
When should I use policing and when shaping?
Shaping is used on your own outbound interface to send at a contracted rate (for example 50 Mbps on a 1 Gbps link to the ISP) without loss, by buffering the excess. Policing is used where you enforce a limit on someone else's traffic, often inbound at the ISP edge, by dropping or re-marking the excess.