A real-life situation
Three months ago you rolled out VLAN 30 for cameras. Since then, someone added it by hand on SW2 with a different name, someone else changed an NTP server on SW1 "just for a test", and a new switch was configured from an old copy of the template. Every switch is now slightly different, and nobody knows which one is right.
This is configuration drift: devices that were meant to match slowly drift apart through manual changes. Configuration management tools stop it by keeping the correct configuration in files and applying it automatically.
What it is
Configuration management means defining how devices should be configured in text files, storing those files in one place (often a Git repository, which records every change and who made it), and using a tool to make the devices match. Treating configuration like program code is called infrastructure as code (IaC).
The tools differ in a few ways the CCNA expects you to know:
- Agent vs. agentless: an agent is software installed on the managed device. Agentless tools use what the device already has (SSH, NETCONF, a REST API). Most network devices can't run agents, so agentless tools fit networking well.
- Push vs. pull: in a push model the central server connects out and sends the configuration. In a pull model each device's agent regularly asks the server for its configuration.
- Declarative vs. procedural: declarative means you describe the end state ("VLAN 30 exists, named CAMERAS") and the tool works out the steps. Procedural (imperative) means you list the steps to run in order.
| Tool | Agent? | Model | Written in | Connects with |
|---|---|---|---|---|
| Ansible (Red Hat) | Agentless | Push | YAML playbooks; Ansible itself is written in Python | SSH (CLI), NETCONF, REST APIs |
| Terraform (HashiCorp) | Agentless | Push, declarative, with a state file | HCL (HashiCorp Configuration Language) | APIs, through providers |
| Puppet | Agent-based (agentless options exist) | Pull | Manifests in Puppet's own language | Agent to server on TCP 8140 |
| Chef | Agent-based | Pull | Recipes and cookbooks in Ruby | Agent to server over HTTPS |
Why it works that way
- One source of truth: the files say what is correct. If a device differs, the device is wrong, not the file.
- Idempotence: good modules check the current state first and only change what differs. Running the same playbook ten times leaves the device exactly as running it once would. That is what makes scheduled "fix the drift" runs safe.
- Review and rollback: because changes are files in Git, a colleague can review them before they run, and you can go back to last week's version.
- Scale: the effort to change 2 switches or 200 is the same: edit one file, run one command.
How Ansible works step by step
Ansible needs three things on the control node (the NetOps PC in the lab):
- An inventory: the list of devices, in groups, with how to reach them.
- A playbook: a YAML file with one or more plays, each a list of tasks.
- Modules: the code each task calls. The
cisco.ioscollection has modules such asios_vlans,ios_interfacesandios_config.
- 1. 1. Connect to every host in the group. ansible-playbook reads the inventory and opens an SSH session to SW1 (10.10.1.11) and SW2 (10.10.1.12) at the same time. Nothing is installed on the switches.
- 2. 2. Read the current state. The ios_vlans module runs show commands and turns the output into structured data about which VLANs exist.
- 3. 3. Send only what differs. SW1 is missing VLAN 30, so Ansible sends the commands to create it. SW2 already matches, so nothing is sent to it.
- 4. 4. Report per device. SW1 reports changed, SW2 reports ok. The play recap sums up every host.
How Terraform works step by step
Terraform is declarative and keeps a state file: its record of every resource it manages. Each run compares three things: your files, the state file, and the real infrastructure.
How to configure it
On the Cisco devices
An agentless tool only needs management access. On SW1 and SW2 give Ansible its own account, so its changes are easy to spot in logs:
username automation privilege 15 secret Auto-Secret-123
ip domain name example.com
crypto key generate rsa modulus 2048
ip ssh version 2
line vty 0 15
transport input sshSSH with a local privilege 15 user. Logins use the AAA local settings from the first lesson (without AAA, add login local under the VTY lines). Terraform's IOS XE provider uses the device API instead, so it also needs RESTCONF or NETCONF enabled.
The steps to turn on RESTCONF are in REST APIs and JSON.
The Ansible inventory and playbook
[access_switches] SW1 ansible_host=10.10.1.11 SW2 ansible_host=10.10.1.12 [access_switches:vars] ansible_network_os=cisco.ios.ios ansible_connection=ansible.netcommon.network_cli ansible_user=automation
[access_switches]: a group; the playbook targets groups, not single IPsansible_network_os: tells Ansible these are Cisco IOS devicesnetwork_cli: connect over SSH and use the CLI
---
- name: Standard VLANs on access switches
hosts: access_switches
gather_facts: false
tasks:
- name: Ensure VLANs exist
cisco.ios.ios_vlans:
config:
- vlan_id: 10
name: USERS
- vlan_id: 20
name: VOICE
- vlan_id: 30
name: CAMERAS
- vlan_id: 99
name: MGMT
state: mergedhosts: which inventory group this play runs againstcisco.ios.ios_vlans: the module; it reads current VLANs and adds what is missingstate: merged: add or update these VLANs and leave others alone (replaced or overridden would remove extras)
YAML uses indentation (spaces, never tabs) instead of braces. A list item starts with - , and key: value pairs work like JSON objects. Unlike JSON, YAML allows comments with #.
Run it first in check mode, which reports what would change without touching the devices, then for real:
ansible-playbook -i inventory.ini vlans.yml --check --diff --ask-passDry run. --ask-pass prompts for the SSH password; in production keep secrets in Ansible Vault, never in the inventory.
ansible-playbook -i inventory.ini vlans.yml --ask-passApply the changes.
The same VLAN with Terraform
terraform {
required_providers {
iosxe = {
source = "CiscoDevNet/iosxe"
}
}
}
provider "iosxe" {
username = "automation"
password = var.device_password
url = "https://10.10.1.11"
}
resource "iosxe_vlan" "cameras" {
vlan_id = 30
name = "CAMERAS"
}provider: the plug-in that knows how to talk to IOS XE through its APIresource: one thing Terraform should create and track: VLAN 30var.device_password: a variable, so the password is not written in the file
How to verify it
netops$ ansible-playbook -i inventory.ini vlans.yml --ask-pass SSH password: PLAY [Standard VLANs on access switches] *************************************** TASK [Ensure VLANs exist] ****************************************************** changed: [SW1] ok: [SW2] PLAY RECAP ********************************************************************* SW1 : ok=1 changed=1 unreachable=0 failed=0 skipped=0 rescued=0 ignored=0 SW2 : ok=1 changed=0 unreachable=0 failed=0 skipped=0 rescued=0 ignored=0
changed=0: that is idempotence.SW1#show vlan brief VLAN Name Status Ports ---- -------------------------------- --------- ------------------------------- 1 default active Gi1/0/13, Gi1/0/14 10 USERS active Gi1/0/1, Gi1/0/2 20 VOICE active 30 CAMERAS active 99 MGMT active 1002 fddi-default act/unsup 1003 token-ring-default act/unsup 1004 fddinet-default act/unsup 1005 trnet-default act/unsup
netops$ terraform plan Terraform will perform the following actions: # iosxe_vlan.cameras will be created + resource "iosxe_vlan" "cameras" { + id = (known after apply) + name = "CAMERAS" + vlan_id = 30 } Plan: 1 to add, 0 to change, 0 to destroy.
+ means create, ~ update in place, - destroy. Nothing changes until you run terraform apply and confirm.What goes wrong and how to troubleshoot it
| Symptom | Likely cause | What to check |
|---|---|---|
Ansible: unreachable=1 | No route, SSH not enabled, wrong password, unknown SSH host key | SSH to the switch by hand from the control node with the same user |
Ansible: failed=1 with a module error | Wrong ansible_network_os, missing collection, or the user lacks privilege 15 | Run again with -vvv for detail; check the user's privilege |
| Changes come back after a run | Someone keeps editing by hand, or two tools manage the same setting | Pick one source of truth; track manual changes in Syslog |
| Terraform plan wants to re-create things that exist | The state file is missing or out of date | Keep state in a shared, backed-up location; import existing resources |
| Terraform: connection or 401 errors | RESTCONF/NETCONF not enabled, wrong URL or credentials | Test the device API with curl first |
Common mistakes
- Calling Ansible agent-based. It is agentless and pushes over SSH; Puppet and Chef are the agent-based, pull-model tools.
- Using tabs in YAML, or getting the indentation wrong: the playbook won't parse.
- Skipping
--check(orterraform plan) before a change to many devices. - Storing device passwords in plain text in the inventory or
.tffiles. - Fixing a device by hand after automating it, which creates the very drift the tool is there to remove.
💡 Exam tip: the CCNA asks you to recognize the capabilities of configuration management tools such as Ansible and Terraform. Know that both are agentless; Ansible uses YAML playbooks and pushes over SSH; Terraform uses HCL, is declarative, and tracks what it built in a state file. If Puppet or Chef come up, they are agent-based and pull their configuration. Also know what configuration drift is and why idempotence matters.
Key takeaways
- Configuration drift: devices drift from the standard through manual changes.
- Infrastructure as code keeps the intended config in files, ideally in Git.
- Ansible: agentless, push, YAML playbooks, inventory, modules such as
cisco.ios.ios_vlans. - Terraform: agentless, declarative HCL, providers, plan then apply, state file.
- Puppet and Chef: agent-based, pull model.
- Verify on the device too:
showcommands still have the final word.
Check yourself
Which tool needs no software installed on the switches and pushes configuration over SSH using YAML files?
You run the same Ansible playbook twice in a row. The first run shows changed=1 on SW1. What should the second run show for SW1?
Which Terraform command shows what would change without changing anything?
What is configuration drift?
Which pair of tools uses agents on the managed device and a pull model?