Routelearn.net
Course menu

Course 15: Automation and ProgrammabilityLesson 1.3 (3 of 4 in this course)90 of 91 in the CCNA series

Ansible, Terraform and configuration management

Why configuration drift happens, and how Ansible, Terraform, Puppet and Chef manage device configurations.

Intermediate · 10 min read

Configuration management is the practice of describing the intended configuration of devices in version-controlled text files and using a tool, such as Ansible or Terraform, to make the devices match that description. It treats configuration as code (infrastructure as code), which makes changes repeatable and lets the tool detect and correct drift.

In simple terms: Instead of typing the same commands on every switch, you write down once how they should look. A tool then sets up every device to match and fixes any that have drifted.

A real-life situation

Three months ago you rolled out VLAN 30 for cameras. Since then, someone added it by hand on SW2 with a different name, someone else changed an NTP server on SW1 "just for a test", and a new switch was configured from an old copy of the template. Every switch is now slightly different, and nobody knows which one is right.

This is configuration drift: devices that were meant to match slowly drift apart through manual changes. Configuration management tools stop it by keeping the correct configuration in files and applying it automatically.

What it is

Configuration management means defining how devices should be configured in text files, storing those files in one place (often a Git repository, which records every change and who made it), and using a tool to make the devices match. Treating configuration like program code is called infrastructure as code (IaC).

The tools differ in a few ways the CCNA expects you to know:

  • Agent vs. agentless: an agent is software installed on the managed device. Agentless tools use what the device already has (SSH, NETCONF, a REST API). Most network devices can't run agents, so agentless tools fit networking well.
  • Push vs. pull: in a push model the central server connects out and sends the configuration. In a pull model each device's agent regularly asks the server for its configuration.
  • Declarative vs. procedural: declarative means you describe the end state ("VLAN 30 exists, named CAMERAS") and the tool works out the steps. Procedural (imperative) means you list the steps to run in order.
ToolAgent?ModelWritten inConnects with
Ansible (Red Hat)AgentlessPushYAML playbooks; Ansible itself is written in PythonSSH (CLI), NETCONF, REST APIs
Terraform (HashiCorp)AgentlessPush, declarative, with a state fileHCL (HashiCorp Configuration Language)APIs, through providers
PuppetAgent-based (agentless options exist)PullManifests in Puppet's own languageAgent to server on TCP 8140
ChefAgent-basedPullRecipes and cookbooks in RubyAgent to server over HTTPS

Why it works that way

  • One source of truth: the files say what is correct. If a device differs, the device is wrong, not the file.
  • Idempotence: good modules check the current state first and only change what differs. Running the same playbook ten times leaves the device exactly as running it once would. That is what makes scheduled "fix the drift" runs safe.
  • Review and rollback: because changes are files in Git, a colleague can review them before they run, and you can go back to last week's version.
  • Scale: the effort to change 2 switches or 200 is the same: edit one file, run one command.

How Ansible works step by step

Ansible needs three things on the control node (the NetOps PC in the lab):

  • An inventory: the list of devices, in groups, with how to reach them.
  • A playbook: a YAML file with one or more plays, each a list of tasks.
  • Modules: the code each task calls. The cisco.ios collection has modules such as ios_vlans, ios_interfaces and ios_config.
Gi0/0/0 .1Gi0/0/1Gi1/0/24Gi0/0/2Gi1/0/24NetOps PC10.10.0.50 (Ansible, Terraform)MGMT-SW10.10.0.0/24Catalyst Center10.10.0.10 cc1.example.comR1mgmt 10.10.0.1SW1mgmt 10.10.1.11SW2mgmt 10.10.1.12
  1. 1. 1. Connect to every host in the group. ansible-playbook reads the inventory and opens an SSH session to SW1 (10.10.1.11) and SW2 (10.10.1.12) at the same time. Nothing is installed on the switches.
  2. 2. 2. Read the current state. The ios_vlans module runs show commands and turns the output into structured data about which VLANs exist.
  3. 3. 3. Send only what differs. SW1 is missing VLAN 30, so Ansible sends the commands to create it. SW2 already matches, so nothing is sent to it.
  4. 4. 4. Report per device. SW1 reports changed, SW2 reports ok. The play recap sums up every host.

How Terraform works step by step

Terraform is declarative and keeps a state file: its record of every resource it manages. Each run compares three things: your files, the state file, and the real infrastructure.

Write .tf files
describe the resources you want, in HCL
terraform init
download the providers the files need
terraform plan
compare files, state and reality; show what would change
terraform apply
make the changes through the provider's API
State file updated
records what now exists, for the next plan
The Terraform workflow.

How to configure it

On the Cisco devices

An agentless tool only needs management access. On SW1 and SW2 give Ansible its own account, so its changes are easy to spot in logs:

username automation privilege 15 secret Auto-Secret-123 ip domain name example.com crypto key generate rsa modulus 2048 ip ssh version 2 line vty 0 15 transport input ssh

SSH with a local privilege 15 user. Logins use the AAA local settings from the first lesson (without AAA, add login local under the VTY lines). Terraform's IOS XE provider uses the device API instead, so it also needs RESTCONF or NETCONF enabled.

The steps to turn on RESTCONF are in REST APIs and JSON.

The Ansible inventory and playbook

Illustrative example · written for this lesson; check current vendor documentation for exact names
inventory.iniINI
[access_switches]
SW1 ansible_host=10.10.1.11
SW2 ansible_host=10.10.1.12

[access_switches:vars]
ansible_network_os=cisco.ios.ios
ansible_connection=ansible.netcommon.network_cli
ansible_user=automation
  • [access_switches]: a group; the playbook targets groups, not single IPs
  • ansible_network_os: tells Ansible these are Cisco IOS devices
  • network_cli: connect over SSH and use the CLI
Illustrative example · written for this lesson; check current vendor documentation for exact names
vlans.ymlYAML
---
- name: Standard VLANs on access switches
  hosts: access_switches
  gather_facts: false

  tasks:
    - name: Ensure VLANs exist
      cisco.ios.ios_vlans:
        config:
          - vlan_id: 10
            name: USERS
          - vlan_id: 20
            name: VOICE
          - vlan_id: 30
            name: CAMERAS
          - vlan_id: 99
            name: MGMT
        state: merged
  • hosts: which inventory group this play runs against
  • cisco.ios.ios_vlans: the module; it reads current VLANs and adds what is missing
  • state: merged: add or update these VLANs and leave others alone (replaced or overridden would remove extras)

YAML uses indentation (spaces, never tabs) instead of braces. A list item starts with - , and key: value pairs work like JSON objects. Unlike JSON, YAML allows comments with #.

Run it first in check mode, which reports what would change without touching the devices, then for real:

ansible-playbook -i inventory.ini vlans.yml --check --diff --ask-pass

Dry run. --ask-pass prompts for the SSH password; in production keep secrets in Ansible Vault, never in the inventory.

ansible-playbook -i inventory.ini vlans.yml --ask-pass

Apply the changes.

The same VLAN with Terraform

Illustrative example · written for this lesson; check current vendor documentation for exact names
main.tfHCL
terraform {
  required_providers {
    iosxe = {
      source = "CiscoDevNet/iosxe"
    }
  }
}

provider "iosxe" {
  username = "automation"
  password = var.device_password
  url      = "https://10.10.1.11"
}

resource "iosxe_vlan" "cameras" {
  vlan_id = 30
  name    = "CAMERAS"
}
  • provider: the plug-in that knows how to talk to IOS XE through its API
  • resource: one thing Terraform should create and track: VLAN 30
  • var.device_password: a variable, so the password is not written in the file

How to verify it

Example output · written for this lesson; exact text varies by tool and version
netops$ ansible-playbook -i inventory.ini vlans.yml --ask-pass
SSH password:

PLAY [Standard VLANs on access switches] ***************************************

TASK [Ensure VLANs exist] ******************************************************
changed: [SW1]
ok: [SW2]

PLAY RECAP *********************************************************************
SW1                        : ok=1    changed=1    unreachable=0    failed=0    skipped=0    rescued=0    ignored=0
SW2                        : ok=1    changed=0    unreachable=0    failed=0    skipped=0    rescued=0    ignored=0
changed on SW1: Ansible added the missing VLAN. ok on SW2: it already matched, so nothing was sent. Run the playbook again and SW1 also reports changed=0: that is idempotence.
Example output · based on Cisco documentation; exact format varies by platform and software version
SW1#show vlan brief
VLAN Name                             Status    Ports
---- -------------------------------- --------- -------------------------------
1    default                          active    Gi1/0/13, Gi1/0/14
10   USERS                            active    Gi1/0/1, Gi1/0/2
20   VOICE                            active
30   CAMERAS                          active
99   MGMT                             active
1002 fddi-default                     act/unsup
1003 token-ring-default               act/unsup
1004 fddinet-default                  act/unsup
1005 trnet-default                    act/unsup
VLAN 30 CAMERAS now exists on SW1. Always confirm on the device: the tool reporting success and the device being right are two different checks.
Example output · written for this lesson; exact text varies by tool and version
netops$ terraform plan
Terraform will perform the following actions:

  # iosxe_vlan.cameras will be created
  + resource "iosxe_vlan" "cameras" {
      + id      = (known after apply)
      + name    = "CAMERAS"
      + vlan_id = 30
    }

Plan: 1 to add, 0 to change, 0 to destroy.
+ means create, ~ update in place, - destroy. Nothing changes until you run terraform apply and confirm.

What goes wrong and how to troubleshoot it

SymptomLikely causeWhat to check
Ansible: unreachable=1No route, SSH not enabled, wrong password, unknown SSH host keySSH to the switch by hand from the control node with the same user
Ansible: failed=1 with a module errorWrong ansible_network_os, missing collection, or the user lacks privilege 15Run again with -vvv for detail; check the user's privilege
Changes come back after a runSomeone keeps editing by hand, or two tools manage the same settingPick one source of truth; track manual changes in Syslog
Terraform plan wants to re-create things that existThe state file is missing or out of dateKeep state in a shared, backed-up location; import existing resources
Terraform: connection or 401 errorsRESTCONF/NETCONF not enabled, wrong URL or credentialsTest the device API with curl first

Common mistakes

  • Calling Ansible agent-based. It is agentless and pushes over SSH; Puppet and Chef are the agent-based, pull-model tools.
  • Using tabs in YAML, or getting the indentation wrong: the playbook won't parse.
  • Skipping --check (or terraform plan) before a change to many devices.
  • Storing device passwords in plain text in the inventory or .tf files.
  • Fixing a device by hand after automating it, which creates the very drift the tool is there to remove.

💡 Exam tip: the CCNA asks you to recognize the capabilities of configuration management tools such as Ansible and Terraform. Know that both are agentless; Ansible uses YAML playbooks and pushes over SSH; Terraform uses HCL, is declarative, and tracks what it built in a state file. If Puppet or Chef come up, they are agent-based and pull their configuration. Also know what configuration drift is and why idempotence matters.

Key takeaways

  • Configuration drift: devices drift from the standard through manual changes.
  • Infrastructure as code keeps the intended config in files, ideally in Git.
  • Ansible: agentless, push, YAML playbooks, inventory, modules such as cisco.ios.ios_vlans.
  • Terraform: agentless, declarative HCL, providers, plan then apply, state file.
  • Puppet and Chef: agent-based, pull model.
  • Verify on the device too: show commands still have the final word.

Check yourself

Predict · scenario 1

Which tool needs no software installed on the switches and pushes configuration over SSH using YAML files?

Predict · scenario 2

You run the same Ansible playbook twice in a row. The first run shows changed=1 on SW1. What should the second run show for SW1?

Predict · scenario 3

Which Terraform command shows what would change without changing anything?

Predict · scenario 4

What is configuration drift?

Predict · scenario 5

Which pair of tools uses agents on the managed device and a pull model?

FAQ

Do I need to install anything on a Cisco switch to use Ansible?
No. Ansible is agentless: it connects to network devices over SSH (or NETCONF or a REST API) from a control node, such as a Linux PC. The switch only needs a reachable management address, SSH and a user account.
What does idempotent mean?
Running the same task again gives the same result and changes nothing if the device is already correct. An idempotent Ansible task reports ok instead of changed on the second run. That makes it safe to run a playbook every day to fix drift.
When would I pick Terraform instead of Ansible?
Terraform is strongest at building and tracking infrastructure through APIs: cloud networks, virtual machines, and controller or device resources, with a state file that records what it created. Ansible is strongest at configuring existing devices and running step-by-step tasks over SSH. Many teams use both.
Are Puppet and Chef still on the CCNA?
The CCNA 200-301 v1.1 blueprint names Ansible and Terraform. Puppet and Chef were in the older version and still appear in study material. It is worth knowing that both are agent-based and use a pull model, unlike Ansible and Terraform.