Routelearn.net
Course menu

Course 15: Automation and ProgrammabilityLesson 1.4 (4 of 4 in this course)91 of 91 in the CCNA series

AI and machine learning in network operations

What generative AI, predictive AI and machine learning can and can't do for running networks.

Intermediate · 8 min read

AI (artificial intelligence) in network operations means using machine learning and other AI techniques to analyse network data such as telemetry, logs and flow records. Predictive AI learns a baseline and flags anomalies or forecasts problems, while generative AI drafts text, summaries or configuration that a person must review.

In simple terms: The network produces far more data than anyone can read. AI tools learn what “normal” looks like, point out what is unusual and help you write things faster, but you still check their answers.

A real-life situation

Monday, 9 a.m. Users in one building say Wi-Fi is "slow". Overnight the network produced thousands of Syslog messages and millions of flow records. Reading them by hand would take a day. Your controller's analytics page already shows a note: "Client onboarding time in Building B is far above its normal Monday level; most failures are DHCP timeouts." That points you straight at the DHCP server.

Ten minutes later a colleague asks an AI chatbot how to add a VLAN to a trunk, pastes the answer into a switch, and cuts off every VLAN on the link. Both stories are AI in network operations. One saved time; the other caused an outage. This lesson explains what these tools are, what they are good at, and where they fail.

What it is

  • Artificial intelligence (AI): the broad field of making computers do tasks that normally need human judgement, such as spotting patterns, making predictions or writing text.
  • Machine learning (ML): a part of AI where the system learns patterns from data instead of following rules a programmer wrote. Feed it months of interface counters and it learns what "normal" looks like.

Machine learning comes in three main types:

TypeHow it learnsNetwork example
SupervisedFrom examples that are already labelled with the right answerClassify traffic flows by application after training on flows labelled "video", "voice", "backup"
UnsupervisedFinds structure in data with no labelsGroup devices that behave alike; flag a device that suddenly behaves unlike its group
ReinforcementTries actions and learns from rewards or penaltiesResearch into tuning radio settings or traffic paths; less common in everyday tools

Network products use AI in two broad ways:

Predictive AIGenerative AI
What it doesAnalyses data to detect, classify and forecastCreates new content: text, summaries, config drafts
InputTelemetry: counters, logs, flows, client eventsA question or instruction in plain language, plus context
Output"This is abnormal", "this link will be full in six weeks""Here is a summary of these logs", "here is a config that might do this"
Network usesDynamic baselines, anomaly detection, capacity planning, failure prediction, root-cause hintsChat assistants, natural-language queries ("which switches run old software?"), log and ticket summaries, drafting scripts and configs
Main riskFalse alarms, or missing a real problemHallucination: fluent, confident, wrong output

Generative AI tools are usually built on a large language model (LLM): a model trained on huge amounts of text to predict likely next words. That is why its answers read well, and also why they can be wrong in ways that sound right.

Why it works that way

Traditional monitoring uses static thresholds: "alert if CPU is above 80%" or "alert if a link is above 90%". These are crude. A backup link at 95% at 2 a.m. may be perfectly normal, while 40% on a link that usually sits at 5% may be an attack.

Machine learning builds a dynamic baseline for each metric: what is normal for this link, at this hour, on this day of the week. It then flags anomalies, values far from that baseline. With thousands of devices, only a machine can watch every metric that closely. It can also correlate events, noticing that 200 separate alerts started within a minute of one switch reboot.

Generative AI works for a different reason: it has seen a lot of text, including documentation and configurations, so it can draft and summarize quickly. But it predicts what is likely, not what is true for your network, platform and software version. It doesn't know your topology unless you tell it.

How it works step by step

A predictive AI feature, such as the assurance analytics in a controller, follows a pipeline:

Collect
SNMP, Syslog, NetFlow, streaming telemetry, client events
Store and clean
consistent timestamps, device names, units
Learn a baseline
what is normal per device, per hour, per day
Detect and correlate
flag anomalies; group related events into one issue
Suggest
likely root cause and a suggested fix
Human decides
verify with show commands; approve or automate the fix
From raw telemetry to an action a person approves.

In the course lab, Catalyst Center plays the analytics role:

Gi0/0/0 .1Gi0/0/1Gi1/0/24Gi0/0/2Gi1/0/24NetOps PC10.10.0.50 (Ansible, Terraform)MGMT-SW10.10.0.0/24Catalyst Center10.10.0.10 cc1.example.comR1mgmt 10.10.0.1SW1mgmt 10.10.1.11SW2mgmt 10.10.1.12
  1. 1. 1. Devices stream data. SW1, SW2 and R1 send Syslog, SNMP data and flow records to Catalyst Center all day.
  2. 2. 2. The controller learns normal. Over weeks it learns each metric's normal range for each hour and day.
  3. 3. 3. An anomaly is raised. SW2's uplink errors jump far above baseline. Related alerts are grouped into one issue with a suggested cause.
  4. 4. 4. The engineer verifies. The engineer checks show interfaces on SW2 before acting. The AI pointed the way; the device confirms it.

Generative AI fits into the same workflow at the ends: a person asks in plain language ("why is Building B slow?"), and the assistant queries the data and summarizes the answer, or drafts a change for review. Cisco, like other vendors, now builds both kinds of feature into its management platforms; the details change quickly from release to release.

How to configure the network for it

There is no "AI" command on a router. What you configure is the data the AI depends on. Bad data means bad answers, so two things matter most: every device must send its data, and every timestamp must be accurate so events can be lined up. On SW1 in the lab (R1 at 10.10.1.1 provides time):

ntp server 10.10.1.1 clock timezone UTC 0 service timestamps log datetime msec show-timezone

Accurate, synchronized time with milliseconds on every log message, so events from different devices can be correlated.

logging host 10.10.0.10 logging trap informational logging source-interface Vlan99

Send Syslog (UDP 514) at levels 0-6 to the analytics platform, always from the management address.

snmp-server enable traps

Send SNMP traps for events (with the SNMPv3 user and a snmp-server host line from the controller lesson).

Details of each are in NTP on Cisco IOS and SNMP and Syslog.

How to verify it

Example output · based on Cisco documentation; exact format varies by platform and software version
SW1#show ntp status
Clock is synchronized, stratum 3, reference is 10.10.1.1
...
synchronized is the line that matters. An unsynchronized clock makes events look like they happened in the wrong order, which misleads both people and AI.
Example output · based on Cisco documentation; exact format varies by platform and software version
SW1#show logging | include Trap|Logging to
    Trap logging: level informational, 58 message lines logged
        Logging to 10.10.0.10  (udp port 514, audit disabled,
Syslog goes to the collector at 10.10.0.10 at level informational (6) and above.

Using generative AI safely: a worked example

The colleague in the opening story asked an assistant: "How do I add VLAN 30 to the trunk on Gi1/0/24?" The answer looked reasonable:

interface GigabitEthernet1/0/24 switchport trunk allowed vlan 30

What the AI suggested. Without the word add, this REPLACES the whole allowed list with just VLAN 30.

Example output · based on Cisco documentation; exact format varies by platform and software version
SW1#show interfaces trunk
Port        Mode             Encapsulation  Status        Native vlan
Gi1/0/24    on               802.1q         trunking      1

Port        Vlans allowed on trunk
Gi1/0/24    30

Port        Vlans allowed and active in management domain
Gi1/0/24    30

Port        Vlans in spanning tree forwarding state and not pruned
Gi1/0/24    30
Only VLAN 30 is left. Users (VLAN 10), phones (VLAN 20) and the switch's own management VLAN 99 are cut off, so the colleague also lost their SSH session and needs the console to fix it.
interface GigabitEthernet1/0/24 switchport trunk allowed vlan add 30

The correct command: add VLAN 30 to the existing list.

A safe routine for any AI-drafted change:

  1. Give context: platform, software version, what the interface does now.
  2. Read every line and make sure you understand what it does. If you can't explain it, don't paste it.
  3. Check the command reference for your platform and version.
  4. Test in a lab, or with a dry run such as Ansible --check.
  5. Apply through normal change control, with a rollback plan.
  6. Verify with show commands afterwards.

What goes wrong and how to handle it

ProblemWhat it looks likeHow to handle it
HallucinationA command or option that doesn't exist, or syntax from another vendorCheck against documentation; use ? on the CLI
Out-of-date knowledgeAdvice for an old software versionState your version; verify in current docs
Bad baselineA model that learned during an outage treats broken as normalReview baselines; retrain after known incidents
False positivesAlerts for harmless changes, leading to alert fatigueTune sensitivity; give feedback where the tool supports it
False negativesA real problem that stays within the "normal" rangeKeep basic hard thresholds and checks too
Data leakageConfigs or logs with secrets sent to an external AI serviceUse approved tools; strip passwords, keys and community strings
Missing dataDevices not sending logs, wrong clocksVerify NTP and Syslog on every device

Common mistakes

  • Treating AI output as fact. It is a suggestion until you have verified it.
  • Mixing up the terms: predictive AI analyses and forecasts; generative AI creates content.
  • Thinking machine learning needs no data work. Without clean, complete, time-synchronized data it learns the wrong things.
  • Letting an AI make changes directly in production with no review or rollback.
  • Pasting configurations with secrets into public AI tools.

💡 Exam tip: the CCNA v1.1 blueprint added AI and machine learning in network operations. Expect conceptual questions: tell predictive AI (baselines, anomaly detection, forecasting) from generative AI (creating text and configs from prompts), know that ML learns from data rather than fixed rules, recognize supervised vs. unsupervised learning, and know the limits, especially hallucination and the need for human verification. No commands are tested.

Key takeaways

  • AI is the broad field; machine learning learns patterns from data.
  • Predictive AI: dynamic baselines, anomaly detection, forecasting, root-cause hints.
  • Generative AI: assistants that summarize, answer questions and draft configs.
  • Good AI needs good data: configure Syslog, SNMP, flows and accurate NTP.
  • AI output can be confidently wrong. Review, test, then verify on the device.

Check yourself

Predict · scenario 1

A platform learns that a link normally runs at 5% on Sunday nights and alerts when it hits 40%. What is this?

Predict · scenario 2

An engineer types 'summarize last night's errors on SW2' and gets a paragraph of text. Which kind of AI is this?

Predict · scenario 3

An ML model is trained on thousands of flows, each already labelled with its application. What type of learning is it?

Predict · scenario 4

An AI assistant suggests 'switchport trunk allowed vlan 30' to add VLAN 30 to a busy trunk. What happens if you paste it?

Predict · scenario 5

Correlation of events across devices keeps giving wrong root causes. Which device setting is most likely to blame?

FAQ

What is the difference between AI and machine learning?
Artificial intelligence is the broad goal of making computers do tasks that normally need human judgement. Machine learning is one way to get there: instead of a programmer writing every rule, the system learns patterns from example data. Most AI features in network products today are built on machine learning.
What is the difference between predictive and generative AI?
Predictive AI analyses existing data to classify it or forecast what comes next: this traffic is abnormal, this link will be full in six weeks, this power supply is likely to fail. Generative AI creates new content, such as text, a configuration snippet or a summary of a log file, usually in answer to a question in plain language.
Will AI replace network engineers?
Not on current evidence. AI tools are good at sifting large amounts of data and drafting text, which saves time. They can also be confidently wrong, and they don't carry responsibility for an outage. Someone still has to understand the network well enough to judge the suggestions, which is exactly what the CCNA teaches.
Is it safe to paste my router configuration into a public AI chatbot?
Usually not. Configurations contain IP plans, usernames, password hashes and sometimes keys. A public service may store or learn from what you send. Follow your organisation's policy, use approved tools, and remove secrets before sharing anything.