A real-life situation
Monday, 9 a.m. Users in one building say Wi-Fi is "slow". Overnight the network produced thousands of Syslog messages and millions of flow records. Reading them by hand would take a day. Your controller's analytics page already shows a note: "Client onboarding time in Building B is far above its normal Monday level; most failures are DHCP timeouts." That points you straight at the DHCP server.
Ten minutes later a colleague asks an AI chatbot how to add a VLAN to a trunk, pastes the answer into a switch, and cuts off every VLAN on the link. Both stories are AI in network operations. One saved time; the other caused an outage. This lesson explains what these tools are, what they are good at, and where they fail.
What it is
- Artificial intelligence (AI): the broad field of making computers do tasks that normally need human judgement, such as spotting patterns, making predictions or writing text.
- Machine learning (ML): a part of AI where the system learns patterns from data instead of following rules a programmer wrote. Feed it months of interface counters and it learns what "normal" looks like.
Machine learning comes in three main types:
| Type | How it learns | Network example |
|---|---|---|
| Supervised | From examples that are already labelled with the right answer | Classify traffic flows by application after training on flows labelled "video", "voice", "backup" |
| Unsupervised | Finds structure in data with no labels | Group devices that behave alike; flag a device that suddenly behaves unlike its group |
| Reinforcement | Tries actions and learns from rewards or penalties | Research into tuning radio settings or traffic paths; less common in everyday tools |
Network products use AI in two broad ways:
| Predictive AI | Generative AI | |
|---|---|---|
| What it does | Analyses data to detect, classify and forecast | Creates new content: text, summaries, config drafts |
| Input | Telemetry: counters, logs, flows, client events | A question or instruction in plain language, plus context |
| Output | "This is abnormal", "this link will be full in six weeks" | "Here is a summary of these logs", "here is a config that might do this" |
| Network uses | Dynamic baselines, anomaly detection, capacity planning, failure prediction, root-cause hints | Chat assistants, natural-language queries ("which switches run old software?"), log and ticket summaries, drafting scripts and configs |
| Main risk | False alarms, or missing a real problem | Hallucination: fluent, confident, wrong output |
Generative AI tools are usually built on a large language model (LLM): a model trained on huge amounts of text to predict likely next words. That is why its answers read well, and also why they can be wrong in ways that sound right.
Why it works that way
Traditional monitoring uses static thresholds: "alert if CPU is above 80%" or "alert if a link is above 90%". These are crude. A backup link at 95% at 2 a.m. may be perfectly normal, while 40% on a link that usually sits at 5% may be an attack.
Machine learning builds a dynamic baseline for each metric: what is normal for this link, at this hour, on this day of the week. It then flags anomalies, values far from that baseline. With thousands of devices, only a machine can watch every metric that closely. It can also correlate events, noticing that 200 separate alerts started within a minute of one switch reboot.
Generative AI works for a different reason: it has seen a lot of text, including documentation and configurations, so it can draft and summarize quickly. But it predicts what is likely, not what is true for your network, platform and software version. It doesn't know your topology unless you tell it.
How it works step by step
A predictive AI feature, such as the assurance analytics in a controller, follows a pipeline:
In the course lab, Catalyst Center plays the analytics role:
- 1. 1. Devices stream data. SW1, SW2 and R1 send Syslog, SNMP data and flow records to Catalyst Center all day.
- 2. 2. The controller learns normal. Over weeks it learns each metric's normal range for each hour and day.
- 3. 3. An anomaly is raised. SW2's uplink errors jump far above baseline. Related alerts are grouped into one issue with a suggested cause.
- 4. 4. The engineer verifies. The engineer checks show interfaces on SW2 before acting. The AI pointed the way; the device confirms it.
Generative AI fits into the same workflow at the ends: a person asks in plain language ("why is Building B slow?"), and the assistant queries the data and summarizes the answer, or drafts a change for review. Cisco, like other vendors, now builds both kinds of feature into its management platforms; the details change quickly from release to release.
How to configure the network for it
There is no "AI" command on a router. What you configure is the data the AI depends on. Bad data means bad answers, so two things matter most: every device must send its data, and every timestamp must be accurate so events can be lined up. On SW1 in the lab (R1 at 10.10.1.1 provides time):
ntp server 10.10.1.1
clock timezone UTC 0
service timestamps log datetime msec show-timezoneAccurate, synchronized time with milliseconds on every log message, so events from different devices can be correlated.
logging host 10.10.0.10
logging trap informational
logging source-interface Vlan99Send Syslog (UDP 514) at levels 0-6 to the analytics platform, always from the management address.
snmp-server enable trapsSend SNMP traps for events (with the SNMPv3 user and a snmp-server host line from the controller lesson).
Details of each are in NTP on Cisco IOS and SNMP and Syslog.
How to verify it
SW1#show ntp status Clock is synchronized, stratum 3, reference is 10.10.1.1 ...
SW1#show logging | include Trap|Logging to Trap logging: level informational, 58 message lines logged Logging to 10.10.0.10 (udp port 514, audit disabled,
10.10.0.10 at level informational (6) and above.Using generative AI safely: a worked example
The colleague in the opening story asked an assistant: "How do I add VLAN 30 to the trunk on Gi1/0/24?" The answer looked reasonable:
interface GigabitEthernet1/0/24
switchport trunk allowed vlan 30What the AI suggested. Without the word add, this REPLACES the whole allowed list with just VLAN 30.
SW1#show interfaces trunk Port Mode Encapsulation Status Native vlan Gi1/0/24 on 802.1q trunking 1 Port Vlans allowed on trunk Gi1/0/24 30 Port Vlans allowed and active in management domain Gi1/0/24 30 Port Vlans in spanning tree forwarding state and not pruned Gi1/0/24 30
interface GigabitEthernet1/0/24
switchport trunk allowed vlan add 30The correct command: add VLAN 30 to the existing list.
A safe routine for any AI-drafted change:
- Give context: platform, software version, what the interface does now.
- Read every line and make sure you understand what it does. If you can't explain it, don't paste it.
- Check the command reference for your platform and version.
- Test in a lab, or with a dry run such as Ansible
--check. - Apply through normal change control, with a rollback plan.
- Verify with
showcommands afterwards.
What goes wrong and how to handle it
| Problem | What it looks like | How to handle it |
|---|---|---|
| Hallucination | A command or option that doesn't exist, or syntax from another vendor | Check against documentation; use ? on the CLI |
| Out-of-date knowledge | Advice for an old software version | State your version; verify in current docs |
| Bad baseline | A model that learned during an outage treats broken as normal | Review baselines; retrain after known incidents |
| False positives | Alerts for harmless changes, leading to alert fatigue | Tune sensitivity; give feedback where the tool supports it |
| False negatives | A real problem that stays within the "normal" range | Keep basic hard thresholds and checks too |
| Data leakage | Configs or logs with secrets sent to an external AI service | Use approved tools; strip passwords, keys and community strings |
| Missing data | Devices not sending logs, wrong clocks | Verify NTP and Syslog on every device |
Common mistakes
- Treating AI output as fact. It is a suggestion until you have verified it.
- Mixing up the terms: predictive AI analyses and forecasts; generative AI creates content.
- Thinking machine learning needs no data work. Without clean, complete, time-synchronized data it learns the wrong things.
- Letting an AI make changes directly in production with no review or rollback.
- Pasting configurations with secrets into public AI tools.
💡 Exam tip: the CCNA v1.1 blueprint added AI and machine learning in network operations. Expect conceptual questions: tell predictive AI (baselines, anomaly detection, forecasting) from generative AI (creating text and configs from prompts), know that ML learns from data rather than fixed rules, recognize supervised vs. unsupervised learning, and know the limits, especially hallucination and the need for human verification. No commands are tested.
Key takeaways
- AI is the broad field; machine learning learns patterns from data.
- Predictive AI: dynamic baselines, anomaly detection, forecasting, root-cause hints.
- Generative AI: assistants that summarize, answer questions and draft configs.
- Good AI needs good data: configure Syslog, SNMP, flows and accurate NTP.
- AI output can be confidently wrong. Review, test, then verify on the device.
Check yourself
A platform learns that a link normally runs at 5% on Sunday nights and alerts when it hits 40%. What is this?
An engineer types 'summarize last night's errors on SW2' and gets a paragraph of text. Which kind of AI is this?
An ML model is trained on thousands of flows, each already labelled with its application. What type of learning is it?
An AI assistant suggests 'switchport trunk allowed vlan 30' to add VLAN 30 to a busy trunk. What happens if you paste it?
Correlation of events across devices keeps giving wrong root causes. Which device setting is most likely to blame?