IT playbook · AI Employee: Ivy

Network Troubleshooting

Median first-response < 3 minutes

The problem

Network issues are the black-box of IT support. "WiFi is slow," "VPN keeps dropping," "can't reach the internal wiki" — every case starts as user-perception, and network diagnostics require correlating data from the user's device + the WiFi controller + the DNS server + the VPN concentrator + the firewall + the internal service. Manual triage takes 30-90 minutes and often ends with "try restarting your router."

At a glance
Trigger
Chat
Approvals
None
What it does
Read-only
Systems
Datadog · Meraki · Cisco · RAG on network runbook
How it feels in production

An hour-by-hour walkthrough.

Priya messages Ivy: "VPN keeps disconnecting during video calls." Ivy pulls the correlated network context in 30 seconds: - Priya's device: connected to home WiFi (SSID Priya_5G), signal strength -47dBm (strong) - VPN client logs: 4 disconnect events in last 2 hours, each after 15-20 minutes of stable connection - Network path: home ISP → VPN concentrator (Berlin PoP) - ISP + region signals: no known outage, but Berlin PoP CPU at 78% - Similar reports: 8 other users at Berlin PoP reporting VPN instability in last 4 hours Ivy diagnoses: "VPN Berlin PoP is under load, causing periodic drops for you and other users. Immediate options: 1. Switch to Frankfurt PoP (currently under-loaded); I can push the config change now 2. Wait — network ops is aware, expected to resolve in 30 minutes For the video-call reliability, I recommend switching. Approve and I'll change your VPN config?" Priya approves; VPN reconfigured; video call resumes without drops. Behind the scenes, Ivy has also promoted the Berlin PoP load issue to network ops with the correlated user impact + suggested load-balancing action.
How it works

Step by step.

  1. 01

    Correlate network signals across every hop

    Device WiFi + VPN client + ISP + VPN concentrator + firewall + destination service. End-to-end path visibility.

    MDM · VPN telemetry · ISP monitoring · Cloud network monitoring
  2. 02

    Diagnose from correlated data

    Not device-only view. Correlates similar reports; identifies infrastructure-side issues that individual devices can't diagnose.

    Reasoning · Similar-report clustering · Infrastructure telemetry
  3. 03

    Present remediation options with tradeoffs

    User remediation (switch PoP, restart, DNS change) with expected impact. Infrastructure escalation to network ops for shared-cause issues.

    Slack · Teams · VPN config · Network ops queue
  4. 04

    Execute approved user remediation

    Config changes via MDM. Restart guidance. DNS cache clear. User visible + approved.

    MDM · VPN client config · Endpoint config
  5. 05

    Escalate shared-cause to network ops

    Infrastructure issues promoted with correlated impact + suggested action. Network ops sees the pattern + user impact together.

    Network ops · PagerDuty · Incident tracker
Systems and wiring

What you connect to make this run.

MDM · Endpoint telemetry

read

Device network state: WiFi signal, IP config, DNS resolution, active connections.

VPN telemetry · Cisco AnyConnect · Palo Alto GlobalProtect · Tailscale

read+write

VPN client logs + concentrator state. Config-push for PoP changes.

Cloud network monitoring · ISP + regional signals

read

Infrastructure-side visibility. PoP load, ISP outages, regional degradations.

Network ops · PagerDuty

write

Escalation for shared-cause issues with correlated impact + suggested action.

What changes

Before and after, honestly.

Time from user report to network diagnosis
Before
30-90 minutes
After
Under 60 seconds
% of network issues resolved without network-ops escalation
Before
20-40%
After
65-80%
Correlated infrastructure issues detected proactively
Before
10-30% via user reports
After
80%+ via correlation
IT hours per week on network triage
Before
15-30 hours
After
3-6 hours
Frequently asked

Answers about this playbook.

What about home-network issues (router, ISP)?

Ivy diagnoses reach beyond corporate network to identify home-side vs. corporate-side. Home-side issues get guidance (router restart, ISP contact) rather than corporate escalation.

How does it handle office WiFi issues?

Same correlation model with office-specific signals (AP density, controller health, DHCP pool). Facilities-team escalation for physical issues.

Can it diagnose latency-specific issues (game lag, video quality)?

Latency + jitter analysis included. Application-specific advice: video-conferencing has different tolerances than file transfers.

What about VPN-specific edge cases (split-tunnel, always-on)?

VPN policy-aware diagnostics. Split-tunnel routing questions, always-on-VPN failures each have specific diagnostic paths.

How does it interact with zero-trust network access (ZTNA)?

ZTNA (Cloudflare Access, Zscaler ZPA) telemetry integrated. Different from traditional VPN diagnostics; posture-aware access decisions surfaced.

See it run on your data.

Free plan, no credit card. Connect the systems this playbook needs and run it against a past event first.