The AIOps engine

Understand incidents, not just alerts

OverWatch baselines every metric, flags real anomalies, reads your logs for known failure patterns, and correlates it all into a single incident — with the entity, probable cause, and blast radius attached.

Adaptive baselines Anomaly detection Root-cause correlation
OverWatch Command — fleet posture, correlated incidents, and live pressure across the estate
30-second tour

Watch the engine work

Live collection, adaptive baselines, a correlated incident, a SOC brute-force caught in seconds, and Walker — end to end, on sample data.

Full walkthrough · 14 minutes

The whole platform, end to end

A narrated walkthrough across a 367-device estate: Command, discovery, device detail, AIOps incidents, the fabric, virtualization, storage, capacity, security, policy, and Walker keeping it stable. Sample data throughout.

The pipeline

Collect → baseline → detect → correlate

A closed pipeline from raw telemetry to an incident with a name, a cause, and a scope.

Collect

Metrics, interfaces, logs, and events stream in from ICMP, SNMP, vendor APIs, syslog, and an HTTP event collector into one model.

Baseline

Each metric learns its own normal rhythm and keeps adapting to it — a self-updating baseline (an exponentially weighted moving average, or EWMA) instead of brittle static thresholds.

Detect

Flags readings that stray far from that normal — scored by how many standard deviations out they sit, with sensible floors for CPU, memory, disk, latency, and event rate — while log signatures catch known failure patterns.

Correlate

Related alerts, anomalies, and log hits group by shared entity and time into one incident with a probable root cause.

Baselines & anomalies

Detection that learns what normal looks like

Static thresholds miss slow drift and cry wolf on spiky-but-fine systems. OverWatch tracks an adapting baseline per stream and scores how far reality has moved from it — measured in standard deviations.

Self-updating baseline (EWMA)

Learns each metric's normal range, warms up quickly, and keeps adapting — no thresholds to hand-tune.

Deviation scoring with guardrails

Scored in standard deviations from baseline, with warning and critical bands, absolute floors, and a cooldown so one blip isn't a hundred alerts.

Threshold rules too

When you want a hard line — device-down, high latency, packet loss, interface-down, high CPU/memory — set it explicitly.

Silence is a signal

A source that suddenly stops reporting becomes an incident on its own — caught within about a minute, and cleared the moment it resumes.

OverWatch device detail — live metrics, adaptive baselines, and interface state
Incident correlation

One incident, not a thousand pings

A flapping uplink can throw hundreds of alerts across neighbors. OverWatch groups them by entity and time window into a single incident, escalates severity as it grows, tracks the blast radius, and auto-resolves when it clears.

Root cause & blast radius

The originating entity plus the count of impacted devices and segments.

Log-aware

Signatures for interface/BGP/OSPF drops, spanning-tree loops, OOM, SMART pre-fail, RAID degraded, thermal, and brute-force auth feed the same incident.

"What's on fire" at a glance

The NOC view ranks open incidents by severity and impact in real time.

OverWatch Incidents — AIOps-correlated incidents with severity, root cause, and blast radius
Walker · natural-language assistant

Ask questions. Get proposed actions.

Walker answers plain-language questions about your estate, incidents, and firewall policy — grounded in a built-in knowledge base of hundreds of metrics across dozens of domains, so it can explain what a number means and whether it's worth worrying about. It answers counts and rankings from your live data, proposes safe, reviewable actions, and runs keyless and local by default.

Grounded, not guessing

A curated metric knowledge base plus your live devices, incidents, and policy — so answers are explanatory and specific, never hallucinated.

Human-in-the-loop

Walker proposes; you approve. Nothing changes without your say-so.

Local by default

Runs without an external key; add an assisted provider only if you want to.

Walker answering from live estate data with evidence, a proposed change, and an honest gap
Discovery & topology

The fabric draws itself

OverWatch reads LLDP, CDP, and ARP and resolves the real hierarchy — Core → MDF → IDF → Edge — with guests routed through their host and only gateway-role devices counted as valid internet exits. Egress is proven from session logs, not from a diagram you drew.

Auto-classified tiers

Every node placed by neighbor depth — no manual topology to maintain.

Egress attribution

Which subnet leaves through which device and NAT IP, from the flow table itself.

IPAM built in

Subnets, host inventory, and DHCP lease ingest alongside the map.

OverWatch fabric topology — Core to MDF to IDF to Edge with session-log egress attribution
Works with your stack

Collects from what you already run

First-party integrations with the platforms you actually operate — deep collection over each vendor's own API — plus SNMP discovery and monitoring across the wider enterprise fleet.

ProxmoxProxmox
Ubiquiti UniFiUbiquiti / UniFi
Palo Alto NetworksPalo Alto
TrueNASTrueNAS
Frigate NVRFrigate
Windows ServerWindows Server
SNMP discovery & monitoring
Cisco VMware NetApp Juniper, Arista, Aruba & more

See the engine on real telemetry

In a 30-minute demo we'll show live collection, baselines and anomalies, a correlated incident, and Walker — all on sample data.