Skip to content

Alerting

Alert quality is a first-class concern in Yagra. The goal is not to send alerts. It is to send the right alerts: one page per real incident, no flood when a site drops, no 3 a.m. noise from a metric grazing its threshold for a single sample.

Hysteresis, flapping detection, dependency suppression, and deduplication are always in the delivery path. They are not optional extras you switch on later, and there is no way to bypass them to “just send the alert”.

Escalation and on-call scheduling are deliberately delegated to external tools (PagerDuty, Jira Service Management). Yagra emits a clean fire/resolve lifecycle signal and holds no escalation scheduler of its own.

Three things can raise an alert:

  • a liveness failure — a device stops answering its ICMP probe (or a URL/DNS monitor stops succeeding). This is a seeded alert rule, not built-in behavior: its breach count is editable, it can be overridden per profile, group or node, and deleting it stops node-down paging (see Monitoring);
  • a threshold breach — a gauge metric crosses a rule’s warning or critical bound (see Monitoring for the threshold model);
  • a matching event rule — a passive event (syslog, trap, webhook) matches an operator-defined rule (see Event-raised alerts below).

Whatever the trigger, every check on every node carries a committed state:

State Meaning
ok The check is healthy
warning A threshold rule’s warning bound is breached
critical A threshold rule’s critical bound is breached
unreachable Liveness has failed — the node does not answer
unknown Yagra has no current opinion (no recent samples)
maintenance The node is inside a maintenance window

Only the problem states — warning, critical, unreachable — carry a severity and produce an alert. ok, unknown, and maintenance never do. A node Yagra can’t see yet, or one under planned work, is not an incident.

You will see unknown briefly on newly added nodes and right after a core restart, before the alert engine has formed an opinion on every node. It resolves itself as samples arrive.

An alert is raised when a check commits a problem state, after the hysteresis below, and clears when the check recovers through the same path.

Three other things clear an alert, and none of them means the fault is fixed. A node entering a maintenance window clears the alerts already open on it. Deleting a threshold rule clears the alert that rule had raised. And an alert clears when the series behind it has gone six hours with no sample while the node itself is demonstrably still reporting — read that one as Yagra can no longer measure this, not as a recovery. The last two are swept once every five minutes. Each is written to alert history as an ordinary resolve, so the incident in PagerDuty or Jira Service Management closes with it.

An interface utilization or traffic alert is not closed by its readings stopping. A port whose values stop arriving keeps its alert open: an SNMP agent that died while the device still answers ping, a walk cut short, or a Meraki window the Dashboard never filled all look like that, and the open alert is what tells you. The alert clears when a reading shows the port back under its bound, through the rule’s dwell. An alert on an interface that was removed stays open until you delete its rule or its node.

Repeated fires of the same problem are deduplicated and grouped into one incident rather than delivered as a stream of duplicates. Every fire and clear is recorded in an append-only alert history, with the measurement that caused it captured at fire time.

A raw sample crossing a threshold does not flip the committed state. The breach must persist for a configured number of consecutive samples — the dwell — before the state commits and an alert fires.

Any sample that returns to the committed state resets the candidate, so an interrupted streak starts the count over.

This is what keeps a metric that grazes its bound once, or a single dropped ping, from paging anyone. A dwell of three, for instance, means three breaching samples in a row at the node’s poll interval. One recovery sample in between and the count starts again.

Interface utilization alerts and alerts on computed metrics (memory and disk used %) are evaluated once a minute. On a node polled less often than that, their count covers that many polls, not that many minutes. Nodes polled faster keep counting minutes.

Counting in samples rather than wall-clock time also has an architectural consequence. Alert evaluation always runs against the live sample stream. History backfilled later — for example after a remote poller reconnects and replays its buffer — is never re-evaluated into alerts.

Replaying a backlog would fabricate transitions for incidents that already resolved. So metrics fill in, and alerts resume from now.

A link that goes up and down every few minutes generates a technically-correct alert each time, and each one is noise.

Yagra counts committed state transitions per check inside a sliding time window. A check reaching the transition threshold within the window is flagged as flapping on its transitions.

The window follows the node’s poll interval. It is 20 polls long, and never shorter than 10 minutes. Five state changes inside it mark the check as flapping. A fixed ten-minute window could never hold five changes on a node polled every five minutes.

Flapping checks surface on the Flapping watchlist dashboard widget, which names the node (the id is on hover) and shows what the flapping check measures. The unstable link then gets fixed as one problem, instead of being re-triaged on every bounce.

For deeper digging, the Troubleshoot catalog includes a flap analysis that scans a node, a group, or the whole fleet for reachability and link-state churn. A companion analysis finds event rules repeatedly firing and clearing on the same node.

When a site router dies, everything behind it goes dark. Without suppression, every one of those nodes pages you.

Yagra drives suppression from the dependency graph. Each node can name its upstream, and the graph is visualized on the Network map.

When a node goes unreachable and its upstream chain is also down, the alert is attributed to the highest down ancestor — the root cause — and rolled into that parent’s incident.

The child’s alert still fires and is still visible in the UI and the history. Only the duplicate notification is suppressed. The topology views and the dependency / root-cause dashboard widget show the attribution, so one incident reads as one incident with its blast radius attached.

Suppression is re-evaluated whenever a parent’s state changes — event-driven, not on a timer — and it works in both directions:

  • A parent that goes down after its children were already alerting re-groups those existing alerts under its incident. Their standalone pages are closed.
  • A parent that recovers while a child is still down re-pages that child on its own. It is no longer explained by the parent.

Two deliberate limits keep suppression honest:

  • Only liveness alerts are ever suppressed. Threshold alerts are not, since a reachable node with a real threshold breach deserves its page.
  • Alerts raised by passive events skip suppression too, because a device that just emitted an event is evidently reachable.

Dependency edges are editable from the WebUI. The node-detail header’s Dependency… action sets or clears a node’s upstream.

Topology ▸ Dependency view lists every node with its upstream, live status, and current root-cause attribution, editable inline. Self-dependencies and cycles are rejected.

Typing an upstream per node does not scale, and a graph nobody maintains is a graph that suppresses the wrong things.

Yagra can derive the dependency graph from the network map instead. Direction comes from each node’s distance to a poller, so whatever sits one hop closer is its upstream.

Topology ▸ Dependencies has a three-position mode switch. Nothing moves between the positions on its own — an upgrade lands on the mode you were already on.

  • The hand-authored graph — the default, and where every existing deployment stays.

  • Comparing — changes nothing about alerting. It shows, node by node, where the derived graph and the one you maintain disagree.

    It also gives the two numbers that decide whether the switch is safe: how many active alerts the derived graph would newly suppress, and how many it would stop suppressing. The first is the risky direction — each one is an alert that would stop being raised.

  • The derived graph — hands suppression over.

There are two things the derived graph can express that a hand-typed parent_id cannot.

The first is multiple parents. A node gets every neighbour one hop closer to a poller, so a server reached through a redundant pair of routers has both as parents, and keeps alerting while either is alive. That is what the “suppressed only when every parent is down” rule was always written for.

The second is Never suppress, which takes a single box out of derived suppression entirely. Its alert then always stands on its own, whatever the graph says. That switch only ever removes suppression, so it cannot cause an outage to go unreported.

Where the derivation gets a link wrong, record the decision rather than working around it. A link can be pinned into existence, hidden, or have its direction declared, and those always beat what was derived on every recomputation.

Both keep a page from reaching anyone, and that is where the resemblance ends. They cut the pipeline at different points, and only one of them changes what the record says afterwards.

  • A maintenance window stops the alert. A node inside an open window observes the maintenance state, which carries no severity, so nothing fires and an alert that was already open resolves.

    This is how you say this is not a fault: planned work, a scheduled reboot, a site being rewired. Opening and closing windows is an Operator-and-up action (the manage-maintenance permission).

  • A mute stops only the delivery. The alert is raised exactly as it otherwise would be. It appears on Active alerts, the node’s state changes, it can be acknowledged, and it is written to alert history. The only thing that does not happen is the notification.

    This is how you say this is a fault, we know, do not wake anyone over it: a device someone is already working on, a noisy check waiting for its threshold to be fixed. Muting, like acknowledging, is part of the Operator role’s respond-to-incidents permission.

Maintenance window Mute
Alert is raised No Yes
Node’s displayed state maintenance Unchanged — a node that is down still reads down
Appears on Active alerts No Yes
Written to alert history No Yes
Can be acknowledged Nothing to acknowledge Yes
Notification delivered No No
Scope A node, a device profile, a tag group, or a folder group (recursively) A node, or a folder group (recursively)
Narrower than a whole node — Yes — a single check on that node
Timing Scheduled: named, with a start and an end, and can be created in advance From now until a time you choose

The choice matters most a month later. Availability figures and the alert history are built from what was recorded.

A maintenance window keeps planned work out of the incident record entirely. A mute leaves the whole story in place, including the fact that nobody was told at the time.

Reach for a window when the outage is expected and should not count. Reach for a mute when the outage is real and you only want the phone to stop.

Two details worth knowing:

  • A mute never holds a resolve back. If an alert fires, is muted afterwards, and then recovers, the resolve is still delivered. Otherwise a mute placed mid-incident would leave a PagerDuty or Jira Service Management incident open with nothing left to close it.
  • Maintenance silences event-raised alerts too. A syslog message or trap arriving from a node inside an open window raises nothing, exactly as a poll result would not.

Both are first-class, audited objects with their own settings pages — Alerts ▸ Maintenance windows and Alerts ▸ Mutes — not a notification filter bolted on at the end.

Setting and releasing suppression from the tree

Section titled “Setting and releasing suppression from the tree”

Both can be set straight from the inventory. Right-click a node or a folder group in Nodes ▸ All nodes and pick a duration, or Custom… for the full form.

A suppressed row carries a marker — 🔧 for maintenance, 🔕 for a mute — and the marker is a button.

Clicking it opens a panel naming what is silencing that row: the window’s name, whether it comes from the row itself or from a folder group above it, and when it stops. A row can be covered by more than one thing at once, and the panel lists each separately.

What the panel offers depends on where the suppression comes from, because the three cases are not the same action:

  • A window that names the row is ended now rather than deleted. Its end time moves to the current moment, so it survives on the Maintenance windows page as a record of the work that actually happened, and is swept later by the “clear ended” button. A mute naming the row is simply lifted.

  • A window or mute the row only inherits cannot be ended from that row without releasing every sibling under it. So a node is offered a release instead: take this node out of maintenance, or unmute this node. That one node returns to normal alerting while the rest of the group stays covered.

    The release expires by itself when the coverage it was carved out of ends, including when that coverage stops earlier than planned. A release can therefore never outlive its reason and silently exclude the node from the next window.

  • A group covered by an ancestor shows the cause read-only and names the group to release it on. Ending an ancestor’s window from a child row would unsilence a set you cannot see from there.

The markers distinguish the two states. A suppression in force is a filled chip — blue for maintenance, yellow for a mute. A row released from one is drained of the colour and outlined in a dashed border instead.

So a released node is visibly not the same as an unsuppressed one. It is still inside a window’s reach, and one click puts it back.

Once an alert survives hysteresis, suppression, and maintenance checks, notification delivery decides where it goes. Yagra delivers to:

Channel Notes
Webhook Generic HTTP webhook, with retry
Email (SMTP) With retry; relay, sender, and credentials configurable
PagerDuty Events API v2
Jira Service Management Alerts API

Every channel receives the full lifecycle — fire and resolve both. The external tool’s incident therefore closes when Yagra’s alert clears, instead of leaving stale incidents to be tidied by hand.

Acknowledgement performed in an external tool is reflected back into Yagra read-only. Routing rules choose which channels are notified, managed on the notification-routing settings page.

Delivery runs off the result-ingest path. A slow or unresponsive notification endpoint can delay its own notifications, which retry, but it can never stall metric ingestion or alert evaluation behind it.

Outbound webhooks are SSRF-guarded the same way monitoring probes are. Loopback, link-local, and cloud-metadata targets are refused.

A node’s tags ride out with the alert, into the field each vendor already provides for them: PagerDuty’s payload.custom_details.yagra_tags (a list) and Jira Service Management’s own tags. “Page the Japan rota for anything tagged JAPAN” is then written once, in those tools’ own routing rules, rather than duplicated in Yagra’s. The tags sent are the node’s effective ones — its own plus everything its inventory folder chain supplies. Webhook and email carry them only if a template asks; see below.

Jira Service Management accepts at most 20 tags per alert, each at most 50 characters. Yagra keeps within those limits. A longer tag is left out rather than shortened, because a shortened tag could match a different rule. Of the rest, the first 20 are sent, the node’s own tags before the ones its folders give it. Keep the tags your JSM rules route on short and few.

Each channel can override the subject and body it sends, written as a Jinja2 template over a fixed set of alert variables — node name, address, group, profile, severity, metric, threshold, the observed value, the node’s effective tags, the name of the table row the alert is about (row_name, such as a memory pool), and whether this is a fire or a resolve. Four more are written for people: title is the alert’s name (SNMP not responding rather than snmp_up), if_name is the name of the port a per-port alert is about, node_label is the node’s name with its address in brackets when the two differ, and alert_label is the alert’s name followed by the port or row. Conditionals work, so wording can differ by severity, by lifecycle event, or by tag:

{% if event == 'resolve' %}Recovered{% else %}{{ severity | upper }}{% endif %}: {{ node_name }} ({{ group }})
{% if 'JAPAN' in tags %}[JP rota]{% endif %} tags: {{ tags | join(', ') }}

A channel with no template sends Yagra’s built-in wording, unchanged. Every kind of channel gets the same title. It names the device and the alert, such as core-sw-01 (192.0.2.11) is critical: SNMP not responding, and a per-port alert also names its port (… : Inbound utilization on GigabitEthernet0/7). The rest depends on who reads it:

  • Webhook and PagerDuty are read by a program, so the body is the alert as JSON. The title is PagerDuty’s summary.
  • Email and Jira Service Management are read by a person, so the body is text. The body repeats the title, then gives one fact per line: node, folder, profile, the alert’s name, metric with its value and threshold, port or row, severity, state, since when, tags, and the alert key.

Jira Service Management also receives those facts as extra properties (details), whether or not the channel has a template. JSM shows them as a table, and its rules can match on them.

Two properties are worth knowing before you write a template:

  • A broken template never costs you a notification. If it cannot be rendered when an alert fires, that field falls back to the built-in wording and the notification still goes out. Causes include a bad filter, an oversized output, or a body that stopped being valid JSON on a channel that sends JSON.

    The fallback is a plain string substitution, not a second trip through the engine that just failed.

    A field whose template renders to nothing, or only to spaces, also keeps the built-in wording. So a template can give the resolve its own text while the fire keeps Yagra’s.

  • Templates are sandboxed. A template can reach nothing but the variables it is handed: no files, no includes, no other templates. Execution is bounded, so a runaway loop cannot stall the notification worker.

The editor lives at Alerts ▸ Notification delivery ▸ a channel ▸ Edit notification template. Its fields follow what the channel sends. Jira Service Management has a title and a body, email a subject and a body, PagerDuty a summary and custom_details, and a webhook one JSON body and no subject.

A channel with no template opens on Yagra’s built-in text, read-only, so you see exactly what is sent today. From there, Edit a copy of this text starts from it and Start from empty fields starts from nothing. A channel with its own template says so at the top. Go back to built-in text… deletes the template after a confirmation.

For email and Jira Service Management, the editor is visual. Variables appear as tags in the text, with a name and an explanation in the operator’s language. The fire, the resolve and the roll-up under an upstream outage each have a tab; a tab you leave alone sends the fire text. A template the tags cannot show opens in the code editor, with the reason. Webhook and PagerDuty bodies must be JSON, so those channels always use the code editor. In the code editor, Free layout makes line breaks and indentation for reading only: they are not sent. Either way, a live preview renders against four sample alerts. The dialog can be resized from its edges.

To check a channel end to end, use its Send test notification action on the same page. Yagra sends one sample alert through that channel, with [TEST] at the start of the subject, and shows whether it was delivered. A PagerDuty or Jira Service Management test opens a real incident, so the on-call is notified once. Yagra closes that incident a few seconds later. The test does not retry, and it works on a disabled channel.

A routing rule can be edited after it is made: its Edit rule action changes the name, the severity and the channels, and leaves the rule enabled or disabled as it was.

The same page has a Delivery log. It holds one row per notification sent to a channel: fire, resolve, roll-up, and test sends. Each row says whether the notification arrived. When it did not, the row says where it failed:

Where it failed Meaning
Yagra Stopped before anything was sent. For example, the target address was refused.
Network Sent, but nothing answered: a timeout, a refused connection, DNS, or TLS.
Receiving service The service answered and refused.

A failed row carries the status the service answered with and the first 512 characters of its answer. The channel’s URL, host and key are replaced by <redacted>. Open a row to see every retry.

Filter by channel, kind, result, where it failed, and time. The environment default route has a filter entry of its own. A channel’s row menu has Show delivery log, which opens the log filtered to that channel.

Reading the log needs Manage the deployment, which only an Admin holds. The same rows are served at GET /api/v1/notification-deliveries and by the MCP tool get_notification_deliveries. They are kept for the alert-history retention window: Settings ▸ Monitoring defaults ▸ Data retention, “Notification delivery log”.

Recording never makes a notification wait. If PostgreSQL falls behind, rows are dropped and counted in yagra_notification_delivery_log_dropped_total. A duplicate the dispatcher suppressed is not recorded, because nothing was sent.

For a fixed deployment-wide route without touching the UI, two environment-level channels exist. YAGRA_WEBHOOK_URL defines an always-on default webhook that fires for every alert alongside the configured channels. The YAGRA_SMTP_* variables define a default email route. See the configuration reference.

Alerts don’t only come from polling. Passive events — syslog messages, SNMP traps, and inbound webhooks — are matched against operator-defined event rules (substring or regex), and a matching rule can raise an alert.

Incoming SNMP traps have their trap OID resolved to a human-readable name, and a set of built-in trap event rules ships out of the box, so common traps arrive as named, alertable events without configuration. See Passive events for the ingestion side.

Event-raised alerts differ from poll-driven ones in two ways:

  • They clear on a TTL. A one-shot event has no “recovery sample” to observe, so each rule carries an auto-clear TTL. The alert closes when it expires, and there is no manual close.
  • They skip dependency suppression. A device that just emitted an event is reachable, so suppressing its alert under a “down” parent would be wrong.

A rule scoped to a single event stream stays scoped to it. A rule whose source kind this core version does not recognise is left out of the matching engine and logged, rather than silently widening to every stream.

Every alert above depends on poll results arriving. That makes the loss of a poller the one failure the alert engine cannot reason about.

When a poller pool loses its last live poller, its jobs are published to a subject nothing is subscribed to and discarded. The nodes then drift to unknown rather than down. An entire site stops being monitored while every dashboard stays calm.

Yagra watches for this directly. A pool that still holds nodes but has no live poller raises a critical alert of its own, delivered over the same notification channels and routing rules as any other, and closed automatically when a poller returns.

  • It is a full alert, not just a notification. It appears on Active alerts and in alert history, streams live, and renders through your notification templates.

  • It waits five minutes by default. A poller announces its own departure, so an ordinary rolling restart raises the condition instantly. The debounce is what stops that paging anyone. Tune or disable it with YAGRA_POOL_COVERAGE_ALERT_AFTER_SECS (configuration).

  • Its subject is a pool, not a node. This is the one alert that belongs to no device, so it never rolls into a node’s displayed state and it cannot be muted — a mute names a node.

    API clients should branch on subject_kind before reading an alert’s node field as a node id.

  • Meraki-managed nodes are excluded. They are collected per organization rather than assigned to a pool’s ring, so they do not depend on one. The alert below is what covers them.

Two gauges are exported whether or not the alert is enabled — yagra_pools_without_live_poller (unlabelled, the one to alert on or drive a scale-up from) and yagra_pool_nodes_without_live_poller{pool}, which reports 0 for a healthy pool rather than disappearing.

A Cisco Meraki organization has the same problem for a different reason. Its devices are never pinged: what the Dashboard API reports is everything Yagra knows about them. A collect that fails — a revoked key, a Dashboard outage, rate limiting, a poller with no way out — used to publish nothing at all, so every device of the organization kept its last state and nothing alerted.

Yagra watches that directly too. After three availability collects in a row go unanswered, one critical alert is raised about the organization.

  • One alert, not one per device. Its subject is the organization, so it belongs to no node either.
  • It closes on evidence, never on silence. An answered collect closes it; reports merely ceasing to arrive does not. After a restart Yagra knows nothing yet, and knowing nothing closes nothing.
  • The devices keep their last collected state, because they did not fail. Each affected node’s Overview says the Meraki API is not answering, names the reason, and says when the state on screen was collected.
  • Only availability raises. An uplink or traffic collect failing while availability still answers costs readings, not liveness; it is shown on the organization’s page and raises nothing.
  • A group-scoped account sees the alert when one of the organization’s nodes is in a group it can see.

That makes three values an API client can meet in subject_kind: a node, a pool, and a Meraki organization.

  • Active alerts is the triage view. It lists every currently-firing alert, newest first, with severity, the node’s name (id on hover), the alert’s name and how it breached — for example, Ping response time above 100 (was 450) — and its age. Every metric Yagra knows has a name in English and Japanese. The raw metric stays at the end of the line, so it can be matched to its rule. A 0/1 check is named by its fault (SNMP not responding) and shows no condition, and a per-port alert names its port.

    The list is virtualized, so it stays smooth even with thousands of simultaneous alerts during a major outage. The notification bell in the top bar shows the active count and opens this view. Each node’s detail page also shows the alerts currently attached to that node.

  • Acknowledging an alert is an operator action — the respond-to-incidents permission, Operator and up — and marks it as being handled. The acknowledgement can also be cleared. Roles and permissions are covered in Users and SSO.

  • Alert history is the append-only record of every fire and clear. Each entry keeps the metric and condition captured at fire time in its “What” column, so history stays truthful even after the threshold rule that raised it is edited or deleted.

    Active alerts and history render the measurement through the same formatter, so the two screens cannot disagree.

Operator actions — acknowledging, muting, opening a maintenance window — are recorded in the audit log like every other state-changing action, so who silenced what, and when, is always answerable.

The same actions are available to AI and automation clients through Yagra’s MCP tool surface, under the same permission checks and the same audit trail.

  • Monitoring — the checks and thresholds that feed this pipeline.
  • Passive events — syslog, SNMP traps, and webhooks, and the event rules that raise alerts.
  • Users and SSO — the roles behind the operator actions on this page.