Monitoring
Yagra actively monitors network devices and servers over ICMP, SNMP, and HTTP(S), watching liveness, performance, and thresholds across tens of thousands of nodes.
This page describes three things: what gets collected, how a node resolves to a set of checks, and how the polling engine behaves. For what happens once a check goes bad, see Alerting.
What gets monitored
Section titled “What gets monitored”Everything Yagra watches is a node, but not every node is a router.
A node’s kind follows from how it is configured. It is derived from the monitoring attached to the node, not a label you maintain separately:
| Kind | What it is | Checks it gets |
|---|---|---|
| Device | A network device or server with an IP address | ICMP liveness always, plus SNMP collection when a credential and a collection set resolve |
| URL | An HTTP(S) endpoint monitor | One HTTP probe — no ICMP (the target may be unpingable behind a CDN) and no SNMP |
| DNS | A name-resolution monitor | One DNS probe — a name has no address of its own |
| Meraki | A device imported from a Cisco Meraki organization | Collected through the Meraki Dashboard API at the organization level, so it emits no per-node poll job |
| Wireless AP | An access point imported from a wireless controller | None of its own — its controller’s poll answers for it (see Wireless controllers) |
When a node matches more than one shape, the more specific kind wins, in the fixed order Wireless AP → Meraki → URL → DNS → Device. All five kinds live in the same inventory tree, share the same dashboards, and feed the same alerting pipeline.
The kind is visible where you work. A short AP / URL / DNS / Meraki badge sits after the name in
the inventory tree and beside the title on the node’s page. Ordinary devices are unmarked, so a
badge means “read this one differently”.
The node’s page then shows only the tabs it can fill. A URL or DNS monitor has no Interfaces, Neighbors, or Flow tab, because it is never SNMP-walked and is never a flow exporter.
Kind is not the only axis. A Device with no SNMP configured — no credential bound to it, and no
deployment-wide YAGRA_SNMP_COMMUNITY — has no Interfaces or Neighbors tab either. Both are built
from SNMP walks, and on such a node no walk ever runs. Events and Flow stay, because syslog, SNMP
traps and NetFlow are attributed by the device’s own address: a host watched only by ping can
genuinely have rows in either.
Its Overview lists only facts that kind can hold, so a DNS monitor shows its resolver rather than four blank rows where a maker and model would be.
A device polled over SNMP also shows its OS version there. The poller re-reads it once an hour,
from a vendor MIB, ENTITY-MIB or sysDescr, whichever that vendor uses. A device the table does not
cover shows —.
The serial number sits beside it. It is read on the same hourly pass, from ENTITY-MIB’s chassis
rows. A stack lists every member, in order. A Huawei stack keeps one chassis row for the whole
stack, so its members are read from their main boards instead. A Juniper device is read from
Juniper’s own MIB first. A
Meraki device shows the serial it was imported with. A device that keeps a serial in neither place
shows —.
Both values normally wait for that hourly read. Press Poll now on the node to read them on that poll instead. The page shows them about 30 seconds later.
A device whose maker Yagra does not know yet is read in full on its first poll, so its version and
serial appear at once. After that it is read in full once an hour, like any other device. Between
those reads, each poll asks it only for sysDescr and sysObjectID. That is enough to name the
maker. The model is taken only from a full read.
The API reports the same value: kind on GET /api/v1/nodes and on the MCP node tools.
Filing nodes into folders
Section titled “Filing nodes into folders”Nodes live in a folder tree, and they can be moved a batch at a time. In the inventory tree, Ctrl-click (⌘ on macOS) picks nodes one at a time and Shift-click picks a run; a bar above the tree then says how many are selected and offers Move…. The same items sit on a node’s right-click menu, so the feature is reachable without knowing the gesture. Dragging any node that is in the selection carries the whole selection; dragging one that is not moves that node alone. Dropping between two rows places the batch there, in the order you selected it. Dropping onto a folder puts it at the end.
The tree also works from the keyboard, as a single Tab stop. ↑ / ↓ move the selection, and the detail pane follows once the key comes to rest. → opens a closed folder or steps into an open one. ← closes an open folder, or steps out to the parent folder. Enter opens or closes a folder, and opens a node’s own page. Home, End, Page Up and Page Down jump through a long tree. Space adds the current node to the selection or takes it out, and Shift+↑ / ↓ extends the selection the way a Shift-click does. The context-menu key or Shift+F10 opens the row’s right-click menu, and the arrow keys walk its items.
Deleting works on the selection the same way: Delete… on the bar, or Delete N selected… on the right-click menu. The confirmation names the nodes. A node in a folder you cannot see is not deleted, and the dialog says how many of how many went.
The same selection drives four more actions. Poller pool, Maintenance, Mute and Poll now act on every selected node when the row you right-clicked is part of the selection, and their headings say how many. The selection bar carries them under More…, along with Tag…. Each one reports how many of how many nodes it reached: a node deleted since the page loaded, or one in a folder you cannot see, is not counted. A maintenance window or a mute is created per node, so each can be ended on its own. A pool chip is shown as already chosen only when every selected node agrees. Two items have no batch form — Edit node… shows one node’s whole binding, and Pin is your own account’s navigation — so while a selection is in play they name the node they act on.
Nodes ▸ Duplicates finds a device that was added more than once. It finds copies at the same address and at two different ones, such as a router added once by its loopback and again by a LAN address. Nodes that share an address, a serial number, an ARP-cache MAC address or an LLDP chassis ID form a likely group. A group joined only by weaker signs, or one whose serial numbers or models disagree, is marked to check. Each group suggests one node to keep, but nothing is selected for you. A delete that would remove every node of a group is refused. The screen needs the manage-configuration permission.
Nodes ▸ Subnet overlaps lists address ranges that devices at more than one site use. A site is the nearest folder of type Site above a device, or the device’s own folder when there is none. The screen compares the interface addresses devices already report, so nothing extra is polled. It shows three kinds:
- Same address. Two sites configure the same IP address. That is almost certainly a reuse.
- Nested. One site’s range sits inside another site’s range.
- Same range. Two sites use the same range, and every address is different. That is either a reuse or a line the sites share, so Yagra asks you.
A link between two sites (a /30 or /31 with two devices on it) is never listed as open. A range that repeats on purpose leaves the list in one of two ways. An exclusion rule names a range, whole words in the port’s name or description, or both; the carrier CGNAT range 100.64.0.0/10 is built in and can be switched off. Mark as intentional takes one overlap off the list until another site starts using it. A hint such as Looks like WAN only suggests a rule. It never applies one. Rules and marks need the manage-configuration permission, and an account limited to some folders cannot change them.
The selection is separate from the row whose detail is open, so you can read one node while collecting a batch. It is deliberately not kept in the URL — a reload never restores a selection of rows that are no longer on screen.
Every filter on the tree sits behind one funnel button beside the search box. It opens a panel with three switches at the top — Pinned only, Needs attention and Hide empty folders — and the State, Kind and Pool filters below them as check boxes. The funnel shows how many filters are in force. Each one is also listed as a chip under the inventory header, with a ✕ to remove it, beside Clear all filters.
Pin a node or a folder from its right-click menu or its detail pane. Pinned only then narrows the tree to what you pinned: the pinned nodes, the pinned folders with everything inside them, and the folders above both. Pins belong to your account, so they follow you to another browser.
Hide empty folders hides every folder that has no node in it or below it. It is kept with your account, like Pinned only. Needs attention narrows the tree to the nodes that are warning, critical or unreachable. It works by ticking those three states in the State filter, so you can see what it did and undo part of it by hand, and it travels in the URL like any other filter.
The folders you close in the tree belong to your account too, so another browser opens the tree the same way. While a search term or a state, kind or pool filter is on, every folder starts open, so a folder you closed earlier cannot hide a match. A folder you close there stays closed after the filter is cleared.
Every folder picker — move, add node, a group’s parent, mute, maintenance window, the discovery site — narrows as you type once there are at least eight folders, matching the whole path so a site’s name keeps the racks under it.
A folder can carry the IP ranges in use at it. Type them into the IP ranges section of the
folder dialog, or let a NetBox sync attach them. Host bits are allowed, so 192.168.1.5/24 is
stored as 192.168.1.0/24. IPv4 and IPv6 are both accepted. Ranges a sync maintains are listed
there too, marked From sync, and cannot be edited or removed by hand.
A folder also lists the subnets its devices use that its ranges do not cover. Open a folder: below its ranges, Subnets missing from the IP prefixes compares every address that the devices in the folder and its subfolders report with those folders’ ranges. Each subnet it lists says why:
- It is in no range at all.
- A range covers only part of it.
- It is in another folder’s range. Either the range is attached to the wrong folder, or two sites reuse one private range.
- Only a parent folder’s range covers it.
It also says how many of the folder’s devices it read addresses from. A device with no SNMP address walk adds nothing to the comparison. A folder of more than 2,000 devices is not compared; open a folder further down, or use the screen below. An account limited to some folders is told that another folder holds a subnet, but not which folder or which range.
Nodes ▸ Missing IP prefixes makes the same comparison for every site at once. A site is the nearest folder of type Site above a device. The screen opens by site. Tabs separate sites with gaps, sites with none, and sites not compared yet, where no device has reported an address. Opening a row lists the subnets with the devices, ports and addresses they were seen on. Subnets lists every gap on its own line, and Export CSV saves that list. At most 2,000 gaps are listed, shared between the sites, with the total beside them.
A gap you have checked and mean to keep can be marked. Press Mark as intentional at the end of its row, and add a note if you like. The gap moves to the Intentional tab. Move back to To check undoes it. A site whose every gap is marked is listed under Intentional, not as complete. When the reason for a subnet changes, for example it becomes partly registered, the mark stops applying and the gap is back to check.
Where a folder carries a range, the right-click menu and the selection bar also offer to file nodes by address. It proposes and never acts: a dialog shows what would move and where, what falls inside no range, and what two folders claim equally well, and nothing is written until you press the button. A node that two folders claim is never moved automatically.
Discovery’s import uses the same ranges, and where a folder carries one it is on by default. Each device goes into the folder whose range contains its address, rather than all of them landing in one folder. The candidate list shows a Folder column before you press Import, and that column is a picker — a device that belongs somewhere else is one click to redirect, including to the tree root. A device no range covers, or one two folders claim equally well, goes into the folder chosen for the sweep. The import never fails because of an address.
Naming a node, and labelling it
Section titled “Naming a node, and labelling it”A node’s name is yours to set. Nothing in the product overwrites it — not a poll, not a discovery sweep, not the classifier — so a typo is fixed in Edit node rather than by deleting the node and creating it again, which would leave its metric and alert history behind.
Notes is free text on the same dialog, for whoever touches the device next: “in the ceiling
void above the east corridor, needs a ladder”. It sits at the top of the node’s Overview tab, and
an AI client reading over /mcp sees it through get_node_status.
Tags are single labels — JAPAN, core, 松山本社 — not key=value pairs. A node
carries as many as it needs, unlike its one inventory folder. Edit them per node in Edit node,
or apply them to a batch from the inventory tree’s right-click menu (Tag N selected…), which
adds to what each node already carries rather than replacing it.
A tag on a folder reaches everything beneath it. Put JAPAN on the Japan site and every
device under it carries it, including ones discovered next month — which a one-time bulk edit
cannot do. It is resolved on every read rather than copied onto each node, so moving a folder or
editing a parent takes effect at once instead of leaving stale copies behind. A device that should
be the exception refuses an inherited tag from its own Edit node dialog, and refusing one on a
folder takes it away from that folder’s whole subtree. Both the node’s Overview tab and the
folder’s detail pane show which tags are the thing’s own and which came from above.
Every device node gets ICMP liveness monitoring. It is always on and requires no credential.
The probe records reachability and round-trip time (icmp_rtt_ms), which is graphable on the node
detail and usable in threshold rules — for example, RTT above 100 ms.
A device that stops answering transitions to the unreachable state and raises a liveness alert, subject to the hysteresis and dependency-suppression logic described under Alerting.
ICMP uses raw sockets, so the poller container needs the NET_RAW capability. The bundled Compose
files grant it to the poller, and only the poller, already.
SNMP is where the depth of device monitoring comes from. Yagra speaks:
- SNMP v2c — community-based.
- SNMP v3 (USM) — authentication and privacy.
Both versions collect scalar metrics (CPU, memory, sessions, and so on) and walk tables, including multi-index tables.
Per-interface metrics come from interface-table walks using GETBULK, on v2c and v3 nodes alike, so per-interface counters are collected even on devices that only permit v3.
A node’s table walks are folded into a single SNMP session per poll, which keeps the cost of a many-port switch to one conversation rather than one per table.
On the node detail, per-interface throughput and status appear as sparklines and time-series charts with a shared range control (1 hour to 7 days, plus custom). Configured-bandwidth reference lines are drawn on the interface charts.
What a port says about itself
Section titled “What a port says about itself”The Interfaces tab lists each port with its speed, its duplex and its media type, all three filterable from the column filter row. That is what makes a mismatch findable: “show me the 100 Mbps ports” is how a gigabit link that negotiated down gets noticed.
Traffic is two columns rather than one, In and Out, each shaded by how full the link is. The shade is the reading against that port’s own advertised rate, so 900 Mbps is red on a gigabit port and green on a ten-gigabit one. A port that advertises no rate is left unshaded, because there is nothing to divide by, and so is a port that is down. Hovering a cell names the figure and the share of the link it represents.
The IP addresses column lists the addresses configured on each port, written 192.168.0.1/24.
Every address the device reports is there, secondaries included. The first is shown, and +N opens
the rest. The column’s filter searches every address of the port: 10.221. finds a port carrying
that range even as a secondary, and /30 finds the point-to-point links. The addresses come from the
hourly address walk the network map already uses, so a change shows within the hour. An address
whose mask the device did not report in a readable form is shown without a prefix. On a phone the
column is not shown, and opening the port lists every address instead.
The Neighbors column names the device each port sees over CDP or LLDP, with +1 when there are
more. Clicking the name opens every neighbor on that port, with its port, capabilities and
management address, without leaving the list.
- A CDP neighbor is placed by the port’s index.
- An LLDP neighbor is placed only when the device names the port the same way the list does. A device that reports the port differently (a Junos switch reporting it by number, for example) shows its neighbors in the Neighbors tab only.
The column is not shown on a phone.
Every column in this list — and in every other list in Yagra — can be resized by dragging the grip on the right edge of its heading, or by focusing that grip and pressing the arrow keys. The widths are stored against your account, so they follow you to the next machine you sign in from.
Speed and duplex ride the interface walk the poller already performs, so they cost no extra SNMP conversation. Media type is read separately, once an hour, since a medium only changes when someone swaps a module. The standard MAU table is asked first; for the ports it does not answer, Cisco’s own port table is read next, and the transceiver’s part number last. Huawei’s YunShan OS implements none of the three, so its ports report the medium on the interface walk instead, and a copper port’s IEEE designation follows from the speed it negotiated.
A blank cell means the device did not answer, not that something is wrong. Duplex is a copper diagnostic — IEEE 802.3 defines no half duplex above 1 Gbit/s, so a fibre port has nothing to negotiate and reports “unknown” — and many devices implement no MAU table at all.
Charts on one port
Section titled “Charts on one port”Selecting a port opens a chart dock below the list, with a shared cursor: hovering one chart moves the crosshair and the legend on all of them, so “traffic spiked — did discards spike with it?” is a single reading.
- Throughput, switchable between bits/sec and packets/sec. A device’s forwarding ceiling is often a packet rate rather than a bit rate, so a link with bandwidth to spare can still be saturated.
- Errors and discards on one axis. The two mean different faults: an error is a frame that arrived damaged (cabling, optics, NIC); a discard is a frame the device dropped although nothing was wrong with it (congestion, queue overflow, ACL).
- Optical power (Rx / Tx, dBm), for any port with a transceiver in it, shaded with the module’s own published acceptable window where the vendor reports one. Fibre degrades gradually — a receive level drifting from -7 dBm toward -18 dBm is a link on its way out. Five vendor dialects are read (ENTITY-SENSOR-MIB, Cisco’s own CISCO-ENTITY-SENSOR-MIB, Huawei, Juniper and H3C) and every reading is normalised to dBm, so the number means the same thing on every vendor. Cisco does not implement the vendor-neutral table at all, which is why its own is read beside it. Nothing alerts on the window: those are the module’s figures, not a threshold anyone configured.
Raw counters, query-time rates
Section titled “Raw counters, query-time rates”Yagra stores interface and error counters raw. It never computes a rate on the poller by subtracting the previous sample.
Rates and utilization are derived at query and evaluation time by the time-series store, which detects counter wrap and counter resets automatically. This has two practical consequences:
- Pollers are stateless. No previous-value bookkeeping, and no wrong rate after a poller restart or failover.
- A counter wrap (32-bit) or a device reboot never produces a phantom traffic spike, because the wrap/reset handling lives in the store’s rate functions, not in a hand-rolled subtraction.
A rate needs at least two samples, so the window it is taken over follows the poll interval. The window always holds at least two of the node’s polls. That applies to the Interfaces tab, a node’s rate charts, the utilization heatmap, Top interfaces, fleet throughput, traffic spikes and drops, and interface utilization alerts.
On the API, the window of GET /api/v1/metrics/interface-delta and the window_secs of the MCP
tool top_interfaces are widened to at least twice the slowest poll interval in the fleet. A
shorter window has nothing to compare.
It also means threshold rules apply to gauges, not raw counters — see Thresholds below.
Credentials
Section titled “Credentials”SNMP communities and v3 credentials are stored encrypted at rest. They are never logged, never returned by the API, and never used as metric labels.
A node without a bound credential can fall back to a deployment-wide v2c community via
YAGRA_SNMP_COMMUNITY — see the configuration reference.
URL monitoring
Section titled “URL monitoring”A URL node monitors an HTTP(S) endpoint the way you’d check it by hand, on a schedule:
-
Liveness and status. The probe succeeds when the endpoint answers with the expected HTTP status. The expectation is configurable: any 2xx, an explicit list of codes, or a range.
-
Response time.
http_response_time_msrecords how long the endpoint took to answer, so “it is up” and “it is slow” are separate questions.It measures time to the response headers, and records nothing at all when the endpoint did not answer, so a timeout does not appear as a flat slow-response line for the whole outage.
No default threshold is seeded. Response latency varies too much between environments for one to be right.
-
A keyword in the response body. A monitor can require that the body contains a keyword, or that it does not. This is the case availability monitoring structurally cannot see: an endpoint answering
200while its body says the service is broken.Matching is plain, case-sensitive substring matching, not a regular expression. A body larger than the read budget reports “not satisfied” rather than risking a false healthy.
-
Numbers out of a JSON body. Up to 8 values per monitor, each recorded under a metric name you choose: a queue depth, a replication lag, a worker count. The path is dot-separated and names exactly one value (
data.queue.depth).A value that is not a number records nothing for that poll rather than a zero, because a zero is indistinguishable from the value genuinely being zero.
-
TLS certificate expiry. Certificate expiry is tracked for HTTPS targets, so a certificate runs down on a chart instead of expiring as a surprise.
-
TLS verification is on by default. Disabling it is an explicit, per-monitor opt-in.
The probe is hardened against SSRF. It resolves and connects only through a filtered resolver, on the initial target and on every redirect hop, refusing loopback, link-local, and cloud-metadata destinations, including DNS-rebinding tricks.
Private and internal ranges stay allowed — an NMS legitimately monitors them.
A URL monitor’s configuration is fully editable after creation. The node’s URL health card has an overflow (⋮) menu with Edit and Remove monitoring, and the editor covers every field, including the expected status.
Removing the monitoring leaves the node in the inventory with its recorded history intact. It simply stops probing. Changing a monitor requires the same permission as any other monitoring change (Admin).
DNS monitoring
Section titled “DNS monitoring”A DNS node monitors a name the way a URL node monitors an endpoint. Bind a node to the built-in DNS name resolution profile and Yagra records:
- whether the name resolves, and how long resolution takes;
- the dig-like recursive CNAME chain the name resolves through.
The chain history appends only when the chain actually changes. TTL countdown and round-robin reordering don’t count as a change, so the history reads as a change log, not noise.
Numeric summaries are graphable and alertable: dns_up, dns_resolve_ms, dns_chain_length, and
dns_answer_count. A default dns_up threshold is seeded, so a new DNS monitor alerts out of the
box.
When resolution fails, the reason is shown as a readable sentence — for example, “No such name (NXDOMAIN)” — in both English and Japanese, rather than an internal error token.
Like URL monitors, a DNS monitor is editable and removable after creation from the node’s DNS health card: resolver, record type, and the rest. Removal preserves the node and its history.
Cisco Meraki
Section titled “Cisco Meraki”Yagra monitors Cisco Meraki estates over the read-only Dashboard API.
Add an organization under Settings ▸ Integrations ▸ Cisco Meraki, with a read-only API key or a key already held under Settings ▸ Credentials. Yagra then reads what the organization holds — its networks and its devices — every five minutes. Each organization has a page of its own, showing that inventory, where each device is filed, and how the last sync ended.
- Devices import themselves. A device becomes a node once it sits in a watched network and Meraki has reported it online at least once. A spare still in its box is left alone until it first comes online, so it does not arrive as an outage. A device you delete stays deleted, and you put it back by hand from the organization’s page. Organizations added from now on start with automatic import on. Organizations that already existed start with it off, because every one of their devices was chosen by hand.
- They are filed by IP range. A device whose address falls inside exactly one folder’s IP range goes into that folder, and the longest range wins — so a Meraki estate lands in the same site tree as everything else. A device with no address, or one that two folders claim equally, goes under an Organization → Network folder instead. Turn File by IP range off for an organization whose sites reuse one private range. Folders the integration owns carry a Meraki badge. With File by IP range on, an access point or switch that has no address yet is not imported. It waits and keeps its place under the device cap. Once Meraki reports an address, it is imported into the folder whose range holds it. A mesh repeater never reports an address, so it is never imported automatically. Import it from the organization’s page; it then goes under its network’s folder.
- An MX is addressed by its LAN, never its WAN. An MX reports no LAN address to the Dashboard, and its WAN address is usually in nobody’s IP range. So Yagra reads each MX network’s VLANs, one network at a time, and gives the MX one of its own VLAN addresses. An address another network of the organization also uses is skipped first. Of the rest, the lowest-numbered VLAN inside a folder’s IP range wins, else the lowest-numbered VLAN. An organization’s first sync reads the VLANs of every MX network in one go — about three minutes for 350 networks at the default rate — and imports its devices only after that. A sync whose reads were cut short imports no MX at all, so an MX never takes an address that a site not yet read turns out to share. An MX that is already a node takes its new address but stays in its folder; select it under Nodes and use Move by IP range. The two MX of a warm-spare pair share one address.
- Collection happens per organization. One paged, org-wide API call covers many devices, so a large estate never trips the organization’s API rate limit. The call asks for the whole organization and keeps the watched networks’ rows, so the request does not grow with the number of networks. A monitored device that ends up outside the watched networks is pointed out on the organization’s page rather than left behind quietly.
- Slow reads do not hold up availability. Each organization’s collects run in two lanes. The fast lane carries availability, uplinks, traffic and the access points’ clients and utilization. The slow lane carries the switch ports, the SSID read and the periodic inventory sync. The request rate you set is the organization’s total, and each lane uses half of it.
- Availability decides whether a device is up. A device the Dashboard reports offline is down and alerts like any other node. The other collections record readings on a Cisco Meraki card on the node detail, and never speak for liveness.
- An MX shows its WAN uplinks, its Auto VPN and its warm-spare pair. Each WAN uplink has its state, loss, latency and traffic, charted per uplink. Two default rules come with the built-in MX profile: a warning when an uplink fails, and an Auto VPN alert that warns when a site loses one of its hubs and turns critical when it reaches none. The card also shows the pair’s primary and spare, and whether the site is running on its spare.
- A switch shows its ports. A Meraki MS gets an Interfaces tab like an SNMP switch: each port’s state, speed, duplex and traffic. Its ports rank beside SNMP ports and carry the same utilization alerts. Traffic is a five-minute average from the Dashboard, so it is drawn 12 to 17 minutes behind the port. Error and discard counts are not collected.
- An access point shows its clients and radios. A Meraki MR reports its connected clients, the SSIDs it broadcasts, and each radio’s channel utilization with its non-Wi-Fi share, under the same names a wireless controller’s access points use. The channel and transmit power are read every twenty minutes. The Dashboard does not say whether a radio is on, so a radio has no up/down state.
- One alert says when the API stops answering. After three availability collects in a row go unanswered — a revoked key, a Dashboard outage, rate limiting — one critical alert is raised about the organization, not one per device. It closes when a collect is answered again, never merely because reports stopped arriving. The devices keep the last state that was collected, and each one’s Overview says why it is not current.
- Per-organization controls let you pause and resume collection, tune per-tier polling cadence and the request-rate budget, and edit which networks are in scope. The switch-port and wireless tiers run every 300 to 600 seconds and are sent only to a poller pool whose pollers all support them, so upgrade the Meraki pool’s pollers to start collecting them. Availability cannot be switched off, because it is the only tier that says whether a device is up. A global kill switch halts all Meraki collection instantly. A new organization may import up to 10,000 devices and collects its MX uplinks’ traffic every five minutes; both can be changed on its page.
- Sync now re-reads the whole organization, in the background. It reads the inventory, the VLANs of every MX network and the warm-spare roles, then imports what the organization imports. It runs in the slow lane only, so availability keeps being collected; the switch-port and SSID reads wait until it ends (about six minutes for 350 networks at the default rate). The organization’s page shows how far the read has got and refreshes itself while it runs. Organizations sync side by side, so one organization’s long read holds up no other.
The integration is read-only by design. It issues HTTP GET only, every request is restricted to allow-listed Meraki API hosts, and the API key is encrypted at rest and never returned or logged. Inside Yagra, every Meraki write refuses an account whose visibility is limited to groups: an organization is monitored as a whole, and a device nobody has imported yet belongs to no group.
In a distributed deployment, Meraki cloud-poll jobs are routed to the poller pool named by
YAGRA_MERAKI_POOL (default default). That is useful when only some sites have internet egress.
See Distributed polling.
Wireless controllers
Section titled “Wireless controllers”Two families of wireless controller are monitored, each as an ordinary SNMP device on its own built-in profile:
- A Huawei wireless controller (AC), on the Huawei wireless controller profile.
- A Cisco wireless LAN controller, on the Cisco wireless controller profile. AireOS controllers and the Catalyst 9800 answer the same tables (AIRESPACE-WIRELESS-MIB), so one set of templates reads both. A controller node that already carries this profile starts on its first poll after the upgrade.
From that one node, Yagra reads the wireless estate behind it:
- Controller totals. Access points joined and clients online. A Huawei AC also reports the APs configured and licensed, the share working normally, clients per band, and successful roams. A Cisco controller has no object for these that AireOS and the 9800 both answer, so Yagra counts its joined APs and its clients from the controller’s own tables.
- SSIDs. One row per SSID, named by the SSID itself, with its clients. On a Huawei AC the row also carries the clients per band, how many access points broadcast the SSID, and its traffic counters. Each SSID has a history you can chart and can carry a threshold rule.
- Access points. Every AP the controller reports, with its state, address, model, serial
number, software version, CPU, memory and clients. They are listed on the controller’s APs
tab, and by
GET /api/v1/wireless/apsand the MCP toollist_wireless_aps, whether or not they are monitored as nodes.
An access point becomes a node of its own, of kind Wireless AP with an AP badge, in one of two
ways:
- Leave automatic import on in the controller’s APs tab. It is on by default, for a controller Yagra starts reading and for one already registered. Within a minute, every AP that has ever been in service becomes a node, up to the controller’s AP cap (1024 unless you set another). Switch it off to keep a controller’s APs out: the AP nodes already created stay, and are yours to delete — a deleted AP node is not created again on its own.
- Press Monitor on one AP in the same tab.
Both need the Manage configuration privilege. An AP node is filed in the controller’s own folder, unless the import settings name another.
An AP node is never polled itself. Its controller’s poll answers for it: whether it is in service, its clients, CPU, memory and temperatures, and its radios. Each radio is a row on the AP’s Interfaces tab (slot 1 is 2.4 GHz, 2 is 5 GHz, 3 is 6 GHz), so a threshold rule can name one radio of one access point. A Cisco controller reports each radio’s clients, channel and channel utilization. A Huawei AC also reports interference, the noise floor, the average client signal and the transmit power.
Behind an HA controller pair, the controller serving an AP speaks for it. The standby’s view is kept in the AP list, but it never replaces the active controller’s values and cannot take an AP down. An AP keeps its node and its history when the pair switches over.
When a controller stops answering, no AP result is produced at all. One controller outage
therefore does not raise one alert per access point. The controller’s wlan_ap_walk_complete
warns instead. An AP that no controller has reported for ten minutes (or for three of the
controller’s poll intervals, if that is longer) reads unknown, and its overview says since when.
An AP that was down when it was last reported stays down, and an alert still open on it still shows.
A Cisco controller has no “down” state for an access point: it removes an AP it has lost from its
table. So on a Cisco controller, and only there, an imported AP that this controller last served
and that is missing from a complete read of the table is recorded as down, and raises its liveness
alert. It comes back up when the controller lists it again. Absence does not count in the first 15
minutes after the controller starts, while its APs rejoin. The controller’s
wlan_controller_aps_missing counts the APs it serves that are not joined, and a built-in rule on
the Cisco profile warns at 1 or more, three polls in a row.
Known limits in this release:
- A controller’s AP list is cut at 1,024 access points.
- It has been measured on controllers with up to 38 APs. By calculation, a controller with more than roughly 800–1,200 APs may not finish its AP walk in time, and its list then stops refreshing.
- The Catalyst 9800 has been checked against a recording of one, not against a running controller.
- An AP that leaves a Cisco controller for a controller Yagra does not monitor reads as down.
NetBox
Section titled “NetBox”Yagra can build its folder tree from NetBox, the network’s source of truth, so nobody maintains the same site list twice.
Register a NetBox deployment under Settings ▸ Integrations ▸ NetBox with its base URL and an API token. Every sync then reads NetBox’s Regions and Sites and mirrors them into the node folder tree, with each Site’s latitude and longitude filled in, so the Geo map’s pins appear without anyone placing them. A sync runs once an hour by default. Sync now on the server’s row asks for one at once: the row says “Sync requested”, then “Syncing…”, and the sync finishes even if you leave the page.
- The sync is read-only, in one direction. Yagra never writes to NetBox. It issues HTTP GET only, follows no redirects, and the API token is encrypted at rest and never returned or logged.
- Ownership is split per column. A sync owns each folder’s name, parent and map position; the poller pool an operator sets on a folder is never touched. A folder’s place in the inventory tree is set only when a sync first creates it. After that, it stays wherever you drag it.
- Only Active sites become folders. A Site whose Status is planned, staging, decommissioning or retired is left out, and so are its prefixes. A folder made for a Site that later stops being Active stays, with its nodes, and is counted as no longer synced from NetBox.
- A Site removed in NetBox is flagged, never deleted. Deleting a folder deletes every folder and node beneath it, so one mistaken click in an external system must not be able to do that.
- Site ID field. Choose which NetBox field holds your own site code —
slug,facility,description, or a custom field — and each Site’s folder is namedJPMYJ01 Matsuyama Homeinstead ofMatsuyama Home. Off by default. The sync result says how many Sites had that field empty, so choosing the wrong field is visible rather than silent. - Subnets. A sync also reads NetBox’s IP prefixes and attaches each one to the folder of the Site or Region it is scoped to. Right-click a folder that has some and choose Run discovery here…, or pick the site on the Discovery screen: its ranges appear as a checkbox each, with the address count per row and a running total, and devices imported from that sweep are filed into the folder. IPv6 prefixes are stored and shown but left out of a sweep. A token that may not read prefixes changes nothing — the folder tree still syncs.
The base URL is operator-supplied, so it is checked before the token is sent anywhere: loopback, link-local, multicast and unspecified addresses are refused, and private addresses are allowed because that is where NetBox usually lives. A NetBox behind a private CA is supported by pasting that CA’s certificate on the same form; only that connection trusts it, and certificate verification is never disabled. Verified against NetBox 4.6; 3.x is untested. Placing nodes into those folders from NetBox, and creating nodes from NetBox, are not part of this integration yet.
GET /api/v1/netbox/servers is also readable over MCP as get_config(kind=netbox_servers). The
API token appears in neither.
Profiles and collection sets
Section titled “Profiles and collection sets”What SNMP data a device yields is decided declaratively, not per node:
- Device profiles form a role × network-OS taxonomy (a core switch is not a firewall is not a Linux host).
- Each profile carries editable collection templates, which draw on a curated, searchable MIB/OID catalog.
- A node bound to a profile resolves — through its templates and its credential — to a concrete collection set: the scalars and tables actually polled on that device.
Built-in profiles and templates cover common vendors out of the box: Cisco, Huawei (including USG firewall sessions and memory), Meraki MX/MS, and A10, plus the standard host and interface MIBs.
All of it is editable and searchable, so extending coverage to a new device family means adding catalog entries and a template, not writing code.
Everything a node collects is visible, not just what the UI anticipated. A node’s Collection tab lists every metric that has arrived for it and charts any of them.
That matters most for the coverage you added yourself, where a vendor table column collects successfully but no screen was written to know its name.
Each entry states which of three states it is in: configured and flowing, configured with nothing arriving yet, or arriving with no collection item behind it.
That last case is normal rather than a fault. Reachability, the URL and DNS monitors, the neighbour count, and values extracted from a monitored JSON body all come from checks rather than from a collection set.
A counter can be charted as a per-second rate. Its stored value is an odometer reading, and charting that raw draws a rising line that looks like traffic and is not.
Discovery and classification
Section titled “Discovery and classification”You don’t have to add nodes one at a time:
- Discovery sweeps IP ranges and address lists, using credentials chosen from the stored credential picker.
- A Credential Finder tries stored credentials by reference against a discovered device to find one that answers. It is rate-limited per device, and the attempted values are never logged.
- Classification rules match on
sysObjectID/sysDescrand apply the matching device profile automatically, so a discovered device lands with the right collection set attached. - Results land in an import grid, so you review and choose what enters the inventory rather than having a sweep add nodes behind your back.
- A candidate whose address is already a device node is shown greyed out with an In tree badge and cannot be ticked, so importing the same sweep twice does not create duplicates.
- A candidate that looks like a device node monitored at another address is muted too. It is
badged Likely the same device when that node’s interface list carries the scanned address
and both report the same model (
sysObjectID). It is badged Maybe the same device when the interface list carries the address but one side reports no model, or when the name and model match. Two devices that report different models are never marked. The row links to the node and says why. It stays selectable: sites that reuse one private address plan can make this guess wrong.
Classification rules are data, not code. They live next to profiles and are editable in the same place.
A corrected rule does not move nodes that are already in the tree. Nodes ▸ Reclassify lists
every device node whose profile differs from the one the rules now choose, with the rule and the
sysObjectID and sysDescr it ran on. Apply to selected moves the nodes you tick. Keep
current profile locks a node, so it is not listed again.
To re-read one node — after changing its SNMP community, say — right-click it and choose Rediscover…. Yagra asks the device again with the node’s own credential, from the node’s own poller pool, and shows the node’s profile, maker and model beside what the device now says. Nothing is written until you tick the rows you want and press Apply, and a locked profile is left alone. While no poller has picked the request up, or the device answers ping but not SNMP, the dialog says so rather than showing “no change”.
Running a sweep
Section titled “Running a sweep”A sweep is a job you can steer, not a page you have to sit on:
- Leave the page and come back. Nodes ▸ Discovery lists the sweeps the core is holding and reattaches to the one you pick. The scan id is in the URL, so a reload or a shared link lands on the same sweep.
- Choose which site it runs from. Without a poller pool the sweep goes to whichever poller answers first — which on a multi-site deployment can mean a remote poller sweeping head office, reaching nothing, and reporting a successful empty scan.
- Stop it mid-run. The poller stops probing, so the ICMP and SNMP traffic actually ceases. Devices already found stay on screen and can still be imported. A probe in flight is left to time out, so the stop takes a few seconds, and the screen says Stopping… until the poller confirms rather than claiming a stop it cannot yet see.
- A sweep nobody has picked up says so. It starts as Waiting for a poller and becomes Running only once a poller reports — so a sweep sent to a pool with nothing alive in it is distinguishable from one that is running and finding nothing.
By default a sweep skips an address that does not answer ping, which is what makes it fast: on a test network a /24 carrying eight devices took 5m21s when every address was asked for its identity with every candidate credential, and almost all of that was spent on the 246 addresses with nothing at them.
The trade is explicit, because discovery sends a single echo request: a device that filters ICMP
but answers SNMP will be missed, and one lost packet is indistinguishable from an empty address.
Turn the checkbox on the scan form back on — or send "snmp_when_unreachable": true to
POST /api/v1/discovery/scan — to probe every address regardless. Liveness monitoring is unaffected;
it waits for three consecutive failures before believing a node is down.
Finished sweeps are kept for six hours (at most twenty), then discarded.
The network map
Section titled “The network map”The Network map draws the connectivity Yagra derived from what the devices themselves report. Nobody types a link. Each edge carries the evidence behind it, and the map labels and legends them:
- CDP / LLDP adjacency, matched to a monitored node through the peer’s management address. A peer that advertises no usable address (a Meraki MX, for one) is matched through its MAC instead, when a Meraki organization lists a device under that MAC and that device is a node here.
- A shared IP subnet — two nodes with an interface address in the same prefix are adjacent as a matter of fact.
- OSPF neighbours, BGP peers, and connected routes, which is how the map sees links that
share no subnet: a point-to-point
/32(a PPPoE dialer, a tunnel endpoint), an unnumbered OSPF link, or a peering across a segment whose addressing has not been collected.
A link seen more than one way is still one link, carrying all of its evidence. Redundant paths are kept rather than collapsed, so a server reached through two routers shows both links.
The map draws one folder at a time. On a level:
- The folder’s own linked nodes are drawn one by one.
- Each subfolder is one box, showing how many nodes it holds and how many are down. Select it to go into it; the breadcrumb takes you back up.
- Links between the same two things are bundled into one line with a count. Select the line to list the ports on both ends.
- A link that leaves the folder ends in a dashed box. It opens the level where both ends are shown.
A Site folder, and every folder beneath one, is drawn flat instead. Every linked device of the site is drawn, with the folders down to the one it is filed in under its name. A line to another floor of the same site ends in a dashed box for the device at the far end.
Devices are stacked in rows by what they do in the network:
- Routers and firewalls.
- Switches that route.
- Switches that do not route.
- Everything else.
The role comes from what Yagra already collects: a Meraki product type, a router or firewall device profile, a default route that leaves the site, an OSPF or BGP adjacency, or addresses in two or more subnets. Select a node to see its role and the reason for it. A level with no router or switch is laid out by its links alone.
A device is a site’s way out when it routes and its IPv4 default route points outside the site. That means the next hop is an address no device of the site holds, and it is not on a subnet another device of the site uses, or the route points out of an interface with no gateway. Such a device goes in the top row even when its profile does not call it a router. The second condition keeps a core switch that points at its routers’ HSRP or VRRP address where it is. A device behind a firewall Yagra does not monitor is taken as the way out, because it is the outermost one Yagra can see.
A Wi-Fi access point is drawn as a small round symbol, not a box. Its rim shows its state. When two or more access points hang off the same device, they are drawn as one bundle:
- The count sits on the bundle, and its rim is split by state in proportion.
- Underneath, it reads how many access points there are and which of them are not OK.
- Select the bundle to list its access points in the side panel, grouped by the port they reach the device through. Several access points on one port usually mean a switch Yagra does not monitor sits between them.
A lone access point is drawn on its own, with its name. When the device also has switches cabled below it, its access points are drawn beside it rather than under it, clear of the lines going down.
To find a device on a level, type in the search box above the map. It matches host names, and it takes a regular expression and NOT the same way the list filters do. Matches are outlined and everything else fades; the layout does not move. A bundle is outlined when any of its access points matches. Press Enter to bring the next match to the middle of the map, or Shift+Enter for the previous one.
The map shown when you select a folder on the Nodes page has the same search box, in the map’s heading row. The search stays while you move between folders. Open full map takes it to the full map.
A level with more than 2,000 linked nodes or 4,000 lines is not drawn. Its subfolders are still shown, and the map says so. The side panel shows what the level holds, and how many observations could not be resolved to links.
The walks that feed it run hourly by default, and each has its own switch at Settings ▸ Monitoring defaults ▸ Discovery walks: neighbours (CDP/LLDP), interface addresses, routing adjacency (OSPF/BGP), and the ARP cache. A fifth switch on the same page controls the interface media type walk, which feeds the Media column on a node’s Interfaces tab rather than the map.
The tables they read are sized by the device’s own peering mesh, not by the network. The routing table is never walked. A router carrying a full table has hundreds of thousands of routes.
Instead Yagra asks about one destination at a time, only of a device that holds a host address of its own, capped at 64 destinations. On a fleet of ordinary devices that issues no such queries at all.
The one route every device is asked about is its default route, which tells a site’s way out (see above). That is one or two short reads per device per hour, about 2 KB.
Two deliberate behaviours are worth knowing.
A down session still draws its link. A BGP session in active is a link with a fault, and that
is usually the thing being investigated.
An iBGP session between loopbacks does not become a link. A peer counts as adjacent only when it sits on a network the reporting device terminates, so a route reflector does not acquire a false star to every client it peers with.
Known limits, stated rather than half-answered. BGP4-MIB is IPv4-only, so IPv6 BGP peers are out of scope. OSPF collection is OSPFv2, and virtual links are not read.
A segment with more than two members, where no member can be identified as routing for the others, produces no links rather than a guessed one. It is counted in the map’s summary instead.
The same graph is what dependency suppression can be switched over to.
Hosts nothing is monitoring
Section titled “Hosts nothing is monitoring”Your monitored devices already see devices that are not in the inventory. Yagra lists them under Nodes ▸ Discovery ▸ Unregistered devices. No range is needed. Four kinds of source feed the list:
- LLDP and CDP neighbors that advertise a management address. A neighbor that says it is only a phone or a station is left out, so the phones behind an access switch do not fill the list.
- OSPF neighbors and BGP peers.
- Syslog and SNMP trap senders that no node claims.
- Hosts in a router’s ARP / IPv6 neighbor cache, when that walk is on (see the box below).
Only the ARP source needs switching on. Each row shows the address, the name a neighbor or a syslog header gave it, and where it was seen: one chip per source, then the device and port that saw it. The list is read 100 rows at a time. Load 100 more reads the next page. While pages remain, the filters search only the rows loaded so far, and the screen says so.
To add one from its row:
- Choose SNMP credentials in Credentials to try above the table. Scan a range uses the same list, so a change on either tab is a change on both.
- Press Detect. Yagra tries those credentials on that one device and picks a profile from what it answers.
- Check the profile and credential it filled in. If nothing answered, pick them by hand.
- Choose where it goes, above the table. Add to folder is the folder it lands in. With File by IP range on (the default), a device goes into the folder whose IP range holds its address instead. The line over Monitor says where it will go. It carries a warning mark when no folder’s range holds the address. While no folder has an IP range at all, File by IP range is not shown.
- Press Monitor. The device becomes a node through the same import path a range scan uses.
Who sees which rows. An account limited to some folders sees a row only when the device that first saw it is in one of those folders. It then sees only the evidence that devices in its folders reported. A row that only a syslog or trap sender vouches for is shown only to an unrestricted account.
Behind NAT, a sender’s address is the translator’s. It is shown as evidence and never imported on its own.
A row only a sender vouches for has no Detect or Monitor. A syslog or trap source address can be forged. Detecting it would send your credentials to whoever forged it, so the row is listed as information only. If the device is yours, add it with Add node. When the list is full, rows a monitored device saw are kept ahead of these, so a flood of forged senders cannot push a real device off it.
An account limited to some folders imports into a folder it can see. It cannot see the root of the tree, so an import that would land there is refused; choose a folder. A range scan’s import is refused for the same reason while any row would land at the root.
No scan is required. It is a by-product of the polling you already do.
Discovered endpoints are deliberately not drawn on the network map. An unmonitored host has no state to show, and a few thousand stateless boxes would bury the nodes that do. Import one and the ordinary derivation picks it up from there.
The same neighbors, from each device’s side. A node’s Neighbors tab lists what each port hears over CDP and LLDP. It also shows:
- Address — the neighbor’s management address. A badge beside the neighbor’s name says what that address is here: Monitored, Monitored (hidden) for a node in a folder you cannot see, Several nodes, or Not monitored. The Monitoring column filters by it.
- An address on more than one node. When exactly one of those nodes bears the name the neighbor sent, that node is the one linked. A node in a folder you cannot see is never picked this way, so such a row still reads Several nodes. Either way a mark beside the name says the address is shared, and opening the row lists the other nodes with whether their port carrying the address has link.
- Model / OS — the CDP platform with its OS version, or the LLDP system description.
- The maker the IEEE registered a MAC-address chassis ID to. That is the maker of the network interface, which may not be who made the device.
A neighbor that is not monitored has a Set up monitoring button. It opens the same Detect, folder and Monitor steps as the Unregistered devices tab. An access point that a wireless controller reports is added through that controller instead, and a device a Meraki organization lists links to the organization’s page. When there is no button, the row says why in its place: for example, the device is not on the Unregistered devices list yet (it is refreshed every five minutes), it announces itself as a phone, or a node outside your folders found it.
Meraki switches (MS), appliances (MX) and access points (MR) get the Neighbors tab too. Their neighbors are read from the Meraki Dashboard. An MX lists what its LAN ports hear, not its internet ports. An MR lists what its wired port hears — usually the switch it is plugged into. A Meraki peer that reports no management address is matched on the MAC its organization lists, so its row still says whether it is monitored.
Switches are read as often as neighbors are collected elsewhere (hourly by default). MX and MR are read one device at a time, within the organization’s Dashboard API rate. In a large organization a full round therefore takes longer than that: about two hours for 2,300 MX and MR at the default rate.
When an upgrade only changes how Yagra writes neighbors, the device’s neighbor history gets one entry marked Recorded differently after an upgrade — not a cabling change. Its rows are folded away, and how long the neighbors have been stable is not reset.
Thresholds
Section titled “Thresholds”Thresholds are per-metric rules that turn a measurement into a warning or critical state. Each has a bound, a direction (above / below), and dwell behavior evaluated by the alert engine — see Alerting.
Two rules of the model are worth knowing before you write any:
-
Thresholds apply to gauge metrics only. Not to counter metrics:
if_hc_in_octets, error or discard counters, and anything else the collection catalog declares a counter.A threshold on a counter would compare a raw monotonic total against the bound. An above rule latches permanently once the counter passes it, and a below rule fires a phantom alert at every reboot’s counter reset.
Yagra therefore rejects creating one:
POST /api/v1/thresholdson a counter metric answers 400counter_metric. It also reads counter samples as OK, which drains any alert an older rule had latched, through the normal recovery path.Rate-style alerting on counters is a query-time concern. Set thresholds on gauges.
-
Reading the threshold list requires the manage-monitoring permission (Operator and up), not just view access. A threshold set describes when and whom Yagra will page, so it stays closed even on a public dashboard. Anonymous requests answer 401.
For automation, GET /api/v1/thresholds returns an envelope,
{ "items": [...], "total": <n>, "truncated": <bool> }, capped at 500 rules per request.
?limit= can narrow that, never widen it, and truncated tells you when the cap bit.
What a rule applies to
Section titled “What a rule applies to”Every rule names a scope, and there are six. From broadest to narrowest they are: every node, a device profile, a node tag, a folder group, one node, and one interface.
The narrowest scope that matches wins. A profile rule sets the house bound; a node rule overrides it for one device; an interface rule overrides that for one port. Where two rules match at the same level, the more restrictive bound is used.
One rule may name several targets at that level — both Cisco IOS profiles, say, rather than one rule each. It may name at most 32, and an interface rule still names exactly one port.
The interface scope exists because utilisation, errors and optical level are per port. A rule written for the node would hold every port on it to the same bound.
A metric reported once per table row is judged per row. A row is a memory pool, a CPU, or a
temperature sensor. Each row is its own check, and its alert names the row:
cisco_mem_used_pct on I/O above 80 (was 83.9). Device health lists the five highest rows by name.
A device with more ends the list with “and N more”.
A rule can name the rows it applies to, in its Row name field (I/O, MPU Board *).
*matches anything, and case is ignored. An empty field means every row.- At the same scope, a rule with a row name wins over the rule for every row, for the rows it names. That is how one pool is loosened or tightened beside the rule for all of them.
- Scope still comes first. A node rule without a row name beats a profile rule with one.
- The poller reads row names from the device once an hour. Until a row’s name has been read, a rule with a row name does not match it, and the rule for every row applies.
The tag scope is not offered for a new rule — a folder group covers the same ground and is what the dialog builds. Existing tag rules keep resolving, and since v0.3.17 they have something to find: tags are editable from the WebUI, and a rule matches what a node effectively carries, its own plus everything its folder chain supplies.
What a new deployment starts with
Section titled “What a new deployment starts with”A fresh install seeds rules for reachability, SNMP reachability, packet loss, latency, SNMP walk
completeness, CPU, memory, temperature, disk, UPS battery and BGP peer state — with vendor-specific
bounds pointed at the profiles they apply to. The built-in URL and DNS monitor profiles carry rules of
their own, such as dns_up for a DNS monitor, so a fresh monitor alerts without configuration.
One default names a row. On Cisco IOS, the I/O memory pool warns at 90% and goes critical at 95%.
That pool holds packet buffers the switch pre-allocates, so it sits high on a healthy switch. Every
other pool keeps the ordinary memory rule.
All 30 are ordinary rules, node up/down included. Reachability alerting used to be a constant inside the alert engine with no way to change it. It is now a rule on Reachability at the every-node scope, seeded with the same three consecutive failures, so nothing about how quickly an existing fleet pages changes on upgrade. What changes is that you can retune it, override it per profile, group or node, or delete it.
Deleting it switches node-down paging off for whatever it covered. The node’s state in the list, the fleet summary, and dependency suppression are unaffected — those are not alerts, and removing an alert rule is not a request to stop tracking the node.
Polling behavior
Section titled “Polling behavior”The polling engine is built to be predictable at fleet scale:
-
Interval. A new installation polls every 300 seconds (five minutes) by default, clamped to 10–3600 s. An existing deployment keeps the interval it has stored.
At five minutes, a node that stops answering is reported down after about 15 minutes. The default reachability rule waits for three missed polls. Where that is too slow, lower the interval at Settings ▸ Monitoring defaults, or on a device profile.
The
YAGRA_POLL_INTERVAL_SECSenvironment variable seeds the value on first boot only. After that, the setting stored in the database is authoritative, and editable in the WebUI. Per-device-profile intervals can override the global default.Adding a node, or changing an interval, takes effect straight away. It does not wait for the scheduler’s next round.
-
Jitter. Polling intervals are jittered so tens of thousands of nodes don’t all fire on the same tick and stampede the network, or the pollers.
-
One probe in flight per device. A slow device never accumulates a pile-up of concurrent probes against itself. The next probe waits for the last.
DNS monitors are the deliberate exception. They share resolver targets by design, so they are governed by the global bound instead of the per-target one.
-
Global concurrency bound.
YAGRA_MAX_CONCURRENT_POLLS(default 256) caps a poller’s total concurrent probes, alongside per-device rate limiting and backpressure. It bounds what is in flight, not a rate: the polls per second it yields is that number divided by how long one probe takes.yagra_poll_cycles_missed_totalis the counter that says a poller is not keeping up.
An immediate, out-of-schedule poll of a single node can also be triggered on demand, without waiting for the next tick. It is a configuration-level (Admin) action, and an AI or automation client can trigger it through Yagra’s MCP tool surface.
Pollers are stateless and horizontally scalable. How work is spread across pools and sites, and how a remote poller rides out a network partition, is covered in Distributed polling.
See also
Section titled “See also”- Alerting — what happens when a check breaches: hysteresis, flapping detection, dependency suppression, and notification channels.
- Distributed polling — poller pools, location affinity, and store-and-forward.
- Configuration reference — the environment variables named on this page, with defaults and clamps.