Distributed pollers
Yagra scales polling horizontally across pools of stateless pollers. Probes stay close to the devices they measure, and monitoring stays alive through network partitions.
This page explains how that behaves: assignment, failover, bus security, and the store-and-forward buffer. For the step-by-step deployment runbook, see the installation guide, section D.
None of this is a separate “distributed mode”. A single-node install already runs one poller in the
pool "default", with exactly the machinery described here. Going distributed means adding pollers
and pools, not changing architecture. The single-node and distributed deployments run the same
images.
Pools and assignment
Section titled “Pools and assignment”Every node carries a pool attribute (default: "default").
A poller identifies itself with a stable id (YAGRA_POLLER_ID) and the pool it serves
(YAGRA_POLLER_POOL), and sends heartbeats over the bus.
The id has to be stable, and since v0.3.2 the shipped composition makes it so for the poller inside
the central deployment: it is called local. Left to itself a poller takes its container hostname,
and Docker invents a new one every time Compose recreates the container — so each upgrade added
another dead row to Settings ▸ Pollers. Those rows are offline and safe to delete; the live
poller is the one named local.
Core’s coordinator tracks those heartbeats and assigns each pool’s nodes across its live pollers using consistent hashing. Adding or losing a poller therefore reshuffles the minimum number of nodes, not the whole pool.
Each poller receives its assignment as a working set pushed over the bus: a snapshot first, then incremental deltas as nodes are added, moved, or reassigned.
The poller schedules its own probes locally from that set — core does not micromanage individual checks. A poller that reconnects after a restart gets a full resync, so no node is left unassigned across a restart.
Failure handling follows from the same machinery:
-
A poller drops out. Its heartbeats stop, and the coordinator fails its nodes over to the surviving pollers in the same pool automatically.
From v0.2.3 the poller taking over starts polling the arriving nodes at a rate it can sustain (
YAGRA_ADOPT_RATE_PER_SEC, default 200 checks/sec), rather than spreading their first poll across the whole interval as if they were new. A handover of a few dozen nodes now costs a fraction of a second rather than up to a full interval. A cold start still spreads out exactly as before. -
A poller joins. The pool rebalances across the new membership, again moving only the minimum share. The rebalance is triggered immediately rather than at the next sweep.
-
A pool has zero live pollers. Core falls back to publishing that pool’s jobs individually over the bus, the path that predates working sets.
This is the rolling-upgrade safety net. A pool whose pollers are mid-restart, or still running the previous release, keeps its nodes monitored instead of going dark.
The liveness and assignment state lives in the optional Redis mirror and is rebuildable — losing it
degrades nothing durable. Scaling a pool is simply running more pollers with the same
YAGRA_POLLER_POOL and distinct YAGRA_POLLER_IDs.
Assigning a pool, and seeing who polls what
Section titled “Assigning a pool, and seeing who polls what”A node’s effective pool resolves in order:
- the node’s own pool
- the nearest ancestor folder that sets one
"default"
Pools are editable from the WebUI: on a node, on a folder, and by right-clicking either in the inventory tree, which offers the pools that already exist plus Inherit and a custom name.
Folders are the important case. The tree has no multi-select, so folder inheritance is how you move a site’s worth of nodes at once.
Each choice shows whether that pool has a live poller. This matters more than it looks. Assigning nodes to a pool with no registered poller publishes their jobs to a bus subject nothing subscribes to, and those nodes quietly stop being monitored.
Node detail answers the reverse question. Pool shows the effective pool and where it came from. Polled by names the poller actually responsible, as one of five states:
- assigned
- pending — assignment not yet published
- legacy fan-out — the zero-poller fallback above
- Meraki — cloud-collected, so not ring-assigned
- unknown
That answer is read from the working set core actually published, not recomputed from the hash ring. So it never invents an owner for a node the sweep dropped, and it cannot disagree with the poller’s own node list. Settings ▸ Pollers goes the other way, drilling into any poller’s node set inline.
Upgrades lean on the same properties. The bus is version-tolerant between adjacent releases, so a new core works with the previous release’s pollers. Upgrade core first, then pollers — including remote sites, one at a time, in any order.
Pollers are stateless, so replacing one is free. A pool briefly without pollers rides the per-job fallback until its upgraded pollers re-register. The full procedure is in upgrades & backups.
Covering a pool that has lost its poller
Section titled “Covering a pool that has lost its poller”When a pool’s nodes have nothing left to poll them, Settings ▸ Pollers offers to have another pool cover it — most obviously after moving a deployment to a new server, where the remote sites do not follow until their bundles are reissued. Every node and folder is recorded before it moves, so Put them back returns each one to exactly the assignment it had, including the ones that were inheriting rather than assigned.
Location affinity
Section titled “Location affinity”Pools map naturally to sites. Pollers are stateless and dial out to the central bus, so a remote site needs one outbound rule and no inbound holes. A poller behind NAT or a firewall works unmodified.
Run a pool per site (“tokyo”, “osaka”, a cloud region) and assign that site’s nodes to it. Probes then stay local: ICMP and SNMP traffic never crosses the WAN, and only results, events, and heartbeats stream back to core.
The intake edge follows the pool. A site poller is also where the site’s passive telemetry should land.
Its optional UDP listeners — syslog, SNMP traps, NetFlow/IPFIX, sFlow — accept the local devices’ output, apply per-source and global rate limits at the edge, and carry the events home over the same authenticated bus.
Flow datagrams are additionally aggregated at the edge before streaming to core: fixed time buckets, top flows by bytes per exporter. That is what keeps a busy site’s flow volume affordable on the uplink.
The remote-poller composition runs on the host network, because passive events correlate to nodes by the datagram’s source IP and bridge NAT would rewrite the address.
Budget the site uplink for what does cross it. Pollers carry the original received bytes alongside the parsed events, so that forwarding can later relay exactly what a device sent. For flow that works out to roughly a megabit per second per 1,000 flows/s.
The per-source and global intake rate limits in the configuration reference cap the bursts.
A typical multi-site fleet looks like:
| Site | Pool | Pollers |
|---|---|---|
| Datacenter / HQ | default |
co-located with core |
| Tokyo branch | tokyo |
1–2 remote, on the branch LAN |
| Osaka branch | osaka |
1–2 remote, on the branch LAN |
Within each pool, redundancy is just headcount. Run two pollers in a site’s pool and either one picks up the other’s nodes on failure.
Do not mistake the zero-live-poller fallback described above for site redundancy. It is an upgrade-compatibility path. Site redundancy is a second poller at the site.
One special case is Cisco Meraki monitoring, which polls the Meraki cloud API at the
organization level rather than probing devices on-site. Those cloud-collection jobs are routed to
the pool named by YAGRA_MERAKI_POOL (default default). Point it at whichever pool has internet
egress.
Registering pollers
Section titled “Registering pollers”Pollers are registered from the WebUI at Settings ▸ Pollers ▸ “Register poller”.
The dialog generates a ready-to-use .env for the remote host, supplying the three variables a
remote poller requires: YAGRA_POLLER_ID (stable and unique), YAGRA_POLLER_POOL, and the tls://
bus URL.
Issuing that poller a bus token of its own goes further: it downloads a single archive holding
the .env, the bus certificate to pin, docker-compose.poller.yml taken from this core’s image, and
a README. The site’s whole procedure becomes unpack and docker compose up -d. The token is stored
only as a SHA-256 digest, so it is shown once and cannot be recovered.
That file drops straight in next to docker-compose.poller.yml, or into the native binary’s
environment. Pin the image with YAGRA_IMAGE_TAG there too.
The poller appears on the page within a few seconds of starting, and core begins assigning that pool’s nodes to it.
The Settings ▸ Pollers page is also the fleet view of your polling layer. It shows, per poller:
- Liveness — current status and the last heartbeat received.
- Assignment — the pool it serves and its current working-set size, so you can see how a pool’s nodes are distributed.
- Version — the poller’s running version, useful mid-rollout.
- Monitoring gaps — the recent windows during which core lost contact with a poller (which poller, which pool, when, and for how long), so you can see at a glance when monitoring was blind and confirm the metrics were backfilled.
It also warns when a pool has nodes but no live poller. The full deployment walkthrough — compose file, host networking, and the privileged-port caveat for the intake listeners — is in the installation guide, section D.
Creating a pool, and moving a poller into it
Section titled “Creating a pool, and moving a poller into it”Until v0.3.4 a pool existed only as a side effect of typing an unknown name into a node’s
assignment. The pool strip at the top of Settings ▸ Pollers now creates one deliberately, and
each card’s ⋮ edits its description, renames it, or removes it. A pool cannot be deleted while
nodes, folders, or pollers still name it. The default pool can be described but not renamed or
deleted, because its name is fixed in the code.
A poller moves between pools from the same page — drag its grip onto a pool card, or use Move in
its Pool cell. Nothing restarts: core owns which pool a poller serves, tells the poller on its
next working-set snapshot, and the poller re-points its bus subscriptions in place.
YAGRA_POLLER_POOL still decides where a poller lands the first time core sees it and is ignored
after that, so recreating a container no longer reverts a move.
If the poller being moved is the last live one of its pool and that pool still has nodes, the move stops and asks which you want: bring those nodes to the destination as well — one transaction, so monitoring never stops — or leave them and accept that they stop being polled. A poller that is offline, or whose build predates v0.3.4, cannot be moved at all; the page says why rather than half-moving it.
Securing the bus
Section titled “Securing the bus”Job messages carry plaintext device credentials, because the poller needs them to probe. On a single host that is fine, since the bus never leaves the internal Docker network.
The moment the bus crosses a trust boundary to a remote site, it must be TLS-encrypted and
authenticated — and that has to happen first. Never publish NATS :4222 plaintext.
The setup is one switch: Settings ▸ Pollers ▸ “Accept remote pollers”, given the addresses your
sites will dial. It reissues the bus certificate for those addresses, turns on TLS and a bus password,
publishes the bus port, and moves the co-located core and poller to tls:// in the same change. It is
installation, section D, step 1.
Each poller pins the server’s certificate via YAGRA_BUS_CA_FILE, so a remote site trusts exactly
one bus endpoint rather than the system CA store.
The bundled NATS configuration gives the core user full access and the poller user least
privilege: publish results, events, and heartbeats; subscribe only to jobs and working-set
assignments.
Be aware of its limit. A poller with no token of its own is admitted by the deployment-wide bootstrap secret, and any poller admitted that way can read any pool’s assignments. The shared secret authenticates the fleet, but it is not a tenant boundary between sites.
Two things narrow it. A poller id that is not in the inventory is refused whatever secret it presents,
so a leaked .env cannot be used to claim another site’s id. And a poller that has been issued its
own token can no longer be admitted by the shared secret at all — which is why issuing tokens is a
per-site reduction in blast radius rather than a formality.
If pool isolation matters for your deployment, scope credentials per poller as described next.
Per-poller credential scoping
Section titled “Per-poller credential scoping”Optionally, core can act as the bus’s authentication service (NATS Auth Callout) and mint each poller a per-connection credential scoped to its own pool’s subjects.
At connect time, core validates the connecting poller against the shared bootstrap secret
(YAGRA_NATS_POLLER_PASSWORD). It then issues a credential granting exactly the subjects for that
poller’s own id and pool — its own job stream, its own assignment subject — and nothing else.
The effect is that a compromised remote site can only ever see the jobs and device credentials for its own pool, instead of holding a shared account with fleet-wide reach.
What each poller can read is decided centrally by core at connect time, not by editing the NATS server configuration per site.
Since v0.3.2 it comes on with remote-poller acceptance, and there is nothing to set up. The account key core signs with is generated on first start and kept sealed in the database — the same envelope encryption every other secret uses — and its public half is written into the bus’s own configuration by the same one-shot that writes the rest of it. Two halves that cannot disagree, instead of two an operator has to keep in step.
On a deployment whose bus never leaves the host, the responder is not started and NATS uses the static accounts above. Nothing changes until you accept remote pollers. Details: security.
Before v0.3.2 this was four manual steps ending in an edit to the shipped NATS configuration — a file that is reinstalled from the image on every start, so the edit lasted until the next one. If you carried them out,
YAGRA_NATS_CALLOUT_SEED_FILEstill takes precedence and nothing breaks.
Store-and-forward
Section titled “Store-and-forward”A remote poller cut off from core — a WAN outage, a firewall blip — keeps monitoring. It continues polling its devices locally and buffers the results instead of dropping them, in two tiers:
- An in-memory ring (20,000 results by default). Buffering never blocks the poll loop.
- An on-disk spill when the ring fills. Segments are written under
/var/lib/yagra/buffer(thepollerbufvolume in Compose), so the buffer survives a poller restart mid-outage.
The buffer is bounded on every axis, and the overflow policy is drop-oldest everywhere. It can never fill the poller’s disk:
| Cap | Default | Variable |
|---|---|---|
| Master switch | on | YAGRA_STORE_FORWARD (off disables) |
| In-memory ring | 20,000 results | YAGRA_STORE_FORWARD_MEM_MAX |
| On-disk spill, total | 512 MB — oldest segment dropped past it | YAGRA_STORE_FORWARD_DISK_MAX_MB |
| Maximum age | 24 h — older results are dropped at replay | YAGRA_STORE_FORWARD_MAX_AGE_SECS |
| Free-space floor | 1 GB — spilling stops below this much free disk | YAGRA_STORE_FORWARD_DISK_FREE_FLOOR_MB |
| Spill segment size | 16 MiB — the granularity of the disk cap | YAGRA_STORE_FORWARD_SEGMENT_MB |
| Spill directory | /var/lib/yagra/buffer |
YAGRA_STORE_FORWARD_DIR |
The buffer also degrades safely rather than failing. If the spill directory cannot be created or read, the poller logs a warning and continues with memory-only buffering. A storage problem never crashes polling.
On reconnect the poller bulk-replays the buffer over a dedicated backfill channel, permitted to the poller account on a secured bus too. Core imports the metrics at their original timestamps, so graphs and history fill in the outage window with no false spike.
What replay carries is exactly the measurement record: metric samples and interface metadata, nothing else. The spill holds no secrets while it waits on disk.
Alerts are never backfilled. Alert evaluation resumes from “now” after a reconnect.
This is deliberate. Dwell-time hysteresis is sample-count based, so replaying a backlog of old samples would fabricate state transitions and flood you with stale, already-resolved alerts.
A recovered link therefore restores your history without re-raising incidents. Each healed outage window is recorded as a monitoring gap on Settings ▸ Pollers, so the blind period stays visible even though it never paged anyone.
Store-and-forward is on by default. Set YAGRA_STORE_FORWARD=off to restore plain publish-live
behavior, where results are dropped if the bus is unreachable. All the caps are in the
configuration reference.
Observing pollers
Section titled “Observing pollers”The polling layer is observable from three angles:
- Settings ▸ Pollers — per-poller health: status, pool, version, working-set size, last heartbeat, pool warnings, and the recent monitoring-gaps list described above.
- Dashboard — a working-set distribution widget shows how each pool’s nodes are spread across its pollers. An unbalanced or degraded pool is visible at a glance.
- Prometheus — every poller serves its own
/metricson:9100, alongside core’s metrics on the API port, so your existing monitoring can watch the monitors. See observability.
Together with the gaps list, this answers the operational questions in order: is every pool covered → who owns which nodes → when was I blind, and was it backfilled?