Skip to content

Architecture

A single-node Yagra runs co-located in a few containers. Growing that into distributed pollers and highly available stores takes configuration, not a rewrite. The pieces are decoupled from the start so that this holds.

Component Role Runs in
Core Orchestration, scheduling, and the northbound REST API core
Poller ICMP / SNMP / API polling and passive intake — stateless and horizontally scalable poller
WebUI Dashboards and visualization (React + TypeScript) browser
Bus Job distribution and poller fan-out over NATS core + poller
Transport The ICMP / SNMP / HTTP abstraction all device I/O goes through poller
Topology Dependency graph powering suppression and the network map core
Discovery Device discovery and classification core + poller
Alert State machine, hysteresis, and dependency suppression core
Ingest Syslog and SNMP-trap parsing, flow decoding, and edge rate limiting poller
Forward Filtered tee of received events and flows to external collectors core
Secrets Envelope-encrypted device credentials — only core holds the key core
Telemetry Structured logs, Prometheus metrics, OpenTelemetry traces core + poller

Only the first three are things you deploy. Two Rust processes carry everything else — core (orchestration-side) and the poller (device-side) — plus the static WebUI served by nginx.

The rest of the table is libraries compiled into one or both of them. There is no ingest container, no alert service, no separate discovery daemon. The “runs in” column says which process each one ends up inside.

One pair of names is worth keeping apart. Bus is the client crate compiled into core and into every poller. NATS is the broker they talk through — not a component but a backing service, listed with the stores below.

Core and pollers communicate only through the bus, never by direct calls. That is what makes pollers stateless, horizontally scalable, and deployable at remote sites across NAT and firewalls.

A remote poller dials out to the central bus rather than accepting inbound connections, so a site needs one outbound rule and no inbound holes.

A poller holds no durable state and touches no database. Everything it needs — job specs, device credentials, working-set assignments — arrives over the bus. Everything it produces — results, events, flow batches, heartbeats — leaves the same way.

Job messages can carry device credentials, so exposing the bus to remote sites requires the bundled TLS + auth configuration. See security.

Each kind of data lives in the store built for it. Of the six backing services, three are required. The other three are optional, and switch their feature on when configured. Same framing as the installation guide:

Backing service Holds Required?
PostgreSQL Metadata: nodes, configuration, thresholds, users, alert history Required
NATS The core ⇄ poller bus: jobs, working sets, results, events Required
VictoriaMetrics Time-series metrics Required
Redis Ephemeral poller liveness/assignment mirror — rebuildable, losing it is safe Optional
VictoriaLogs Passive-event log store — full-text syslog/trap search Optional
ClickHouse Traffic-flow records Optional

Two rules sit behind the table. High-cardinality time-series never goes into PostgreSQL. Durable configuration never goes into Redis.

Without VictoriaLogs, passive events stay entirely in PostgreSQL. Without ClickHouse, flow monitoring is off.

Pollers store raw SNMP counters. Rates and utilization are derived at query time, which handles counter wrap and resets correctly — and means pollers hold no previous-value state.

Core schedules every check and publishes the work to the bus. It adds jitter to the schedule, so thousands of nodes don’t probe on the same tick.

The assigned poller runs the probe through the transport abstraction (ICMP, SNMP, or HTTP) and publishes the result back.

Core evaluates state and thresholds, drives the alerting pipeline, and writes the raw metric samples to VictoriaMetrics.

Nothing rate-related is computed at collection time. Rates come out of the time-series store at query and evaluation time.

A device sends syslog or an SNMP trap to a poller’s UDP listener. The poller parses it, applies per-source and global rate limits at the edge, and forwards the event over the bus — along with a copy of the original datagram bytes.

Core correlates it to a node by source IP, matches it against event rules to raise or clear alerts, and persists it to PostgreSQL. With VictoriaLogs configured, it goes there too, for full-text search.

Inbound webhooks skip the poller. They arrive directly on core’s API with a per-source bearer token. Details: passive events.

A flow exporter sends NetFlow, IPFIX, or sFlow datagrams to a poller’s flow listener.

The poller decodes them and aggregates at the edge: fixed time buckets, top flows by bytes per exporter. It streams that aggregate to core, alongside the verbatim datagrams, so that forwarding can relay exactly what the device sent.

Core enriches flows with AS numbers from the offline IP→ASN dataset and writes them to ClickHouse. Details: traffic flow.

Forwarding relays received syslog, traps, and flow exports onward to external collectors. It egresses from core, not the pollers.

Pollers carry the original received bytes to core over the bus. The active core applies the per-destination filters and does all relaying and BigQuery streaming. That leaves one egress point to firewall, regardless of how many sites feed it.

Poller pools. Every node carries a pool attribute, and each pool’s nodes are spread across its live pollers by consistent hashing. Add a poller to a pool and core rebalances. Lose one and its nodes fail over to the survivors.

Pools map naturally to sites: a Tokyo pool’s pollers monitor Tokyo’s devices locally. See distributed polling.

Working sets. Core pushes each poller its assigned node set over the bus, as a snapshot plus incremental deltas. The poller schedules its own probes locally.

Polling capacity therefore scales horizontally without core micromanaging every check. A pool that briefly has no pollers falls back to publishing jobs one at a time, so no node goes dark.

High availability. Multiple cores can run against the same stores with automatic leader election. One leader does the work, standbys take over within seconds, and /readyz tells the load balancer which is which. See high availability.

Version upgrades are designed to be low-effort and non-destructive. Three things make that true:

  • Database migrations run automatically on core startup, adding first and removing later, in two stages.
  • Bus messages tolerate one version of skew. A new core works with the previous release’s pollers during a rolling upgrade, so upgrade core first, then pollers in any order.
  • Persistent stores are preserved across upgrades, with a pre-upgrade backup step.

The procedure lives in the installation guide.

Yagra is AGPL-3.0 and the source is public, so it is worth saying how it is laid out — and what changed in v0.3.0.

Through the v0.2 series the backend grew a handful of very large Rust files in which several unrelated jobs sat side by side. Core’s startup file had reached 5,066 lines and the poller’s 1,097. A file of that shape is a hazard rather than an inconvenience: nothing tells a reader that the piece they are editing is also read by something they have never opened.

v0.3.0 took those files apart. Each was split along the line its own contents already followed, rather than by size:

Area Split by
analysis/ which store a Troubleshoot analysis reads
scheduler/ what each part needs in order to run — the pure half never awaits
events/ the program each part belongs to (query, storage, matching, ingest)
repo/ the database table each method’s SQL names
reports/ the stage of producing one report
config_bundle/ the direction a configuration travels
mcp/tools/ the REST domain each tool mirrors
poller worker/ how the check talks to the device

Each boundary is held by a test that fails the build when something crosses it — a query naming a table its file does not declare, a scheduler function that awaits, a poller check reaching for the transport directly. The rules are executable, so they cannot rot the way a written convention does.

Core’s startup file is now 2,100 lines and holds starting up and nothing else; the poller’s is 450. None of this changed what Yagra does. It changes what the next change costs.