Architecture
A single-node Yagra runs co-located in a few containers. Growing that into distributed pollers and highly available stores takes configuration, not a rewrite. The pieces are decoupled from the start so that this holds.
Components
Section titled “Components”| Component | Role | Runs in |
|---|---|---|
| Core | Orchestration, scheduling, and the northbound REST API | core |
| Poller | ICMP / SNMP / API polling and passive intake — stateless and horizontally scalable | poller |
| WebUI | Dashboards and visualization (React + TypeScript) | browser |
| Bus | Job distribution and poller fan-out over NATS | core + poller |
| Transport | The ICMP / SNMP / HTTP abstraction all device I/O goes through | poller |
| Topology | Dependency graph powering suppression and the network map | core |
| Discovery | Device discovery and classification | core + poller |
| Alert | State machine, hysteresis, and dependency suppression | core |
| Ingest | Syslog and SNMP-trap parsing, flow decoding, and edge rate limiting | poller |
| Forward | Filtered tee of received events and flows to external collectors | core |
| Secrets | Envelope-encrypted device credentials — only core holds the key | core |
| Telemetry | Structured logs, Prometheus metrics, OpenTelemetry traces | core + poller |
Only the first three are things you deploy. Two Rust processes carry everything else — core (orchestration-side) and the poller (device-side) — plus the static WebUI served by nginx.
The rest of the table is libraries compiled into one or both of them. There is no ingest container, no alert service, no separate discovery daemon. The “runs in” column says which process each one ends up inside.
One pair of names is worth keeping apart. Bus is the client crate compiled into core and into every poller. NATS is the broker they talk through — not a component but a backing service, listed with the stores below.
The core ⇄ poller boundary
Section titled “The core ⇄ poller boundary”Core and pollers communicate only through the bus, never by direct calls. That is what makes pollers stateless, horizontally scalable, and deployable at remote sites across NAT and firewalls.
A remote poller dials out to the central bus rather than accepting inbound connections, so a site needs one outbound rule and no inbound holes.
A poller holds no durable state and touches no database. Everything it needs — job specs, device credentials, working-set assignments — arrives over the bus. Everything it produces — results, events, flow batches, heartbeats — leaves the same way.
Job messages can carry device credentials, so exposing the bus to remote sites requires the bundled TLS + auth configuration. See security.
Store separation
Section titled “Store separation”Each kind of data lives in the store built for it. Of the six backing services, three are required. The other three are optional, and switch their feature on when configured. Same framing as the installation guide:
| Backing service | Holds | Required? |
|---|---|---|
| PostgreSQL | Metadata: nodes, configuration, thresholds, users, alert history | Required |
| NATS | The core ⇄ poller bus: jobs, working sets, results, events | Required |
| VictoriaMetrics | Time-series metrics | Required |
| Redis | Ephemeral poller liveness/assignment mirror — rebuildable, losing it is safe | Optional |
| VictoriaLogs | Passive-event log store — full-text syslog/trap search | Optional |
| ClickHouse | Traffic-flow records | Optional |
Two rules sit behind the table. High-cardinality time-series never goes into PostgreSQL. Durable configuration never goes into Redis.
Without VictoriaLogs, passive events stay entirely in PostgreSQL. Without ClickHouse, flow monitoring is off.
Pollers store raw SNMP counters. Rates and utilization are derived at query time, which handles counter wrap and resets correctly — and means pollers hold no previous-value state.
Data paths
Section titled “Data paths”The poll lifecycle
Section titled “The poll lifecycle”Core schedules every check and publishes the work to the bus. It adds jitter to the schedule, so thousands of nodes don’t probe on the same tick.
The assigned poller runs the probe through the transport abstraction (ICMP, SNMP, or HTTP) and publishes the result back.
Core evaluates state and thresholds, drives the alerting pipeline, and writes the raw metric samples to VictoriaMetrics.
Nothing rate-related is computed at collection time. Rates come out of the time-series store at query and evaluation time.
Passive events
Section titled “Passive events”A device sends syslog or an SNMP trap to a poller’s UDP listener. The poller parses it, applies per-source and global rate limits at the edge, and forwards the event over the bus — along with a copy of the original datagram bytes.
Core correlates it to a node by source IP, matches it against event rules to raise or clear alerts, and persists it to PostgreSQL. With VictoriaLogs configured, it goes there too, for full-text search.
Inbound webhooks skip the poller. They arrive directly on core’s API with a per-source bearer token. Details: passive events.
Traffic flow
Section titled “Traffic flow”A flow exporter sends NetFlow, IPFIX, or sFlow datagrams to a poller’s flow listener.
The poller decodes them and aggregates at the edge: fixed time buckets, top flows by bytes per exporter. It streams that aggregate to core, alongside the verbatim datagrams, so that forwarding can relay exactly what the device sent.
Core enriches flows with AS numbers from the offline IP→ASN dataset and writes them to ClickHouse. Details: traffic flow.
Forwarding egress
Section titled “Forwarding egress”Forwarding relays received syslog, traps, and flow exports onward to external collectors. It egresses from core, not the pollers.
Pollers carry the original received bytes to core over the bus. The active core applies the per-destination filters and does all relaying and BigQuery streaming. That leaves one egress point to firewall, regardless of how many sites feed it.
Scaling model
Section titled “Scaling model”Poller pools. Every node carries a pool attribute, and each pool’s nodes are spread across its
live pollers by consistent hashing. Add a poller to a pool and core rebalances. Lose one and its
nodes fail over to the survivors.
Pools map naturally to sites: a Tokyo pool’s pollers monitor Tokyo’s devices locally. See distributed polling.
Working sets. Core pushes each poller its assigned node set over the bus, as a snapshot plus incremental deltas. The poller schedules its own probes locally.
Polling capacity therefore scales horizontally without core micromanaging every check. A pool that briefly has no pollers falls back to publishing jobs one at a time, so no node goes dark.
High availability. Multiple cores can run against the same stores with automatic leader
election. One leader does the work, standbys take over within seconds, and /readyz tells the load
balancer which is which. See high availability.
Upgrades
Section titled “Upgrades”Version upgrades are designed to be low-effort and non-destructive. Three things make that true:
- Database migrations run automatically on core startup, adding first and removing later, in two stages.
- Bus messages tolerate one version of skew. A new core works with the previous release’s pollers during a rolling upgrade, so upgrade core first, then pollers in any order.
- Persistent stores are preserved across upgrades, with a pre-upgrade backup step.
The procedure lives in the installation guide.
How the code is organised
Section titled “How the code is organised”Yagra is AGPL-3.0 and the source is public, so it is worth saying how it is laid out — and what changed in v0.3.0.
Through the v0.2 series the backend grew a handful of very large Rust files in which several unrelated jobs sat side by side. Core’s startup file had reached 5,066 lines and the poller’s 1,097. A file of that shape is a hazard rather than an inconvenience: nothing tells a reader that the piece they are editing is also read by something they have never opened.
v0.3.0 took those files apart. Each was split along the line its own contents already followed, rather than by size:
| Area | Split by |
|---|---|
analysis/ |
which store a Troubleshoot analysis reads |
scheduler/ |
what each part needs in order to run — the pure half never awaits |
events/ |
the program each part belongs to (query, storage, matching, ingest) |
repo/ |
the database table each method’s SQL names |
reports/ |
the stage of producing one report |
config_bundle/ |
the direction a configuration travels |
mcp/tools/ |
the REST domain each tool mirrors |
poller worker/ |
how the check talks to the device |
Each boundary is held by a test that fails the build when something crosses it — a query naming a table its file does not declare, a scheduler function that awaits, a poller check reaching for the transport directly. The rules are executable, so they cannot rot the way a written convention does.
Core’s startup file is now 2,100 lines and holds starting up and nothing else; the poller’s is 450. None of this changed what Yagra does. It changes what the next change costs.