System requirements
Yagra runs as a set of Docker containers on one host, and it is lighter than its feature list suggests. A complete stack with every optional feature switched on holds 1.0–1.4 GB of RAM.
What moves that number is not how many nodes you monitor. It is which optional stores you run, and how many metric series you retain and for how long.
The tables below separate what has been measured from what has not. Where a figure is an estimate, it says so.
Recommended host
Section titled “Recommended host”| Nodes | vCPU | RAM | Disk | Basis | |
|---|---|---|---|---|---|
| Minimum — evaluation | up to ~50 | 2 | 4 GB | 20 GB | Measured: a stock install holds 1.0–1.4 GB and pulls ~2.3 GB of images |
| Recommended — full stack, every feature on | up to ~500 | 4 | 8 GB | 100 GB | Measured on a 4 vCPU / 7.6 GiB host running everything |
| Medium | ~5,000 | 8 | 16 GB | 500 GB – 1 TB+ | Estimated. Disk is decided by collected interfaces, not nodes — use the formula below |
| Large | ~50,000 | 16 | 32 GB | 1 TB+ | Estimated, except core itself (see below) |
The footprint columns are measured; the node column is not. The RAM and disk figures in the two
smaller rows come from docker stats and docker images on hosts of that size. The node ceilings in
every row are extrapolation: the largest fleet we run continuously is 32 nodes, and the 50,000-node
figures further down come from a synthetic replay rather than a polled fleet.
Treat the node column as a starting point, then watch Settings ▸ Yagra health, which shows the real CPU, load, memory and per-filesystem usage of core and every poller.
Any x86-64 or ARM64 host with Docker and the Compose plugin will do. On Linux the poller uses
raw-socket ICMP and needs the NET_RAW capability. The bundled Compose files grant it.
What a stock install actually starts
Section titled “What a stock install actually starts”docker-compose.deploy.yml defines no Compose profiles, so docker compose up -d starts every
service — ClickHouse, VictoriaLogs and the IP→ASN updater included. The optional stores are
opt-out: you remove the service from the file (and unset the matching URL on core) if you do not
want it.
So unless you have edited that file, size for the “everything on” row — 1.0–1.4 GB of RAM and ~2.3 GB of images — not for the required-services subtotal.
Per-container memory, measured
Section titled “Per-container memory, measured”Measured with docker stats --no-stream against two v0.3.3 deployments, each monitoring 32 nodes
with every optional feature enabled. Host A is 4 vCPU / 7.6 GiB, host B is 8 vCPU / 7.8 GiB:
| Container | Host A | Host B | Role |
|---|---|---|---|
clickhouse |
872 MiB | 696 MiB | Traffic flow (NetFlow/IPFIX/sFlow) — optional, started by default |
core |
141 MiB | 111 MiB | Required — orchestration, scheduler, API |
victoriametrics |
104 MiB | 117 MiB | Required — metrics store |
victorialogs |
98 MiB | 11 MiB | Passive events (syslog, traps) — optional, started by default |
poller |
58 MiB | 62 MiB | Required — ICMP/SNMP/API probing |
postgres |
48 MiB | 39 MiB | Required — configuration, users, alert history |
nats |
17 MiB | 6 MiB | Required — core⇄poller bus |
yagra-updater |
7 MiB | 1 MiB | Bundled — in-place upgrades |
web |
6 MiB | 8 MiB | Required — WebUI (nginx) |
redis |
6 MiB | 3 MiB | Bundled — poller liveness and assignment |
ipasn-updater |
5 MiB | 8 MiB | IP→ASN enrichment for flow — optional, started by default |
The two columns are the same version and the same inventory, so the spread is what these services do
rather than what they are. victorialogs is the widest: it is largely cache, and it grows with
ingest and query volume. Do not plan on its low figure.
Which adds up to:
| Configuration | Memory |
|---|---|
| Required services only — after editing them out of the Compose file | ~0.4 GB |
| …plus passive event monitoring | ~0.4–0.5 GB |
| Everything on — what a stock install runs | ~1.0–1.4 GB |
ClickHouse is more than half the stack on its own. If you are not collecting traffic flow, removing it is the single largest saving available, in both memory and CPU.
It is also the first service to move to its own host as you grow, because it is I/O and page-cache
heavy. Only YAGRA_CLICKHOUSE_URL has to change.
CPU has not been the binding constraint at any size we have measured.
On the two 32-node deployments above, all Yagra containers together used 6% and 38% of one
core in a docker stats snapshot. On both hosts, more than 70% of that was ClickHouse and
VictoriaLogs housekeeping. In the 50,000-node test below, core averaged 11.8% of one core.
Size for RAM and disk. Give CPU enough headroom for polling bursts and report rendering, but it is not the number to plan around.
Disk is driven by metric series count × retention, not by node count:
bytes ≈ series × (86400 ÷ interval_seconds) × retention_days × ~1 byte per sampleMeasured on the two hosts above, a stored sample costs 0.72 and 0.78 bytes including the index. The sample data alone compresses to 0.39–0.41 bytes; the inverted index is the rest, and on these hosts it is 46–48% of the bytes on disk. One byte per sample is therefore still a safe ceiling — but the margin is the index, not compression headroom. (These are lab hosts with heavy series churn, which inflates the index share. A stable fleet should sit lower.)
How many series is a node?
Section titled “How many series is a node?”Measured, per collected object:
| Object | Series | Which |
|---|---|---|
| ICMP check, node answering | 2 | icmp_loss_pct and icmp_rtt_ms — an unreachable node stores loss only |
| SNMP node, base | ~4 | snmp_up, snmp_neighbor_count, snmp_l3_address_count, snmp_routing_adjacency_count |
| Each collected interface | 11 | admin and oper status, HC in/out octets, HC in/out unicast packets, in/out errors, in/out discards, high speed |
| …with optics | +2 | if_rx_power_dbm, if_tx_power_dbm |
Vendor health metrics — CPU, memory, temperature — add a handful more per node.
Derived values (utilisation, bps, packet rates) are computed at query time and stored nowhere, so they cost no series at all.
One interface costs more series than the node it hangs off. That is the number to plan around.
Two worked examples, at the shipped defaults — a 60-second interval and 12-month VictoriaMetrics retention:
| Fleet | Series | Metrics disk |
|---|---|---|
| 50,000 ICMP-only nodes | ~100,000 | ~53 GB |
| 5,000 SNMP nodes × 24 collected interfaces | ~1,350,000 | ~710 GB |
The second row is over thirteen times the disk of the first, with a tenth of the nodes. That is the point: interface and table monitoring, not node count, is what fills a volume. Raising metrics retention to 24 months doubles both figures.
Watch series cardinality. An unbounded label is the fastest way to fill one.
Budget on top of that:
- Container images — measured on v0.3.3: Yagra’s three come to 620 MB, and the infrastructure
images another 910 MB (PostgreSQL 424 MB, the updater’s
dockerCLI 291 MB, Redis 58 MB, VictoriaMetrics 54 MB, VictoriaLogs 42 MB, NATS 41 MB). ClickHouse adds 807 MB on top. A stock install pulls about 2.3 GB. - Upgrade headroom — an in-place upgrade pulls a new set of images while keeping the previous ones for rollback. Leave 2 GB free.
- Passive events — VictoriaLogs retains 30 days by default, and flow data in ClickHouse expires after 30 days. (The 90-day default is the alert-linked event rows kept in PostgreSQL, not the log store.)
- Remote pollers — each buffers results to disk when it cannot reach core, capped at 512 MB, and stops when free space falls to 1 GB.
Retention is not all in one place. Settings ▸ System sets what Yagra itself prunes — alert-linked
events (90 days), unmatched events (24 hours), report runs (90 days), diagnostics (90 days) and flow
data (30 days). The two Victoria stores are different: neither has a runtime retention API, so their
--retentionPeriod flags live in the Compose file, and Settings ▸ System shows those values
read-only with a note saying so. Metrics retention — the one that drives the table above — is
changed by editing that flag and recreating the container.
What scales, and what does not
Section titled “What scales, and what does not”Core has been measured against a 50,000-node fleet:
| Measured | |
|---|---|
| Core memory, average | 183 MiB |
| Core memory, peak | 245 MiB |
| Core CPU, average | 11.8% of one core |
| Ingest lag | ~1.0 s |
Core holding fifty thousand nodes therefore costs a little over what core holding a few dozen costs.
Core is not what you size for — the stores are. Plan around VictoriaMetrics’ series count and ClickHouse’s flow volume, and treat core and the pollers as close to fixed cost.
Remote poller hosts
Section titled “Remote poller hosts”A poller at a remote site is small. These figures are now measured, on two remote-site pollers running v0.3.3 against a core on another host:
| Measured | Recommended | |
|---|---|---|
| Memory | 28 MiB and 34 MiB resident | 1 GB |
| CPU | 0.2% of one core | 1 vCPU |
| Disk | 2.0 GB used, buffer empty | 4 GB |
Each was holding a small working set — 59 and 68 check specs — so read those as a floor rather than a ceiling. The recommendation stays well above them because the store-and-forward buffer is what actually sizes a poller: it holds up to 20,000 results in memory before spilling to disk, and the spill file is capped at 512 MB with a 1 GB free-space floor it will not cross.
The disk recommendation is the poller image (132 MB) plus that cap and that floor.
A poller keeps no persistent state beyond the buffer, so it is always safe to rebuild.
How these numbers were measured
Section titled “How these numbers were measured”Per-container memory and CPU come from docker stats --no-stream against two running v0.3.3
deployments with every feature enabled, each monitoring 32 nodes. Image sizes come from
docker images on the same hosts.
Series counts per object come from VictoriaMetrics’ own /api/v1/status/tsdb on those deployments.
Bytes per sample is vm_data_size_bytes over vm_rows, storage and index together.
The 50,000-node figures come from a seeded fleet driven by a synthetic result generator, sampled over a 20-minute window after a 5-minute settle.
The remote-poller figures come from docker stats and the poller’s own /metrics on two
remote-site hosts.
Disk totals for the worked examples are calculated from the formula rather than measured. Estimates are labelled as such throughout.