Installation
Yagra is two long-running binaries plus a static WebUI:
- Yagra-core — orchestration, scheduling, the northbound REST API
- Yagra-poller — stateless ICMP/SNMP/API workers
- Yagra-web — the WebUI, served by nginx, which reverse-proxies
/apito core
This page covers every supported way to install that stack, from a one-command evaluation box to remote-site pollers streaming results home over a TLS bus.
Behind them sit up to six backing services:
| Backing service | Role | Needed by |
|---|---|---|
| PostgreSQL | Metadata: nodes, configuration, thresholds, users, alert history | core (required) |
| NATS (JetStream) | The core⇄poller bus: jobs, working sets, results, events | core + poller (required) |
| VictoriaMetrics | Time-series store — the metrics body | core (required) |
| Redis | Ephemeral poller liveness/assignment mirror — rebuildable | core (optional) |
| VictoriaLogs | Passive-event log store — full-text syslog/trap search | core (optional) |
| ClickHouse | Traffic-flow store — NetFlow/IPFIX/sFlow records | core (optional) |
Pollers talk to NATS only. Device credentials, job specs, and results all flow over the bus, and a poller never touches the databases. That is what makes pollers stateless, horizontally scalable, and deployable behind a remote site’s NAT.
Choosing a deployment
Section titled “Choosing a deployment”Two axes: single-node vs. distributed polling, and Docker vs. native processes. Whichever you pick, size the host from System requirements first.
| Docker Compose | Native (no Docker) | |
|---|---|---|
| Single node | A — pre-built images · B — build from source | C |
| Distributed pollers | D | E |
Start with A. It pulls the published images and needs no checkout and no build. It is also the only single-node deployment that can upgrade itself from the WebUI.
Reach for D or E once you need pollers at remote sites. They connect out to the central bus, so they cross NAT and firewalls without any inbound port at the site.
The rest are for narrower audiences, and all are supported. B builds from source, for developing on Yagra, auditing it, or making a custom build. C and E run the binaries directly, for hosts where Docker is not an option.
Scaling out is a config change, not a rewrite: the single-node and distributed deployments run the same images. You add remote pollers, and you turn on NATS TLS + auth before the bus leaves the machine — job messages carry device credentials.
High availability is orthogonal to this matrix. Any deployment with more than one core against
the same stores can enable leader election (YAGRA_ENABLE_HA) for automatic core failover.
docker-compose.ha.yml in the repository is a ready-made two-core overlay for trying it:
docker compose -f docker-compose.yml -f docker-compose.ha.yml up --buildExactly one core (the leader) answers 200 on /readyz. The standby answers 503 until it takes
over, which happens within seconds of the leader stopping. Route API and WebUI traffic on
/readyz. Details: high availability.
Images & tags
Section titled “Images & tags”Three images are published to GitHub Container Registry:
ghcr.io/horryworks/yagra-coreghcr.io/horryworks/yagra-pollerghcr.io/horryworks/yagra-webOnly releases are published. Development builds never reach the registry, so every tag below is a release you can run.
| Tag | Meaning | Use it for |
|---|---|---|
:v<version> |
One release (e.g. :v0.1.18) |
Production — pin these |
:latest |
The latest stable release. Pre-releases (-beta, -rc) never move it |
Evaluation, staying current |
:<git-sha> |
Immutable reference to one release | Reproducible deploys, rollback |
The Compose files select the tag via YAGRA_IMAGE_TAG (default latest). Upgrading means changing
the tag and recreating — see Upgrades & backups. Rollback is the same
operation with an older tag.
A — Single node, Docker (pre-built images)
Section titled “A — Single node, Docker (pre-built images)”The recommended deployment. docker-compose.deploy.yml does four things:
- Pulls the published images from GHCR. Nothing is built locally.
- Takes its settings from
.env. - Runs a one-shot
kek-initservice that writes a persistent key-encryption key, so stored monitoring credentials survive redeploys. - Ships the
yagra-updatersidecar, which is what makes Settings ▸ Upgrade work.
It needs no repository checkout. The composition is a single self-contained file, with a default for every variable it interpolates and no bind mounts outside the Docker socket.
-
Make a directory for the deployment and fetch the composition:
Terminal window mkdir yagra && cd yagracurl -fsSL -o docker-compose.deploy.yml \https://github.com/horryworks/Yagra/releases/latest/download/docker-compose.deploy.ymlTake the composition from the release, not from
main. It and the images it pulls are one artifact — a composition can require a container command or an init step that only exists in an image which has not been published yet.releases/latest/download/resolves to the latest stable release, which is the same thing the:latestimage tag means, so the two always match.Keep the file there, under that name. In-place upgrades read the directory back from a label on their own container, and refuse to run if it no longer holds a
docker-compose.deploy.yml. -
Create
.env. Only the database password really needs choosing, and it needs choosing now. PostgreSQL writes it into the data volume when it initialises, so changing it later takes anALTER ROLEas well as an edit.Use only characters that are safe in a URL. The password is placed into a connection URL and cannot be percent-encoded there, so one holding
/,@,:,?or#ends the URL early and core refuses to start.openssl rand -hex 16produces none of them.openssl rand -base64produces/most of the time, so it is the wrong generator here.POSTGRES_PASSWORD=change-me # set this before the first start# YAGRA_IMAGE_TAG=v0.2.5 # pin a release for production (default: latest)# YAGRA_API_PORT=8080 # host port for the API (plaintext)# YAGRA_WEB_PORT=443 # host port for the WebUI (HTTPS)# YAGRA_ADMIN_PASSWORD=choose-a-strong-password # else a one-time random one is loggedEverything else has a working default. Full variable list: the configuration reference.
-
Start it:
Terminal window docker compose -f docker-compose.deploy.yml up -dThat pulls as it goes. The images carry
pull_policy: always, so there is no separatepullstep and nothing to build. -
Retrieve the one-time
adminpassword (unless you setYAGRA_ADMIN_PASSWORD):Terminal window docker compose -f docker-compose.deploy.yml logs core | grep -i passwordOpen the WebUI at https://localhost, sign in as
admin, and change the password. Your browser will warn about the self-signed certificate — see WebUI TLS for how to replace it.
What’s running. Core, one poller, and the WebUI, plus:
- PostgreSQL, Redis, NATS, VictoriaMetrics
- VictoriaLogs, for passive-event search
- ClickHouse, for traffic flows
- an
ipasn-updatersidecar that keeps the offline IP→ASN dataset fresh. It is the only container with internet egress.
Migrations run automatically on core startup. Named volumes preserve all persistent data across
down/up.
| Purpose | Host default | Change via |
|---|---|---|
| WebUI (HTTPS) | 443 |
YAGRA_WEB_PORT |
REST API + /metrics (plaintext) |
8080 |
YAGRA_API_PORT |
| syslog intake (UDP) | 514 |
YAGRA_SYSLOG_PORT |
| SNMP trap intake (UDP) | 162 |
YAGRA_TRAP_PORT |
| NetFlow v5/v9 / IPFIX intake (UDP) | 2055 |
YAGRA_FLOW_PORT |
sFlow intake (UDP; listener off until YAGRA_SFLOW_BIND is set) |
6343 |
YAGRA_SFLOW_PORT |
The stores stay on the internal Docker network. Optional features switch off by setting their
variable empty in .env — YAGRA_CLICKHOUSE_URL= disables flow monitoring, and
YAGRA_SYSLOG_BIND= disables syslog intake. Full matrix:
ports · configuration.
Point your devices at it. Send syslog to the host’s :514/udp, SNMP traps (v1/v2c, informs
included) to :162/udp, and NetFlow/IPFIX export to :2055/udp.
Passive events correlate to nodes by the datagram’s source IP. If Docker’s bridge networking
rewrites source addresses on your host, switch the poller service to network_mode: host so the
real address survives.
Verify. All containers up, and core answering:
docker compose -f docker-compose.deploy.yml pscurl -fsS http://localhost:8080/healthzSettings ▸ Yagra health in the WebUI then shows each backing store’s reachability and the server’s own overall verdict.
Credential persistence — the KEK. Stored monitoring credentials (SNMP communities, SNMPv3 credentials, API tokens) are envelope-encrypted with a master key (KEK).
The kek-init service generates a 32-byte KEK into the kekdata volume once and never overwrites
it. Core mounts it read-only.
Without a persistent KEK, core falls back to an ephemeral key, regenerated on every restart. Every stored credential then becomes undecryptable after a redeploy.
B — Single node, Docker (build from source)
Section titled “B — Single node, Docker (build from source)”Yagra is AGPL-3.0 and builds from a clean checkout. This is the path for developing on Yagra,
auditing it, or making a custom build. docker-compose.yml builds the images locally, tags them
:dev, and runs the whole stack on one host:
git clone https://github.com/horryworks/Yagra.gitcd Yagradocker compose up --buildThen open the WebUI at https://localhost:8443 and sign in with the one-time admin password
printed in the core logs.
It publishes 8443 rather than 443 for two reasons: a laptop usually has 443 taken, and
rootless Docker cannot bind below 1024 at all. The certificate is self-signed on a first start, so
your browser will warn — see WebUI TLS.
WebUI TLS
Section titled “WebUI TLS”The WebUI is HTTPS out of the box. There is no plain-HTTP listener to publish by accident.
On a first start, core generates a self-signed certificate covering loopback and the container’s hostname. Your browser will warn, and will usually object to the name as well — nothing inside the container can know the address you will type.
Fix it from Settings ▸ TLS, which is reachable through the warning. There are two ways: import a PEM certificate chain and private key (pasted or from a file), or regenerate the self-signed one with the hostnames and IP addresses you actually use.
An imported certificate is live within seconds, with nothing restarted. Yagra refuses a mismatched pair, an expired certificate, or one with no subject alternative name, and says which — rather than failing at the next handshake.
Set YAGRA_WEB_TLS=off in .env when an external reverse proxy or load balancer already
terminates HTTPS in front of the container.
C — Single node, native
Section titled “C — Single node, native”Running the binaries directly, no Docker. You provision the stores yourself, build the workspace,
and run yagra-core + yagra-poller as services (e.g. systemd).
1. Provision the backing stores
Section titled “1. Provision the backing stores”Install and start each of these, reachable from the host that will run core:
-
PostgreSQL 17 — create a database and role. Core runs its migrations itself, but it does not create the database:
CREATE ROLE yagra LOGIN PASSWORD 'yagra';CREATE DATABASE yagra OWNER yagra; -
NATS 2.x with JetStream —
nats-server -js -
VictoriaMetrics —
victoria-metrics-prod --retentionPeriod=12(12 months of metrics) -
Redis 7 (optional) — enables the poller liveness/assignment mirror. Its absence only degrades, and never blocks startup
-
VictoriaLogs (optional) — full-text passive-event search. Without it, events stay entirely in PostgreSQL
-
ClickHouse 24.x (optional) — the traffic-flow store; without it, flow monitoring is off and the flow API answers 503
2. Build the workspace
Section titled “2. Build the workspace”Requires Rust 1.90 and (for the WebUI) Node 22. Build from the repository root so the workspace’s vendored dependency patches apply:
git clone https://github.com/horryworks/Yagra.gitcd Yagracargo build --release --workspace # → target/release/yagra-core, target/release/yagra-pollercd web && npm ci && npm run build # → web/dist/ (static SPA bundle)3. Provision the KEK (before first core start)
Section titled “3. Provision the KEK (before first core start)”This is the envelope-encryption master key: a persistent 32-byte file. Without it, core boots with an ephemeral dev key, and stored credentials will not survive a restart:
sudo install -d -m 0700 /etc/yagrahead -c 32 /dev/urandom | sudo tee /etc/yagra/kek > /dev/nullsudo chmod 0400 /etc/yagra/kekBack this file up. Losing it makes every stored credential permanently undecryptable.
4. Run core
Section titled “4. Run core”export YAGRA_DATABASE_URL="postgres://yagra:yagra@localhost:5432/yagra"export YAGRA_BUS_URL="nats://localhost:4222"export YAGRA_TSDB_URL="http://localhost:8428"export YAGRA_REDIS_URL="redis://localhost:6379" # optionalexport YAGRA_LOGS_URL="http://localhost:9428" # optional: passive-event log storeexport YAGRA_CLICKHOUSE_URL="http://localhost:8123" # optional: traffic-flow storeexport YAGRA_KEK_FILE="/etc/yagra/kek"export YAGRA_API_ADDR="0.0.0.0:8080" # default# export YAGRA_ADMIN_PASSWORD="choose-a-strong-password" # else a one-time random one is loggedexport RUST_LOG=info
./target/release/yagra-coreOn startup core does four things: connects to the stores, runs its migrations automatically,
seeds the built-in profiles and catalog, and serves /api/v1 + Prometheus /metrics on
YAGRA_API_ADDR. Check it with curl -fsS http://localhost:8080/healthz.
If YAGRA_ADMIN_PASSWORD was unset, grep the logs for the one-time admin password.
Three URLs are required: YAGRA_DATABASE_URL, YAGRA_BUS_URL and YAGRA_TSDB_URL. If any is
missing, core runs in an in-memory skeleton mode instead of live.
For a permanent installation, wrap this in a service unit (systemd or equivalent) with the same environment. Core is safe to restart at any time. Full variable list: the configuration reference.
5. Serve the WebUI
Section titled “5. Serve the WebUI”web/dist/ is a static bundle. Serve it with any web server and reverse-proxy /api to core.
Mirror the shipped nginx config (web/nginx.conf). Two things matter: SSE needs
proxy_buffering off and a long proxy_read_timeout, and the SPA needs a try_files fallback:
server { listen 80; root /var/www/yagra; # the contents of web/dist/ index index.html;
location /api/ { proxy_pass http://localhost:8080; # → core's YAGRA_API_ADDR proxy_set_header Host $host; proxy_set_header X-Real-IP $remote_addr; proxy_http_version 1.1; proxy_buffering off; # stream SSE immediately proxy_read_timeout 1h; # keep SSE connections open }
location / { try_files $uri $uri/ /index.html; # SPA client-side routing fallback }}6. Run the poller
Section titled “6. Run the poller”The poller needs raw sockets for ICMP. Either grant the capability to the binary, so it can run non-root, or run it as root:
sudo setcap cap_net_raw+ep ./target/release/yagra-poller
export YAGRA_BUS_URL="nats://localhost:4222"export YAGRA_POLLER_ID="poller-1" # unique per poller; defaults to hostnameexport YAGRA_POLLER_POOL="default"# Optional intake listeners, each off until bound (binding :514/:162 directly# would additionally need root or CAP_NET_BIND_SERVICE — these high ports don't):# export YAGRA_SYSLOG_BIND="0.0.0.0:1514"# export YAGRA_TRAP_BIND="0.0.0.0:1162"# export YAGRA_FLOW_BIND="0.0.0.0:2055"export RUST_LOG=info
./target/release/yagra-pollerThe poller exposes its own Prometheus /metrics on 0.0.0.0:9100.
D — Distributed pollers, Docker
Section titled “D — Distributed pollers, Docker”Run the full stack centrally (as in A) and add pollers at remote sites.
Each remote poller polls its site’s devices locally and streams results back over the bus. Nodes
carry a pool attribute. Core assigns each pool’s nodes across its live pollers by consistent
hashing, and fails them over automatically. Concepts and behavior in depth:
distributed polling.
Step 1 — Turn on NATS TLS + auth on the central stack
Section titled “Step 1 — Turn on NATS TLS + auth on the central stack”This is one switch in the WebUI. It replaces what used to be five manual steps — generating a
certificate with openssl, hand-editing two blocks of docker-compose.deploy.yml, and handing every
site the same password.
-
Go to Settings ▸ Pollers and find the Remote pollers panel. While the bus has never been exposed it reads Internal only.
-
Press “Accept remote pollers” and give the addresses remote pollers will dial — hostnames or IP addresses, comma-separated. A site cannot connect unless the exact address it dials is listed, so name every one.
-
Confirm the outage. The bus, this server and the poller running here are all restarted, so monitoring stops for about a minute. A fleet-wide maintenance window is opened first, so nothing pages. The page disconnects while it happens.
One switch does all of it: the bus certificate is reissued for the addresses you gave, TLS and a bus
password are turned on, the bus port is published, and the co-located core and poller move to
tls:// in the same change. Server-wide TLS leaves no plaintext port, which is why the last part is
not optional.
The panel then reads Encrypted and shows the certificate — what it covers, when it expires, and its fingerprint. Reissue certificate… mints a new one when a site’s address changes; every site has to be given the new file before it can reconnect.
The certificate is generated by Yagra and kept in PostgreSQL, the same way the WebUI’s own certificate is: the private key envelope-encrypted, the certificate itself plaintext because it is public. There is no import — a bus certificate only has to be trusted by pollers Yagra also configures.
The NATS configuration gives the core user full access and the poller user least privilege:
publish results, events and heartbeats; subscribe only to jobs and working-set assignments. It ships
inside the core image and is placed on the bus volume for you.
By default the one poller account is shared. Step 2 issues each poller a token of its own,
which is what makes a leak at one site stop being a key to every site. To narrow it further — so a
compromised site cannot even read another pool’s assignments — enable the optional Auth Callout step
described in the same Compose block. Details: security.
Step 2 — Register the poller in the WebUI
Section titled “Step 2 — Register the poller in the WebUI”Go to Settings ▸ Pollers ▸ “Register poller”. Give the poller a stable id and assign the pool you want it to serve, then press Issue token & download kit. The site does not have to exist yet — this is what registers it — and a single archive comes down holding everything the machine at the site needs:
| Member | What it is |
|---|---|
.env |
This poller’s id, pool, bus URL and its token. Written mode 0600. |
certs/server-cert.pem |
The bus certificate to pin. Public certificate only, never the key. |
docker-compose.poller.yml |
Taken from this core’s image, so it matches the version you are running. |
README.txt |
The site’s procedure, and what to do if the archive is lost. |
The token is stored only as a SHA-256 digest, so it is shown once, exists only inside that archive, and cannot be recovered. If the archive is lost, issue a new token — which invalidates the old one.
For a site that is already connected, the same archive is reachable from its row in the fleet table: the Token column shows which pollers still use the deployment-wide shared secret, and Issue token & download gives that one its own.
Step 3 — Run the remote poller
Section titled “Step 3 — Run the remote poller”On the remote-site machine, unpack the archive and bring it up. There is nothing to edit:
tar xzf yagra-poller-edge-tokyo-1.tar.gzdocker compose -f docker-compose.poller.yml up -dThe .env in the archive supplies everything Compose requires:
YAGRA_BUS_URL=tls://poller:<this poller's token>@core.example.com:4222YAGRA_POLLER_ID=edge-tokyo-1 # stable, unique per pollerYAGRA_POLLER_POOL=tokyo # the pool this poller servesYAGRA_BUS_CA_FILE=/etc/yagra/certs/server-cert.pemPin the image with YAGRA_IMAGE_TAG here too. Per-scenario extras — intake rate caps, flow
collection, the store-and-forward buffer — are all in the
configuration reference.
docker-compose.poller.yml uses host networking and grants only NET_RAW. Host networking is
needed for two reasons: passive syslog/trap correlation keys on the datagram source IP, which
bridge NAT would rewrite, and raw-socket ICMP wants the host’s interfaces directly.
The poller appears on Settings ▸ Pollers within a few seconds of starting, and core begins assigning that pool’s nodes to it.
A few properties worth knowing:
- Scaling a pool means running more pollers with the same
YAGRA_POLLER_POOLand distinctYAGRA_POLLER_IDs. Core rebalances the pool across them and fails over on loss. - A pool with zero live pollers falls back to publishing jobs one at a time, so no nodes go dark during a rollout.
- WAN outages don’t hole your history. The poller keeps polling locally and buffers results,
in memory and then spilling to the
pollerbufvolume, and replays them on reconnect. Metrics are backfilled at their original timestamps. Alerts are deliberately not backfilled — they resume from “now”. - Flow at the edge: with host networking the poller binds
:2055directly, so point the site’s NetFlow/IPFIX exporters at the poller. Flows are aggregated at the edge and streamed to core over the same TLS bus.
E — Distributed pollers, native
Section titled “E — Distributed pollers, native”Same as D, but the remote poller is the native binary instead of a container. The central bus TLS + auth setup (D, step 1) is unchanged.
On the remote host, build (or copy) the yagra-poller binary, drop the CA cert somewhere
readable, and run:
sudo setcap cap_net_raw+ep ./yagra-poller
export YAGRA_BUS_URL="tls://poller:a-strong-poller-bus-password@core.example.com:4222"export YAGRA_POLLER_ID="edge-tokyo-1" # unique per pollerexport YAGRA_POLLER_POOL="tokyo"export YAGRA_BUS_CA_FILE="/etc/yagra/certs/server-cert.pem"# Optional intake listeners (see the privileged-port caveat in D):# export YAGRA_SYSLOG_BIND="0.0.0.0:1514"# export YAGRA_TRAP_BIND="0.0.0.0:1162"# export YAGRA_FLOW_BIND="0.0.0.0:2055"export RUST_LOG=info
./yagra-pollerRun it on the host network, not a private network namespace. Passive-event source-IP correlation and raw ICMP need the site’s real interfaces.
Everything else — pools, registration, failover, buffering — behaves exactly as in D.
Upgrades & backups
Section titled “Upgrades & backups”Upgrades are designed to be low-effort, and to not lose or corrupt data. That is the design goal every migration and every message format is held to — not a promise that each release has been proven against.
So read the notes for the release you are moving to first, not only for behavior changes an operator or API client could notice, but because a release can withdraw the guarantee for itself and ask you to install fresh instead. When one does, it says so under Breaking changes and says what a fresh install costs. The changelog points at the notes for every version.
From the WebUI (v0.2.2+, deployment A)
Section titled “From the WebUI (v0.2.2+, deployment A)”This is the ordinary way to upgrade. Settings ▸ Upgrade lists the releases this deployment can move to.
Pressing Upgrade starts nothing. It opens the list of what this deployment is made of — core, the co-located poller, and every remote site — each with the version it runs now, the version it would move to, and a checkbox, all ticked. Pressing Upgrade again starts the work on the rows you left ticked.
For every row that moves, the page does the whole thing — six steps, in this order:
- Check there is room, before anything is written
- Take a backup
- Pull the images
- Install the composition carried inside the target image
- Recreate the containers
- Verify that it came back up
Step 1 is what stops a full host from being made worse. The backup is a full PostgreSQL dump plus a
VictoriaMetrics snapshot, and it lands on this host before any image is fetched — so without the
check, a host that was already full got several hundred MB written onto it and then failed to
pull. The run now stops with the two numbers in the message and the deployment untouched. The floor
is YAGRA_UPGRADE_MIN_FREE_BYTES — 3 GiB by default.
Once the new version is up and healthy, the sidecar clears up after itself: the release you came
from is kept, so the one-hop rollback still needs no download, and anything older goes, along with
the older backups. Only Yagra’s own three images are ever touched, and one backup is always kept.
YAGRA_UPGRADE_KEEP_RELEASES changes how many releases stay.
The work is carried out by a yagra-updater sidecar, the only container holding the Docker socket.
Core never has it. Requesting an upgrade needs manage-the-deployment, which only an Admin
holds; it is audited, and it is deliberately not on the MCP surface.
A switch on the same page turns the mechanism off. The setting lives in PostgreSQL, so it survives the upgrades it governs. That is unlike deleting the service from your compose file: each version installs its own composition, so a service deleted that way comes back on the next upgrade.
Step 4 replaces docker-compose.deploy.yml for the same reason, so an edit made to that file does
not survive an upgrade either. Put the changes this particular host needs in a file called
docker-compose.local.yml, beside it in the same directory. An upgrade never replaces that file,
and from v0.3.11 it is handed to Compose as a second -f whenever it exists. Compose merges a later
-f over an earlier one, so the overlay names only what it changes:
# docker-compose.local.yml — an upgrade never replaces this fileservices: poller: networks: [default, lab-devices]networks: lab-devices: external: trueStart the stack the way the upgrade does:
docker compose -p yagra -f docker-compose.deploy.yml -f docker-compose.local.yml up -dA deployment without such a file is unaffected — no file, no argument. The upgrade says which of the
two happened, with this deployment's docker-compose.local.yml or no docker-compose.local.yml here, so a deployment whose changes were not carried across reads differently from one that had
none to carry. Settings ▸ Pollers ▸ “Accept remote pollers” passes the overlay too, because that
switch recreates the poller by name.
Only deployment A has this. Every other way of installing Yagra — from source, natively, or from
a composition without the yagra-updater sidecar — has no such mechanism. The page says so plainly
instead of offering controls that would fail.
Pollers at remote sites come along too, and from v0.3.3 they do so by default.
A poller running in its own compose project at a monitored site is not part of the deployment’s project, so nothing central could restart it on its own.
The site bundle — Settings ▸ Pollers ▸ “Issue token & download” — writes
COMPOSE_PROFILES=self-upgrade into the .env it generates, and the checkbox that does so is
ticked. That profile starts a yagra-poller-updater sidecar at the site, and the site then declares
itself able to install a release. Core hands it the release it just installed. There are rules about
how:
- It goes out one poller at a time per pool, so a pool with two or more pollers keeps monitoring throughout.
- Images are fetched before anything stops, so a single-poller site is out for the container recreate rather than for the download.
- The poller never touches the Docker socket. It validates the command and writes it into the shared volume for its own sidecar, exactly as core does centrally.
To keep a site out, untick the box before issuing its bundle, or empty COMPOSE_PROFILES in that
site’s .env afterwards. Change it in .env rather than in the composition: an upgrade reinstalls
the composition from the release being installed, and never touches .env. A site issued its bundle
before v0.3.3 keeps its current behavior until it is handed a re-issued one.
Read what this grants before you roll it out. The sidecar runs as root with that site host’s Docker socket — the same trade the central one makes, and bounded the same way. See the upgrade sidecar.
A site that drifted can be brought back without upgrading core. Commands used to go out only as the tail of core’s own upgrade, so a site stood up afterwards — or one that failed and was fixed, or one that only enabled its updater today — stayed behind until the next release. From v0.3.4 you untick core in that list and press Upgrade: nothing in the deployment restarts, and each pool is still crossed one poller at a time. A poller ahead of core is moved back to it, because that is the skew direction the bus does not promise.
Before you press Upgrade, the list names every poller that will not come along and why, and every pool that has only one poller and so stops monitoring while its container is recreated. That warning follows your ticks, so unticking a row changes it. A release that would leave a poller two versions behind is shown with the reason and no checkbox, since the bus supports one release of skew.
The poller half of an upgrade has its own progress track. Core’s own steps end at “Verifying”; what follows — each site fetching the image over its own link, then one container recreated at a time per pool — can take another half hour. Each site now shows where it has got to, and the record stays on the page after the run, so a site that did not come back is named instead of quietly dropping off the live poller list.
Going back is a supported move within a declared window. A migration can declare a compatibility floor: the oldest version that can still run once that migration has been applied. A release below the current floor is shown greyed out with the reason rather than hidden.
Nothing is deleted by going back. Columns the newer version added stay in place, unread, and become visible again on the next upgrade.
There is a path for a site with no reachable registry at all. Run docker save on the three
release images where you can reach them, and upload the archive at the same page.
That path needs a second, deliberate opt-in on the host (YAGRA_UPGRADE_ALLOW_BUNDLE=1), because
docker load installs whatever the archive contains. Read the configuration
reference before turning it on.
From the command line
Section titled “From the command line”Upgrading from the command line is still supported. Use it when the sidecar is switched off or cannot run.
All you are doing is changing the image tag to the new version and recreating the containers.
-
Take a backup.
Always take one before a major upgrade. There are two things to save.
Terminal window # PostgreSQL — nodes, configuration, users, alert historydocker compose -f docker-compose.deploy.yml exec postgres \pg_dump -U yagra yagra > yagra-backup.sql# The KEK — tiny, and irreplaceabledocker run --rm -v yagra_kekdata:/kek busybox cat /kek/key > kek-backup.keyMetrics live in the
vmdatavolume. Snapshot it with your usual volume tooling, or with VictoriaMetrics’ own snapshot API. Redis needs no backup — it is rebuildable, so losing it is non-fatal. -
Choose the release you are moving to.
Skim the changelog. Anything whose behavior changed is listed there.
-
Pull the images at the new tag.
Terminal window # in .env: YAGRA_IMAGE_TAG=v0.2.2# or override it inline, as belowYAGRA_IMAGE_TAG=v0.2.2 docker compose -f docker-compose.deploy.yml pull -
Recreate the containers on those images.
Terminal window YAGRA_IMAGE_TAG=v0.2.2 docker compose -f docker-compose.deploy.yml up -dSchema changes run automatically when core starts. There is nothing to apply by hand.
Why this cannot corrupt your data
- Schema changes add first and remove later, in two stages. Stopping partway loses nothing, and
N→N+1 is always supported. To see what a release will do before it does it, run
yagra-core migrationsinside the target image. It prints the migration set that binary embeds as JSON, with no database and no configuration, so an upgrade can be planned before anything is touched. - Core and pollers can talk across one version of skew. A new core works with old pollers, so upgrade core first and pollers after — including remote sites, one at a time, in any order.
- Pollers are stateless. Replace them freely. A pool briefly without pollers falls back to publishing jobs one at a time, so no node goes dark mid-rollout.
- Persistent data is preserved. The
pgdata(PostgreSQL),vmdata(VictoriaMetrics), andkekdata(KEK) volumes — or their native equivalents — survive image upgrades untouched.
To go back, re-run the same steps with the previous tag. The immutable :<git-sha> tags exist
for exactly this.
Uninstalling
Section titled “Uninstalling”Yagra installs nothing on the host — no packages, no system services, no files outside the deployment directory and Docker’s own storage. Removing it is the install run backwards.
Run everything below from the directory holding docker-compose.deploy.yml.
Stop it, keep the data. Containers and the network go; every named volume stays, so up -d
brings the deployment back with its configuration, alert history and metrics intact.
docker compose -p yagra -f docker-compose.deploy.yml downRemove it completely. -v destroys the eleven named volumes — pgdata (nodes, users,
thresholds, alert history, every setting), vmdata (all metric history), vldata (all passive
events), chdata (all flow records), kekdata (the key-encryption key), and the certificate, log,
buffer and hand-off volumes.
docker compose -p yagra -f docker-compose.deploy.yml down -v --remove-orphansTo take the images and the directory too:
docker compose -p yagra -f docker-compose.deploy.yml down -v --rmi all --remove-orphanscd .. && rm -rf yagra # the compose file and .env, which holds POSTGRES_PASSWORD--rmi all also removes the backing-store images (postgres, redis, nats, victoria-metrics,
victoria-logs, clickhouse); Docker skips any another project is still using.
Remote pollers are separate stacks. A poller at a remote site (deployment D) is its own
Compose project on its own host. Nothing above touches it, and it will retry the bus forever after
the central stack is gone. Remove each on its own host with
docker compose -f docker-compose.poller.yml down -v.
Checking for leftovers. Compose labels everything it created, so nothing depends on still having the compose file:
docker ps -a --filter label=com.docker.compose.project=yagradocker volume ls --filter label=com.docker.compose.project=yagradocker network ls --filter label=com.docker.compose.project=yagraThere is no uninstall action in the WebUI — Settings ▸ Upgrade only moves a deployment between releases. Uninstalling is deliberately a host-side act.
Next steps
Section titled “Next steps”- Every environment variable, default, and clamp — the configuration reference
- Which port carries what, and what to firewall — the port matrix
- How pools, working sets, and failover actually behave — distributed polling
- Running more than one core — high availability
- TLS everywhere, the KEK, and per-poller bus credentials — security
- Structured logs, Prometheus metrics, and distributed tracing — observability