Skip to content

Installation

Yagra is two long-running binaries plus a static WebUI:

  • Yagra-core — orchestration, scheduling, the northbound REST API
  • Yagra-poller — stateless ICMP/SNMP/API workers
  • Yagra-web — the WebUI, served by nginx, which reverse-proxies /api to core

This page covers every supported way to install that stack, from a one-command evaluation box to remote-site pollers streaming results home over a TLS bus.

Behind them sit up to six backing services:

Backing service Role Needed by
PostgreSQL Metadata: nodes, configuration, thresholds, users, alert history core (required)
NATS (JetStream) The core⇄poller bus: jobs, working sets, results, events core + poller (required)
VictoriaMetrics Time-series store — the metrics body core (required)
Redis Ephemeral poller liveness/assignment mirror — rebuildable core (optional)
VictoriaLogs Passive-event log store — full-text syslog/trap search core (optional)
ClickHouse Traffic-flow store — NetFlow/IPFIX/sFlow records core (optional)

Pollers talk to NATS only. Device credentials, job specs, and results all flow over the bus, and a poller never touches the databases. That is what makes pollers stateless, horizontally scalable, and deployable behind a remote site’s NAT.

Two axes: single-node vs. distributed polling, and Docker vs. native processes. Whichever you pick, size the host from System requirements first.

Docker Compose Native (no Docker)
Single node A — pre-built images · B — build from source C
Distributed pollers D E

Start with A. It pulls the published images and needs no checkout and no build. It is also the only single-node deployment that can upgrade itself from the WebUI.

Reach for D or E once you need pollers at remote sites. They connect out to the central bus, so they cross NAT and firewalls without any inbound port at the site.

The rest are for narrower audiences, and all are supported. B builds from source, for developing on Yagra, auditing it, or making a custom build. C and E run the binaries directly, for hosts where Docker is not an option.

Scaling out is a config change, not a rewrite: the single-node and distributed deployments run the same images. You add remote pollers, and you turn on NATS TLS + auth before the bus leaves the machine — job messages carry device credentials.

High availability is orthogonal to this matrix. Any deployment with more than one core against the same stores can enable leader election (YAGRA_ENABLE_HA) for automatic core failover. docker-compose.ha.yml in the repository is a ready-made two-core overlay for trying it:

Terminal window
docker compose -f docker-compose.yml -f docker-compose.ha.yml up --build

Exactly one core (the leader) answers 200 on /readyz. The standby answers 503 until it takes over, which happens within seconds of the leader stopping. Route API and WebUI traffic on /readyz. Details: high availability.

Three images are published to GitHub Container Registry:

ghcr.io/horryworks/yagra-core
ghcr.io/horryworks/yagra-poller
ghcr.io/horryworks/yagra-web

Only releases are published. Development builds never reach the registry, so every tag below is a release you can run.

Tag Meaning Use it for
:v<version> One release (e.g. :v0.1.18) Production — pin these
:latest The latest stable release. Pre-releases (-beta, -rc) never move it Evaluation, staying current
:<git-sha> Immutable reference to one release Reproducible deploys, rollback

The Compose files select the tag via YAGRA_IMAGE_TAG (default latest). Upgrading means changing the tag and recreating — see Upgrades & backups. Rollback is the same operation with an older tag.

A — Single node, Docker (pre-built images)

Section titled “A — Single node, Docker (pre-built images)”

The recommended deployment. docker-compose.deploy.yml does four things:

  • Pulls the published images from GHCR. Nothing is built locally.
  • Takes its settings from .env.
  • Runs a one-shot kek-init service that writes a persistent key-encryption key, so stored monitoring credentials survive redeploys.
  • Ships the yagra-updater sidecar, which is what makes Settings ▸ Upgrade work.

It needs no repository checkout. The composition is a single self-contained file, with a default for every variable it interpolates and no bind mounts outside the Docker socket.

  1. Make a directory for the deployment and fetch the composition:

    Terminal window
    mkdir yagra && cd yagra
    curl -fsSL -o docker-compose.deploy.yml \
    https://github.com/horryworks/Yagra/releases/latest/download/docker-compose.deploy.yml

    Take the composition from the release, not from main. It and the images it pulls are one artifact — a composition can require a container command or an init step that only exists in an image which has not been published yet. releases/latest/download/ resolves to the latest stable release, which is the same thing the :latest image tag means, so the two always match.

    Keep the file there, under that name. In-place upgrades read the directory back from a label on their own container, and refuse to run if it no longer holds a docker-compose.deploy.yml.

  2. Create .env. Only the database password really needs choosing, and it needs choosing now. PostgreSQL writes it into the data volume when it initialises, so changing it later takes an ALTER ROLE as well as an edit.

    Use only characters that are safe in a URL. The password is placed into a connection URL and cannot be percent-encoded there, so one holding /, @, :, ? or # ends the URL early and core refuses to start. openssl rand -hex 16 produces none of them. openssl rand -base64 produces / most of the time, so it is the wrong generator here.

    POSTGRES_PASSWORD=change-me # set this before the first start
    # YAGRA_IMAGE_TAG=v0.2.5 # pin a release for production (default: latest)
    # YAGRA_API_PORT=8080 # host port for the API (plaintext)
    # YAGRA_WEB_PORT=443 # host port for the WebUI (HTTPS)
    # YAGRA_ADMIN_PASSWORD=choose-a-strong-password # else a one-time random one is logged

    Everything else has a working default. Full variable list: the configuration reference.

  3. Start it:

    Terminal window
    docker compose -f docker-compose.deploy.yml up -d

    That pulls as it goes. The images carry pull_policy: always, so there is no separate pull step and nothing to build.

  4. Retrieve the one-time admin password (unless you set YAGRA_ADMIN_PASSWORD):

    Terminal window
    docker compose -f docker-compose.deploy.yml logs core | grep -i password

    Open the WebUI at https://localhost, sign in as admin, and change the password. Your browser will warn about the self-signed certificate — see WebUI TLS for how to replace it.

What’s running. Core, one poller, and the WebUI, plus:

  • PostgreSQL, Redis, NATS, VictoriaMetrics
  • VictoriaLogs, for passive-event search
  • ClickHouse, for traffic flows
  • an ipasn-updater sidecar that keeps the offline IP→ASN dataset fresh. It is the only container with internet egress.

Migrations run automatically on core startup. Named volumes preserve all persistent data across down/up.

Purpose Host default Change via
WebUI (HTTPS) 443 YAGRA_WEB_PORT
REST API + /metrics (plaintext) 8080 YAGRA_API_PORT
syslog intake (UDP) 514 YAGRA_SYSLOG_PORT
SNMP trap intake (UDP) 162 YAGRA_TRAP_PORT
NetFlow v5/v9 / IPFIX intake (UDP) 2055 YAGRA_FLOW_PORT
sFlow intake (UDP; listener off until YAGRA_SFLOW_BIND is set) 6343 YAGRA_SFLOW_PORT

The stores stay on the internal Docker network. Optional features switch off by setting their variable empty in .env — YAGRA_CLICKHOUSE_URL= disables flow monitoring, and YAGRA_SYSLOG_BIND= disables syslog intake. Full matrix: ports · configuration.

Point your devices at it. Send syslog to the host’s :514/udp, SNMP traps (v1/v2c, informs included) to :162/udp, and NetFlow/IPFIX export to :2055/udp.

Passive events correlate to nodes by the datagram’s source IP. If Docker’s bridge networking rewrites source addresses on your host, switch the poller service to network_mode: host so the real address survives.

Verify. All containers up, and core answering:

Terminal window
docker compose -f docker-compose.deploy.yml ps
curl -fsS http://localhost:8080/healthz

Settings ▸ Yagra health in the WebUI then shows each backing store’s reachability and the server’s own overall verdict.

Credential persistence — the KEK. Stored monitoring credentials (SNMP communities, SNMPv3 credentials, API tokens) are envelope-encrypted with a master key (KEK).

The kek-init service generates a 32-byte KEK into the kekdata volume once and never overwrites it. Core mounts it read-only.

Without a persistent KEK, core falls back to an ephemeral key, regenerated on every restart. Every stored credential then becomes undecryptable after a redeploy.

B — Single node, Docker (build from source)

Section titled “B — Single node, Docker (build from source)”

Yagra is AGPL-3.0 and builds from a clean checkout. This is the path for developing on Yagra, auditing it, or making a custom build. docker-compose.yml builds the images locally, tags them :dev, and runs the whole stack on one host:

Terminal window
git clone https://github.com/horryworks/Yagra.git
cd Yagra
docker compose up --build

Then open the WebUI at https://localhost:8443 and sign in with the one-time admin password printed in the core logs.

It publishes 8443 rather than 443 for two reasons: a laptop usually has 443 taken, and rootless Docker cannot bind below 1024 at all. The certificate is self-signed on a first start, so your browser will warn — see WebUI TLS.

The WebUI is HTTPS out of the box. There is no plain-HTTP listener to publish by accident.

On a first start, core generates a self-signed certificate covering loopback and the container’s hostname. Your browser will warn, and will usually object to the name as well — nothing inside the container can know the address you will type.

Fix it from Settings ▸ TLS, which is reachable through the warning. There are two ways: import a PEM certificate chain and private key (pasted or from a file), or regenerate the self-signed one with the hostnames and IP addresses you actually use.

An imported certificate is live within seconds, with nothing restarted. Yagra refuses a mismatched pair, an expired certificate, or one with no subject alternative name, and says which — rather than failing at the next handshake.

Set YAGRA_WEB_TLS=off in .env when an external reverse proxy or load balancer already terminates HTTPS in front of the container.

Running the binaries directly, no Docker. You provision the stores yourself, build the workspace, and run yagra-core + yagra-poller as services (e.g. systemd).

Install and start each of these, reachable from the host that will run core:

  • PostgreSQL 17 — create a database and role. Core runs its migrations itself, but it does not create the database:

    CREATE ROLE yagra LOGIN PASSWORD 'yagra';
    CREATE DATABASE yagra OWNER yagra;
  • NATS 2.x with JetStream — nats-server -js

  • VictoriaMetrics — victoria-metrics-prod --retentionPeriod=12 (12 months of metrics)

  • Redis 7 (optional) — enables the poller liveness/assignment mirror. Its absence only degrades, and never blocks startup

  • VictoriaLogs (optional) — full-text passive-event search. Without it, events stay entirely in PostgreSQL

  • ClickHouse 24.x (optional) — the traffic-flow store; without it, flow monitoring is off and the flow API answers 503

Requires Rust 1.90 and (for the WebUI) Node 22. Build from the repository root so the workspace’s vendored dependency patches apply:

Terminal window
git clone https://github.com/horryworks/Yagra.git
cd Yagra
cargo build --release --workspace # → target/release/yagra-core, target/release/yagra-poller
cd web && npm ci && npm run build # → web/dist/ (static SPA bundle)

3. Provision the KEK (before first core start)

Section titled “3. Provision the KEK (before first core start)”

This is the envelope-encryption master key: a persistent 32-byte file. Without it, core boots with an ephemeral dev key, and stored credentials will not survive a restart:

Terminal window
sudo install -d -m 0700 /etc/yagra
head -c 32 /dev/urandom | sudo tee /etc/yagra/kek > /dev/null
sudo chmod 0400 /etc/yagra/kek

Back this file up. Losing it makes every stored credential permanently undecryptable.

Terminal window
export YAGRA_DATABASE_URL="postgres://yagra:yagra@localhost:5432/yagra"
export YAGRA_BUS_URL="nats://localhost:4222"
export YAGRA_TSDB_URL="http://localhost:8428"
export YAGRA_REDIS_URL="redis://localhost:6379" # optional
export YAGRA_LOGS_URL="http://localhost:9428" # optional: passive-event log store
export YAGRA_CLICKHOUSE_URL="http://localhost:8123" # optional: traffic-flow store
export YAGRA_KEK_FILE="/etc/yagra/kek"
export YAGRA_API_ADDR="0.0.0.0:8080" # default
# export YAGRA_ADMIN_PASSWORD="choose-a-strong-password" # else a one-time random one is logged
export RUST_LOG=info
./target/release/yagra-core

On startup core does four things: connects to the stores, runs its migrations automatically, seeds the built-in profiles and catalog, and serves /api/v1 + Prometheus /metrics on YAGRA_API_ADDR. Check it with curl -fsS http://localhost:8080/healthz.

If YAGRA_ADMIN_PASSWORD was unset, grep the logs for the one-time admin password.

Three URLs are required: YAGRA_DATABASE_URL, YAGRA_BUS_URL and YAGRA_TSDB_URL. If any is missing, core runs in an in-memory skeleton mode instead of live.

For a permanent installation, wrap this in a service unit (systemd or equivalent) with the same environment. Core is safe to restart at any time. Full variable list: the configuration reference.

web/dist/ is a static bundle. Serve it with any web server and reverse-proxy /api to core. Mirror the shipped nginx config (web/nginx.conf). Two things matter: SSE needs proxy_buffering off and a long proxy_read_timeout, and the SPA needs a try_files fallback:

server {
listen 80;
root /var/www/yagra; # the contents of web/dist/
index index.html;
location /api/ {
proxy_pass http://localhost:8080; # → core's YAGRA_API_ADDR
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_http_version 1.1;
proxy_buffering off; # stream SSE immediately
proxy_read_timeout 1h; # keep SSE connections open
}
location / {
try_files $uri $uri/ /index.html; # SPA client-side routing fallback
}
}

The poller needs raw sockets for ICMP. Either grant the capability to the binary, so it can run non-root, or run it as root:

Terminal window
sudo setcap cap_net_raw+ep ./target/release/yagra-poller
export YAGRA_BUS_URL="nats://localhost:4222"
export YAGRA_POLLER_ID="poller-1" # unique per poller; defaults to hostname
export YAGRA_POLLER_POOL="default"
# Optional intake listeners, each off until bound (binding :514/:162 directly
# would additionally need root or CAP_NET_BIND_SERVICE — these high ports don't):
# export YAGRA_SYSLOG_BIND="0.0.0.0:1514"
# export YAGRA_TRAP_BIND="0.0.0.0:1162"
# export YAGRA_FLOW_BIND="0.0.0.0:2055"
export RUST_LOG=info
./target/release/yagra-poller

The poller exposes its own Prometheus /metrics on 0.0.0.0:9100.

Run the full stack centrally (as in A) and add pollers at remote sites.

Each remote poller polls its site’s devices locally and streams results back over the bus. Nodes carry a pool attribute. Core assigns each pool’s nodes across its live pollers by consistent hashing, and fails them over automatically. Concepts and behavior in depth: distributed polling.

Step 1 — Turn on NATS TLS + auth on the central stack

Section titled “Step 1 — Turn on NATS TLS + auth on the central stack”

This is one switch in the WebUI. It replaces what used to be five manual steps — generating a certificate with openssl, hand-editing two blocks of docker-compose.deploy.yml, and handing every site the same password.

  1. Go to Settings ▸ Pollers and find the Remote pollers panel. While the bus has never been exposed it reads Internal only.

  2. Press “Accept remote pollers” and give the addresses remote pollers will dial — hostnames or IP addresses, comma-separated. A site cannot connect unless the exact address it dials is listed, so name every one.

  3. Confirm the outage. The bus, this server and the poller running here are all restarted, so monitoring stops for about a minute. A fleet-wide maintenance window is opened first, so nothing pages. The page disconnects while it happens.

One switch does all of it: the bus certificate is reissued for the addresses you gave, TLS and a bus password are turned on, the bus port is published, and the co-located core and poller move to tls:// in the same change. Server-wide TLS leaves no plaintext port, which is why the last part is not optional.

The panel then reads Encrypted and shows the certificate — what it covers, when it expires, and its fingerprint. Reissue certificate… mints a new one when a site’s address changes; every site has to be given the new file before it can reconnect.

The certificate is generated by Yagra and kept in PostgreSQL, the same way the WebUI’s own certificate is: the private key envelope-encrypted, the certificate itself plaintext because it is public. There is no import — a bus certificate only has to be trusted by pollers Yagra also configures.

The NATS configuration gives the core user full access and the poller user least privilege: publish results, events and heartbeats; subscribe only to jobs and working-set assignments. It ships inside the core image and is placed on the bus volume for you.

By default the one poller account is shared. Step 2 issues each poller a token of its own, which is what makes a leak at one site stop being a key to every site. To narrow it further — so a compromised site cannot even read another pool’s assignments — enable the optional Auth Callout step described in the same Compose block. Details: security.

Step 2 — Register the poller in the WebUI

Section titled “Step 2 — Register the poller in the WebUI”

Go to Settings ▸ Pollers ▸ “Register poller”. Give the poller a stable id and assign the pool you want it to serve, then press Issue token & download kit. The site does not have to exist yet — this is what registers it — and a single archive comes down holding everything the machine at the site needs:

Member What it is
.env This poller’s id, pool, bus URL and its token. Written mode 0600.
certs/server-cert.pem The bus certificate to pin. Public certificate only, never the key.
docker-compose.poller.yml Taken from this core’s image, so it matches the version you are running.
README.txt The site’s procedure, and what to do if the archive is lost.

The token is stored only as a SHA-256 digest, so it is shown once, exists only inside that archive, and cannot be recovered. If the archive is lost, issue a new token — which invalidates the old one.

For a site that is already connected, the same archive is reachable from its row in the fleet table: the Token column shows which pollers still use the deployment-wide shared secret, and Issue token & download gives that one its own.

On the remote-site machine, unpack the archive and bring it up. There is nothing to edit:

Terminal window
tar xzf yagra-poller-edge-tokyo-1.tar.gz
docker compose -f docker-compose.poller.yml up -d

The .env in the archive supplies everything Compose requires:

YAGRA_BUS_URL=tls://poller:<this poller's token>@core.example.com:4222
YAGRA_POLLER_ID=edge-tokyo-1 # stable, unique per poller
YAGRA_POLLER_POOL=tokyo # the pool this poller serves
YAGRA_BUS_CA_FILE=/etc/yagra/certs/server-cert.pem

Pin the image with YAGRA_IMAGE_TAG here too. Per-scenario extras — intake rate caps, flow collection, the store-and-forward buffer — are all in the configuration reference.

docker-compose.poller.yml uses host networking and grants only NET_RAW. Host networking is needed for two reasons: passive syslog/trap correlation keys on the datagram source IP, which bridge NAT would rewrite, and raw-socket ICMP wants the host’s interfaces directly.

The poller appears on Settings ▸ Pollers within a few seconds of starting, and core begins assigning that pool’s nodes to it.

A few properties worth knowing:

  • Scaling a pool means running more pollers with the same YAGRA_POLLER_POOL and distinct YAGRA_POLLER_IDs. Core rebalances the pool across them and fails over on loss.
  • A pool with zero live pollers falls back to publishing jobs one at a time, so no nodes go dark during a rollout.
  • WAN outages don’t hole your history. The poller keeps polling locally and buffers results, in memory and then spilling to the pollerbuf volume, and replays them on reconnect. Metrics are backfilled at their original timestamps. Alerts are deliberately not backfilled — they resume from “now”.
  • Flow at the edge: with host networking the poller binds :2055 directly, so point the site’s NetFlow/IPFIX exporters at the poller. Flows are aggregated at the edge and streamed to core over the same TLS bus.

Same as D, but the remote poller is the native binary instead of a container. The central bus TLS + auth setup (D, step 1) is unchanged.

On the remote host, build (or copy) the yagra-poller binary, drop the CA cert somewhere readable, and run:

Terminal window
sudo setcap cap_net_raw+ep ./yagra-poller
export YAGRA_BUS_URL="tls://poller:a-strong-poller-bus-password@core.example.com:4222"
export YAGRA_POLLER_ID="edge-tokyo-1" # unique per poller
export YAGRA_POLLER_POOL="tokyo"
export YAGRA_BUS_CA_FILE="/etc/yagra/certs/server-cert.pem"
# Optional intake listeners (see the privileged-port caveat in D):
# export YAGRA_SYSLOG_BIND="0.0.0.0:1514"
# export YAGRA_TRAP_BIND="0.0.0.0:1162"
# export YAGRA_FLOW_BIND="0.0.0.0:2055"
export RUST_LOG=info
./yagra-poller

Run it on the host network, not a private network namespace. Passive-event source-IP correlation and raw ICMP need the site’s real interfaces.

Everything else — pools, registration, failover, buffering — behaves exactly as in D.

Upgrades are designed to be low-effort, and to not lose or corrupt data. That is the design goal every migration and every message format is held to — not a promise that each release has been proven against.

So read the notes for the release you are moving to first, not only for behavior changes an operator or API client could notice, but because a release can withdraw the guarantee for itself and ask you to install fresh instead. When one does, it says so under Breaking changes and says what a fresh install costs. The changelog points at the notes for every version.

This is the ordinary way to upgrade. Settings ▸ Upgrade lists the releases this deployment can move to.

Pressing Upgrade starts nothing. It opens the list of what this deployment is made of — core, the co-located poller, and every remote site — each with the version it runs now, the version it would move to, and a checkbox, all ticked. Pressing Upgrade again starts the work on the rows you left ticked.

For every row that moves, the page does the whole thing — six steps, in this order:

  1. Check there is room, before anything is written
  2. Take a backup
  3. Pull the images
  4. Install the composition carried inside the target image
  5. Recreate the containers
  6. Verify that it came back up

Step 1 is what stops a full host from being made worse. The backup is a full PostgreSQL dump plus a VictoriaMetrics snapshot, and it lands on this host before any image is fetched — so without the check, a host that was already full got several hundred MB written onto it and then failed to pull. The run now stops with the two numbers in the message and the deployment untouched. The floor is YAGRA_UPGRADE_MIN_FREE_BYTES — 3 GiB by default.

Once the new version is up and healthy, the sidecar clears up after itself: the release you came from is kept, so the one-hop rollback still needs no download, and anything older goes, along with the older backups. Only Yagra’s own three images are ever touched, and one backup is always kept. YAGRA_UPGRADE_KEEP_RELEASES changes how many releases stay.

The work is carried out by a yagra-updater sidecar, the only container holding the Docker socket. Core never has it. Requesting an upgrade needs manage-the-deployment, which only an Admin holds; it is audited, and it is deliberately not on the MCP surface.

A switch on the same page turns the mechanism off. The setting lives in PostgreSQL, so it survives the upgrades it governs. That is unlike deleting the service from your compose file: each version installs its own composition, so a service deleted that way comes back on the next upgrade.

Step 4 replaces docker-compose.deploy.yml for the same reason, so an edit made to that file does not survive an upgrade either. Put the changes this particular host needs in a file called docker-compose.local.yml, beside it in the same directory. An upgrade never replaces that file, and from v0.3.11 it is handed to Compose as a second -f whenever it exists. Compose merges a later -f over an earlier one, so the overlay names only what it changes:

# docker-compose.local.yml — an upgrade never replaces this file
services:
poller:
networks: [default, lab-devices]
networks:
lab-devices:
external: true

Start the stack the way the upgrade does:

Terminal window
docker compose -p yagra -f docker-compose.deploy.yml -f docker-compose.local.yml up -d

A deployment without such a file is unaffected — no file, no argument. The upgrade says which of the two happened, with this deployment's docker-compose.local.yml or no docker-compose.local.yml here, so a deployment whose changes were not carried across reads differently from one that had none to carry. Settings ▸ Pollers ▸ “Accept remote pollers” passes the overlay too, because that switch recreates the poller by name.

Only deployment A has this. Every other way of installing Yagra — from source, natively, or from a composition without the yagra-updater sidecar — has no such mechanism. The page says so plainly instead of offering controls that would fail.

Pollers at remote sites come along too, and from v0.3.3 they do so by default.

A poller running in its own compose project at a monitored site is not part of the deployment’s project, so nothing central could restart it on its own.

The site bundle — Settings ▸ Pollers ▸ “Issue token & download” — writes COMPOSE_PROFILES=self-upgrade into the .env it generates, and the checkbox that does so is ticked. That profile starts a yagra-poller-updater sidecar at the site, and the site then declares itself able to install a release. Core hands it the release it just installed. There are rules about how:

  • It goes out one poller at a time per pool, so a pool with two or more pollers keeps monitoring throughout.
  • Images are fetched before anything stops, so a single-poller site is out for the container recreate rather than for the download.
  • The poller never touches the Docker socket. It validates the command and writes it into the shared volume for its own sidecar, exactly as core does centrally.

To keep a site out, untick the box before issuing its bundle, or empty COMPOSE_PROFILES in that site’s .env afterwards. Change it in .env rather than in the composition: an upgrade reinstalls the composition from the release being installed, and never touches .env. A site issued its bundle before v0.3.3 keeps its current behavior until it is handed a re-issued one.

Read what this grants before you roll it out. The sidecar runs as root with that site host’s Docker socket — the same trade the central one makes, and bounded the same way. See the upgrade sidecar.

A site that drifted can be brought back without upgrading core. Commands used to go out only as the tail of core’s own upgrade, so a site stood up afterwards — or one that failed and was fixed, or one that only enabled its updater today — stayed behind until the next release. From v0.3.4 you untick core in that list and press Upgrade: nothing in the deployment restarts, and each pool is still crossed one poller at a time. A poller ahead of core is moved back to it, because that is the skew direction the bus does not promise.

Before you press Upgrade, the list names every poller that will not come along and why, and every pool that has only one poller and so stops monitoring while its container is recreated. That warning follows your ticks, so unticking a row changes it. A release that would leave a poller two versions behind is shown with the reason and no checkbox, since the bus supports one release of skew.

The poller half of an upgrade has its own progress track. Core’s own steps end at “Verifying”; what follows — each site fetching the image over its own link, then one container recreated at a time per pool — can take another half hour. Each site now shows where it has got to, and the record stays on the page after the run, so a site that did not come back is named instead of quietly dropping off the live poller list.

Going back is a supported move within a declared window. A migration can declare a compatibility floor: the oldest version that can still run once that migration has been applied. A release below the current floor is shown greyed out with the reason rather than hidden.

Nothing is deleted by going back. Columns the newer version added stay in place, unread, and become visible again on the next upgrade.

There is a path for a site with no reachable registry at all. Run docker save on the three release images where you can reach them, and upload the archive at the same page.

That path needs a second, deliberate opt-in on the host (YAGRA_UPGRADE_ALLOW_BUNDLE=1), because docker load installs whatever the archive contains. Read the configuration reference before turning it on.

Upgrading from the command line is still supported. Use it when the sidecar is switched off or cannot run.

All you are doing is changing the image tag to the new version and recreating the containers.

  1. Take a backup.

    Always take one before a major upgrade. There are two things to save.

    Terminal window
    # PostgreSQL — nodes, configuration, users, alert history
    docker compose -f docker-compose.deploy.yml exec postgres \
    pg_dump -U yagra yagra > yagra-backup.sql
    # The KEK — tiny, and irreplaceable
    docker run --rm -v yagra_kekdata:/kek busybox cat /kek/key > kek-backup.key

    Metrics live in the vmdata volume. Snapshot it with your usual volume tooling, or with VictoriaMetrics’ own snapshot API. Redis needs no backup — it is rebuildable, so losing it is non-fatal.

  2. Choose the release you are moving to.

    Skim the changelog. Anything whose behavior changed is listed there.

  3. Pull the images at the new tag.

    Terminal window
    # in .env: YAGRA_IMAGE_TAG=v0.2.2
    # or override it inline, as below
    YAGRA_IMAGE_TAG=v0.2.2 docker compose -f docker-compose.deploy.yml pull
  4. Recreate the containers on those images.

    Terminal window
    YAGRA_IMAGE_TAG=v0.2.2 docker compose -f docker-compose.deploy.yml up -d

    Schema changes run automatically when core starts. There is nothing to apply by hand.

Why this cannot corrupt your data

  • Schema changes add first and remove later, in two stages. Stopping partway loses nothing, and N→N+1 is always supported. To see what a release will do before it does it, run yagra-core migrations inside the target image. It prints the migration set that binary embeds as JSON, with no database and no configuration, so an upgrade can be planned before anything is touched.
  • Core and pollers can talk across one version of skew. A new core works with old pollers, so upgrade core first and pollers after — including remote sites, one at a time, in any order.
  • Pollers are stateless. Replace them freely. A pool briefly without pollers falls back to publishing jobs one at a time, so no node goes dark mid-rollout.
  • Persistent data is preserved. The pgdata (PostgreSQL), vmdata (VictoriaMetrics), and kekdata (KEK) volumes — or their native equivalents — survive image upgrades untouched.

To go back, re-run the same steps with the previous tag. The immutable :<git-sha> tags exist for exactly this.

Yagra installs nothing on the host — no packages, no system services, no files outside the deployment directory and Docker’s own storage. Removing it is the install run backwards.

Run everything below from the directory holding docker-compose.deploy.yml.

Stop it, keep the data. Containers and the network go; every named volume stays, so up -d brings the deployment back with its configuration, alert history and metrics intact.

Terminal window
docker compose -p yagra -f docker-compose.deploy.yml down

Remove it completely. -v destroys the eleven named volumes — pgdata (nodes, users, thresholds, alert history, every setting), vmdata (all metric history), vldata (all passive events), chdata (all flow records), kekdata (the key-encryption key), and the certificate, log, buffer and hand-off volumes.

Terminal window
docker compose -p yagra -f docker-compose.deploy.yml down -v --remove-orphans

To take the images and the directory too:

Terminal window
docker compose -p yagra -f docker-compose.deploy.yml down -v --rmi all --remove-orphans
cd .. && rm -rf yagra # the compose file and .env, which holds POSTGRES_PASSWORD

--rmi all also removes the backing-store images (postgres, redis, nats, victoria-metrics, victoria-logs, clickhouse); Docker skips any another project is still using.

Remote pollers are separate stacks. A poller at a remote site (deployment D) is its own Compose project on its own host. Nothing above touches it, and it will retry the bus forever after the central stack is gone. Remove each on its own host with docker compose -f docker-compose.poller.yml down -v.

Checking for leftovers. Compose labels everything it created, so nothing depends on still having the compose file:

Terminal window
docker ps -a --filter label=com.docker.compose.project=yagra
docker volume ls --filter label=com.docker.compose.project=yagra
docker network ls --filter label=com.docker.compose.project=yagra

There is no uninstall action in the WebUI — Settings ▸ Upgrade only moves a deployment between releases. Uninstalling is deliberately a host-side act.