Compare commits

...
59 Commits
Author SHA1 Message Date
dannygroenewegen 382f6353fc fix: strip monitoring. prefix from instance labels for cleaner dashboard views 2026-08-20 16:50:25 +02:00
dannygroenewegen c8f74e2a3c fix: refresh dashboard variables on time range change, not just page load 2026-08-20 14:07:10 +02:00
dannygroenewegen 6dbb6c5f55 fix: mount containerd socket read-only, avoid abra#900 2026-08-17 16:09:44 +02:00
dannygroenewegen 17c4f6237b docs: rewrite README and release notes for the Alloy migration 2026-08-17 15:58:04 +02:00
dannygroenewegen 50fc916107 feat: add healthchecks to Alloy, Prometheus and Pushgateway 2026-08-17 15:08:08 +02:00
dannygroenewegenandClaude Sonnet 5 7480a8d6ab fix abra linting secret length
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-17 14:59:03 +02:00
dannygroenewegen 067f013fb7 fix: don't expose Alloy's web UI by default, but optionally with compose.alloy-webui.yml 2026-08-17 14:48:44 +02:00
dannygroenewegen 536fa7f37d fix Traefik dashboard 2026-08-17 13:37:27 +02:00
dannygroenewegen fafd1f89f0 fix: discover containers by task instead of service for scraping
Swarm service discovery misses dnsrr-mode services like traefik; switch
to swarm task discovery, scoped to this host, with a stable instance
label across redeploys and a default port when none is labeled.
2026-08-10 21:24:38 +02:00
dannygroenewegen 7ab6cc74be fix: make swarm/stacks dashboards work with old and new metric data
Replace node_meta-based joins with label_replace() of the instance
label, so old (pull-model) and new (Alloy push) series for the same
host render as one continuous series. Each rewritten query has a TODO
noting the simpler form to use once old data ages out of retention.
2026-08-07 16:57:39 +02:00
dannygroenewegen 0f989c4567 fix: drop unfiltered discovery.docker.containers.targets from default scrape
prometheus.exporter.cadvisor.docker.targets already reports resource
metrics for every container on the host. The unfiltered discovery just
added targets with internal overlay IPs, most of which failed to
scrape anything useful.
2026-08-07 15:52:31 +02:00
dannygroenewegen 402e022114 feat: separate scrape_auth secret, configurable usernames
Added an optional scrape_auth secret for authenticating scrapes of
containers that opt in via prometheus.io/auth labels, no longer
reusing the basic_auth secret meant for Prometheus/Loki writes.

Made usernames for both write endpoints and authenticated scraping
configurable in .env.
2026-08-07 15:52:19 +02:00
dannygroenewegen 7f3d98fd9f - Restore GRAFANA_DOMAIN
- Start documenting the migration in release/next
- Various cleanups
2026-08-07 14:47:22 +02:00
fauno 627d988902 fix: upgrade to grafana 13.0.6 2026-08-07 09:24:53 -03:00
fauno e75fa4487c fix: upgrade alloy to 1.18.1 2026-08-06 12:37:25 -03:00
fauno 5f3f9c957c fix: upgrade to grafana 13.0.5 2026-08-04 10:48:47 -03:00
fauno 7818c89ad7 fix: provide access to containerd socket toolshed/abra#900 2026-07-31 13:57:43 -03:00
fauno b47f1b021f fix: upgrade grafana to v13.0.3 2026-06-27 14:43:43 -03:00
fauno c6efc69859 fix: upgrade alloy to v1.17.0 2026-06-27 14:42:11 -03:00
fauno 6c529f3527 fix: upgrade to grafana v12.4.5 2026-06-27 14:37:58 -03:00
fauno 01b398ddea fix: unneeded flag 2026-06-27 14:24:59 -03:00
fauno 30ad8e54f6 fix: upgrade to alloy v1.16.3 2026-06-27 14:20:32 -03:00
fauno c9910eabf4 Merge branch 'main' of https://git.coopcloud.tech/coop-cloud/monitoring-ng into alloy 2026-06-20 23:59:20 -03:00
fauno 1d9eb10004 improve alloy config and match with main branch (#24)
Reviewed-on: #24
Reviewed-by: fauno <fauno@sutty.coop.ar>
2026-06-17 12:53:08 +00:00
fauno 23acf56637 fix: filter by proxy network 2026-06-16 22:01:50 -03:00
fauno 03227f1907 fix: send container metrics directly to prometheus 2026-06-16 22:01:05 -03:00
fauno d085c66d68 fix: needed labels come from docker swarm 2026-06-16 22:00:42 -03:00
fauno 1970061ff8 feat: live debugging alloy 2026-06-16 21:59:50 -03:00
fauno fa76179987 feat: enable alloy web ui 2026-06-16 21:20:29 -03:00
fauno 64cb07a4a2 feat: bearer auth support 2026-06-16 20:03:17 -03:00
dannygroenewegen e247677433 feat: scrape metrics from containers via Docker label discovery
Containers opt in with prometheus.io/scrape=true and optionally set
prometheus.io/port, prometheus.io/path, and prometheus.io/auth=basic.
2026-06-14 21:38:43 +02:00
dannygroenewegen f2310f2b86 improve alloy config and match with main branch
- Restrict Alloy UI to loopback
- Narrow volume mounts: drop /dev, reduce /var/run to docker.sock:ro
- Replace HTTP scrape of :12345 with prometheus.exporter.self
Match with main branch (node-exporter / promtail / cadvisor):
- Add docker_only and explicit enabled_metrics to cadvisor exporter
- Match node-exporter collector config
- Match promtail relabeling (container_name, container_id, stack_namespace,
  service_name) and external hostname label
- Add SYSLOG_FILES option to tail /var/log/*log (matches promtail)
- Fix journal path and syslog listener address
2026-06-13 22:07:55 +02:00
fauno f2711fa16e fix: upgrades 2026-06-03 00:19:41 -03:00
fauno 2870b9486c fix: use the actual health check path 2026-06-02 21:06:05 -03:00
fauno 3a1fabe4f9 fix: prevent redirections on health check 2026-06-02 21:04:56 -03:00
fauno a358837922 wip: relabel syslog according to docs 2026-06-02 21:00:34 -03:00
fauno dd0a0c1bb0 fixup! feat: read syslog 2026-06-02 20:19:45 -03:00
fauno 31cabc36ae fix: prevent traefik deprecation warnings 2026-06-02 19:16:49 -03:00
fauno d25986d5cb fix: README 2026-06-02 18:51:10 -03:00
fauno f8f8004445 feat: read syslog 2026-06-02 18:50:41 -03:00
fauno aa05d022da feat: optionally push to prometheus and loki 2026-06-02 18:50:20 -03:00
fauno fb52a76247 BREAKING CHANGE: deprecate node-exporter 2026-06-02 18:49:05 -03:00
fauno 2e2a52eae0 BREAKING CHANGE: deprecate promtail 2026-06-02 18:48:20 -03:00
fauno 48419d5afa fixup! BREAKING CHANGE: no need to expose exporters 2026-06-02 18:46:02 -03:00
fauno a0a6e2c509 fix: basic auth secret is always needed 2026-06-02 18:44:32 -03:00
fauno 024f2a8aec feat: send docker logs to loki 2026-06-02 18:39:24 -03:00
fauno 38095e23fa BREAKING CHANGE: no need to expose exporters 2026-06-02 18:37:56 -03:00
fauno 641161329e fix: grafana alternate domain doesn't work
the variable is not expanded and the domain name label ends up as a
literal "$DOMAIN".
2026-06-02 18:00:00 -03:00
fauno cdacfd035e fix: prometheus querying panel is accessible through basic auth 2026-06-02 17:52:25 -03:00
fauno b2d3901f61 fix: bind mounts recommended by docs 2026-06-02 13:24:28 -03:00
fauno 8becf1c1d6 fixup! feat: node exporter 2026-05-29 16:16:37 -03:00
fauno 777b1355dd fixup! feat: node exporter 2026-05-29 16:16:08 -03:00
fauno e83433cebd feat: node exporter 2026-05-29 16:04:19 -03:00
fauno a713f98ffb feat: instance name is domain 2026-05-29 16:03:59 -03:00
fauno 8dc84c591c fixup! feat: enable prometheus remote write receiver 2026-05-29 15:38:52 -03:00
fauno d9aa05a4b5 feat: send metrics to prometheus 2026-05-28 21:00:10 -03:00
fauno 349df12204 feat: enable prometheus remote write receiver 2026-05-28 20:44:00 -03:00
fauno 6c33089078 feat: cadvisor 2026-05-28 20:38:50 -03:00
fauno 4bedebfab1 BREAKING CHANGES: replace promtail and cadvisor for alloy 2026-05-28 20:33:36 -03:00
24 changed files with 737 additions and 407 deletions
+45 -10
View File
@@ -5,16 +5,51 @@ DOMAIN=monitoring-ng.example.com
#TIMEOUT=120
ENABLE_BACKUPS=true
## Enable this secret for Promtail / Prometheus
#COMPOSE_FILE="$COMPOSE_FILE:compose.basic-auth.yml"
#SECRET_BASIC_AUTH_VERSION=v1
#
# Promtail (Gathering Logs)
# COMPOSE_FILE="$COMPOSE_FILE:compose.promtail.yml"
# LOKI_PUSH_URL=https://loki.monitoring.example.org/loki/api/v1/push
# Secret Alloy authenticates with when writing metrics/logs to Prometheus/
# Loki and what Traefik's basicauth middleware expects for that
SECRET_BASIC_AUTH_VERSION=v1
# Username sent along with SECRET_BASIC_AUTH_VERSION above (default: admin)
# WRITE_BASIC_AUTH_USERNAME=admin
## Expose node and cadvisor ports instead of traefik
# COMPOSE_FILE="$COMPOSE_FILE:compose.expose-ports.yml"
# Expose Alloy's web UI publicly (behind basic-auth) at alloy.$DOMAIN.
# COMPOSE_FILE="$COMPOSE_FILE:compose.alloy-webui.yml"
# Enable Live Debugging in web ui
# LIVE_DEBUGGING=false
# Enable this to send metrics to a Prometheus server, adapt DOMAIN if
# server is remote
# PROMETHEUS_REMOTE_WRITE_URL=https://prometheus.$DOMAIN/api/v1/write
# Enable authenticated scraping of containers that opt in via
# prometheus.io/auth=basic or prometheus.io/auth=bearer labels (used as
# password/bearer token respectively). Insert it with:
# abra app secret insert <domain> scrape_auth v1 <password>
# COMPOSE_FILE="$COMPOSE_FILE:compose.scrape-auth.yml"
# SECRET_SCRAPE_AUTH_VERSION=v1
#
# Username sent along with it (default: alloy)
# SCRAPE_BASIC_AUTH_USERNAME=alloy
# Edit this if your distribution sets this socket to another location
# CONTAINERD_SOCKET=/var/run/containerd/containerd.sock
# Enable this to send logs to a Loki server, adapt DOMAIN if server is
# remote
# LOKI_PUSH_URL=https://loki.$DOMAIN/loki/api/v1/push
# Enable on systemd hosts to read logs from the journal
# JOURNALD=1
#
# Enable on non-systemd hosts (Alpine, older Debian/Ubuntu) to tail
# /var/log/*log files (syslog, auth.log, kern.log, etc.) that a local
# syslogd writes. No syslogd reconfiguration needed.
# SYSLOG_FILES=1
#
# Enable to receive syslog messages over the network on port 514/tcp.
# Use for remote devices that push syslog to this host, or for a
# local syslogd configured to forward over the network.
# Not needed if you just want to read local log files — use SYSLOG_FILES instead.
# SYSLOG=1
# COMPOSE_FILE="$COMPOSE_FILE:compose.syslog.yml"
# Monitoring Server
#
@@ -70,7 +105,7 @@ ENABLE_BACKUPS=true
# GF_SMTP_ENABLED=true
# GF_SMTP_FROM_ADDRESS=grafana@example.com
# GF_SMTP_SKIP_VERIFY=false
# SECRET_GF_SMTP_PASSWD_VERSION=v1
# SECRET_GF_SMTP_PASS_VERSION=v1
#
## Grafana Matrix Contact Point (optional)
+83 -62
View File
@@ -1,9 +1,10 @@
# monitoring-ng
Yet another monitoring stack ...
This time its a all-in-one grafana/prometheus/loki/node_exporter/cadvisor/promtail stack.
It's based heavily on the [monitoring-lite](https://git.coopcloud.tech/coop-cloud/monitoring-lite) stack, but has everything in one recipe included now. So you can deploy monitoring instances to only gather metrics / logs (node_exporter/cadvisor/promtail) and also deploy instances with the full monitoring stack (grafana/prometheus/loki) with the same recipe and just different .env configuration.
This time its a all-in-one grafana/prometheus/loki/alloy stack.
It's based heavily on the [monitoring-lite](https://git.coopcloud.tech/coop-cloud/monitoring-lite) stack, but has everything in one recipe included now. So you can deploy monitoring instances to only gather metrics / logs (alloy) and also deploy instances with the full monitoring stack (grafana/prometheus/loki) with the same recipe and just different .env configuration.
Metrics and logs are collected by [Grafana Alloy](https://grafana.com/docs/alloy/latest/) and pushed to a central Prometheus/Loki (via `remote_write`/`loki push`). Every `monitoring-ng` instance runs its own Alloy, whether or not it also runs the central stack.
<!-- metadata -->
@@ -18,62 +19,93 @@ It's based heavily on the [monitoring-lite](https://git.coopcloud.tech/coop-clou
<!-- endmetadata -->
## Setup Metrics Gathering
## Setup: gathering-only host
Where gathering.org is the node you want to gather metrics from.
Deploys just Alloy, pushing this host's own node/container metrics and logs to a central instance. Use this on every host you want metrics/logs from.
1. Configure DNS
- cadvisor.monitoring.gathering.org
- node.monitoring.gathering.org
2. [Configure Traefik to use BasicAuth](https://git.coopcloud.tech/coop-cloud/traefik#configuring-wildcard-ssl-using-dns)
3. `abra app new monitoring-ng`
4. `abra app config monitoring.gathering.org` (for gathering only the main `compose.yml` is needed, nothing more.)
1. `abra app new monitoring-ng --server gathering.org`
2. `abra app config monitoring.gathering.org`
3. Point it at your central instance:
```
PROMETHEUS_REMOTE_WRITE_URL=https://prometheus.example.org/api/v1/write
LOKI_PUSH_URL=https://loki.example.org/loki/api/v1/push
JOURNALD=1 # or SYSLOG_FILES=1 / SYSLOG=1, see .env.sample
```
4. `abra app secret insert monitoring.gathering.org basic_auth v1 <password>`. Same username/password as the `usersfile` credential configured for Traefik's basicauth on the central node (see below). This is what Alloy authenticates with when pushing metrics/logs. Alloy defaults to username `admin` for this. Uf the Traefik `usersfile` uses a different username, set `WRITE_BASIC_AUTH_USERNAME` in this recipe to match.
5. `abra app deploy monitoring.gathering.org`
6. check that endpoints are up and basic-auth works
- cadvisor.monitoring.gathering.org
- node.monitoring.gathering.org
### Expose node and cadvisor via ports instead of traefik
## Setup: full monitoring stack (metrics/logs browser)
In case you have no traefik running on the machine, you can expose the ports directly by uncommenting the following line:
```
# COMPOSE_FILE="$COMPOSE_FILE:compose.expose-ports.yml"
This is what a gathering host pushes into. It also runs its own Alloy, so it monitors itself too.
1. Configure DNS: `monitoring.example.org`, plus `prometheus.`/`loki.`/`pushgateway.` subdomains for whichever of those you enable below
2. Traefik on this node needs basic auth configured (`BASIC_AUTH=1`, see the Traefik recipe's "Configuring basic auth" section) — Prometheus/Loki/Pushgateway route through its `basicauth@file` middleware, so without it those endpoints won't work. Use the same username (default `admin`, see `WRITE_BASIC_AUTH_USERNAME`) and password you'll insert as the `basic_auth` secret below when generating Traefik's `usersfile`.
3. `abra app config monitoring.example.org` Uncomment `compose.prometheus.yml` (metrics), `compose.loki.yml` (logs) and `compose.grafana.yml` (dashboard)
4. `abra app secret insert monitoring.example.org basic_auth v1 <password>` — this is the password every gathering host's Alloy (including this instance's own) authenticates with; also what Traefik's basicauth expects on the public Prometheus/Loki/Pushgateway endpoints
5. `abra app secret insert monitoring.example.org gf_adminpasswd v1 <password>`
6. `abra app deploy monitoring.example.org`
### Post-setup guide
- configure the SMTP mailer under `Alerting > Contact points`
- edit the default contact point, choose "Alertmanager" as type & `http://alertmanager:9093` as URL
- use the "Test" button to send a test mail. It should fire a request at the alertmanager & that should send a mail
- from your dashboard panels, choose `Edit > Alert` to create alerts based on those panels
## Additional features
### Discovering metrics from other apps
Alloy auto-discovers and scrapes other Docker Swarm services running on the same host, on the `proxy` network, that opt in via labels. No manual scrape config needed. On the app's `compose.yml`:
```yaml
deploy:
labels:
- "prometheus.io/scrape=true" # required: opt in
# - "prometheus.io/port=8082" # optional: defaults to 80
# - "prometheus.io/path=/metrics" # optional: defaults to /metrics
# - "prometheus.io/auth=basic" # optional: basic auth, see below
# - "prometheus.io/auth=bearer" # optional: bearer token auth
```
## Setup Metrics Browser
Each scraped target gets `instance` (`<service_name>.<slot>`, stable across redeploys), `domain` (the service's stack namespace with underscores converted back to dots, e.g. `traefik.example.com`) and `task_slot` labels attached automatically.
This builds upon [Setup Metrics Gathering](#setup-metrics-grathering) so make sure you did that first.
If the target needs authentication, configure the monitoring-ng instance with a scrape-auth secret that will be used for targets having the auth label set:
```
COMPOSE_FILE="$COMPOSE_FILE:compose.scrape-auth.yml"
SECRET_SCRAPE_AUTH_VERSION=v1
```
`abra app secret insert <domain> scrape_auth v1 <password-or-token>`, then set the scraped app's `prometheus.io/auth` label to `basic` or `bearer` to match how it checks the secret.
1. Configure DNS
- monitoring.example.org
2. Setup monitoring stack
- `abra app config monitoring.example.org` Uncomment prometheus, loki and grafana
- `abra app secret insert monitoring.example.org basic_auth v1 <password>`
this needs the plaintext traefik basic-auth secret, not the hashed one!
- `abra app secret ls monitoring.example.org`
- `abra app deploy monitoring.example.org`
3. Add scrape config to prometheus
- `abra app cmd monitoring.example.org prometheus gathering.org`
- or manually
```
cp scrape-config.example.yml gathering.org.yml
# adjust domain
# mkdir scrape_configs
abra app cp monitoring.dev.local-it.cloud gathering.org.yml prometheus:/prometheus/scrape_configs/
```
Check discovered targets via `alloy.example.org` (needs `compose.alloy-webui.yml`, see below), or query the central Prometheus for `up{job="<service_name>"}`.
* check that all configured targets are up:
https://prometheus.monitoring.example.org/targets
### Manual scraping
For targets where auto-discovery doesn't work (e.g. not a Docker Swarm service on the `proxy` network, or missing labels): add them directly to Prometheus instead. Create a scrape config file:
```yaml
- targets:
- 'metrics.something-external.example.org'
- 'app-without-labels.example.org'
```
and copy it into Prometheus' scrape config directory:
```
abra app cp monitoring.gathering.org targets.yml prometheus:/prometheus/scrape_configs/
```
Prometheus picks up files there automatically.
### Alloy Web UI / Live Debugging
| Service | Authentication | Domain |
| ------------- | ------------------ | --------------------------------- |
| Grafana | Email / SSO | monitoring.example.org |
| Prometheus | traefik basic-auth | prometheus.monitoring.example.org |
| loki | traefik basic-auth | loki.monitoring.example.org |
| Cadvisor | traefik basic-auth | cadvisor.monitoring.example.org |
| Node Exporter | traefik basic-auth | node.monitoring.example.org |
Alloy's own web UI isn't exposed by default, its HTTP server only listens
on localhost inside its own container. To reach it from outside (e.g. to
browse its component graph or use live debugging), add:
```
COMPOSE_FILE="$COMPOSE_FILE:compose.alloy-webui.yml"
```
This exposes it (behind the same basic-auth) at `alloy.$DOMAIN`.
To actually see what's being collected (live-tailing the metrics/logs
flowing through each component, not just their config) also set
`LIVE_DEBUGGING=true`. Only enable this while troubleshooting.
### Logging from a docker host to loki server without anything else
@@ -90,36 +122,25 @@ $ echo '{
$ systemctl restart docker.service
```
## Setup Push Gateway
### Setup Push Gateway
1. Enable in the env fiöle by uncommenting the following lines:
1. Enable in the env file by uncommenting the following lines:
```
## Prometheus Pushgateway
# COMPOSE_FILE="$COMPOSE_FILE:compose.pushgateway.yml"
```
2. `abra app deploy monitoring.example.org`
This will expose the pushgateway at `https://pushgateway.${DOMAIN}`.
It is secured behind the same basic auth as the other services.
After that you need to add the `pushgateway.${DOMAIN}` to the scare config.
This will expose the pushgateway at `https://pushgateway.${DOMAIN}`, secured behind the same basic auth as the other services.
After that you need to add the `pushgateway.${DOMAIN}` to the scrape config of Prometheus.
## Post-setup guide
- configure prometheus/loki/alertmanager as data sources in grafana under `Configuration > Data sources`
- for loki, you need to set a "Custom HTTP Header": `X-Scope-OrgID: fake`
- configure the SMTP mailer under `Alerting > Contact points`
- edit the default contact point, choose "Alertmanager" as type & `http://alertmanager:9093` as URL
- use the "Test" button to send a test mail. It should fire a request at the alertmanager & that should send a mail
- `abra app cp` your `scrap_configs: ...` into `/prometheus/scrape_configs` & log into your prometheus web UI to ensure they're working
- load your dashboards in manually under `Create > Dashboard`
- from your dashboard panels, choose `Edit > Alert` to create alerts based on those panels
---
THX to the previous work of @decentral1se @knooflok @3wc @cellarspoon @mirsal
## Adding Matrix as Alert Contact point
### Adding Matrix as Alert Contact point
1. Enable the [matrix-alertmanager-receiver](https://github.com/metio/matrix-alertmanager-receiver/):
```
@@ -139,7 +160,7 @@ GF_MATRIX_HOME_SERVER_URL=
```
4. Configure Alertmanager webhook and set the url to `http://matrix-alertmanager-receiver:12345/alerts/<room-id>`
## Alerts
### Alerts
It is possible to enable the following alerts, by uncommenting the corresponding env variable:
+4 -29
View File
@@ -1,26 +1,16 @@
export ENTRYPOINT_VERSION=v1
export GF_DATASOURCES_VERSION=v1
export GF_DASHBOARDS_VERSION=v2
export GF_SWARM_DASH_VERSION=v2
export GF_STACKS_DASH_VERSION=v2
export GF_TRAEFIK_DASH_VERSION=v2
export GF_TRAEFIK_DASH_VERSION=v3
export GF_BACKUP_DASH_VERSION=v1
export GF_CUSTOM_INI_VERSION=v4
export PROMTAIL_YML_VERSION=v3
export LOKI_YML_VERSION=v3
export PROMETHEUS_YML_VERSION=v2
export MATRIX_ALERTMANAGER_CONFIG_VERSION=v1
export MATRIX_ALERTMANAGER_ENTRYPOINT_VERSION=v1
export GRAFANA_ALERTS_NODE_VERSION=v2
# creates a default prometheus scrape config for a given node
add_node(){
name=$1
add_domain "$name" "metrics.traefik.$name"
add_domain "$name" "node.monitoring.$name"
add_domain "$name" "cadvisor.monitoring.$name"
cat "/prometheus/scrape_configs/$name.yml"
}
export GF_ALERTS_NODE_VERSION=v2
export CONFIG_ALLOY_VERSION=v11
# migrates secrets from old names to new names by reading values from the
# running containers on the server and re-inserting them under the new names.
@@ -37,7 +27,7 @@ migrate_secret_names() {
# Hardcoded migration mappings: old_secret_name|new_secret_name
MIGRATIONS="
grafana_admin_password|gf_adminpasswd
grafana_smtp_password|gf_smtp_passwd
grafana_smtp_password|gf_smtp_pass
grafana_oidc_client_secret|gf_oidc_secret
matrix_access_token|matrix_token
loki_aws_secret_access_key|loki_aws_key
@@ -116,18 +106,3 @@ loki_aws_secret_access_key|loki_aws_key
echo ""
echo "Done."
}
# adds a domain to a scrape config or creates a new one
add_domain(){
name=$1
domain=$2
if [ ! -d "/prometheus/scrape_configs/" ]; then
mkdir -p /prometheus/scrape_configs/
fi
cd /prometheus/scrape_configs/ || exit 1
if [ ! -f "$name.yml" ]; then
echo -e "- targets:\n - '$domain'" > "$name.yml"
else
echo " - '$domain'" >> "$name.yml"
fi
}
+1 -1
View File
@@ -29,7 +29,7 @@ groups:
datasourceUid: PBFA97CFB590B2093
model:
editorMode: code
expr: (node_filesystem_free_bytes{fstype="ext4"} / node_filesystem_size_bytes{fstype="ext4"}) * 100
expr: (node_filesystem_free_bytes{fstype=~"ext4|xfs"} / node_filesystem_size_bytes{fstype=~"ext4|xfs"}) * 100
instant: true
intervalMs: 1000
legendFormat: __auto
+16
View File
@@ -0,0 +1,16 @@
version: "3.8"
services:
app:
environment:
- ALLOY_HTTP_LISTEN_ADDR=0.0.0.0
deploy:
labels:
- "traefik.enable=true"
- "traefik.swarm.network=proxy"
- "traefik.http.services.${STACK_NAME}-alloy.loadbalancer.server.port=12345"
- "traefik.http.routers.${STACK_NAME}-alloy.rule=Host(`alloy.${DOMAIN}`)"
- "traefik.http.routers.${STACK_NAME}-alloy.entrypoints=web-secure"
- "traefik.http.routers.${STACK_NAME}-alloy.tls=true"
- "traefik.http.routers.${STACK_NAME}-alloy.tls.certresolver=${LETS_ENCRYPT_ENV}"
- "traefik.http.routers.${STACK_NAME}-alloy.middlewares=basicauth@file"
-7
View File
@@ -1,7 +0,0 @@
---
version: "3.8"
secrets:
basic_auth:
external: true
name: ${STACK_NAME}_basic_auth_${SECRET_BASIC_AUTH_VERSION}
-13
View File
@@ -1,13 +0,0 @@
---
version: "3.8"
services:
app:
ports:
- "9100:9100"
deploy:
cadvisor:
ports:
- "9101:8080"
deploy:
+4 -4
View File
@@ -3,16 +3,16 @@ version: '3.8'
services:
grafana:
secrets:
- gf_smtp_passwd
- gf_smtp_pass
environment:
- GF_SMTP_HOST
- GF_SMTP_USER
- GF_SMTP_PASSWORD__FILE=/run/secrets/gf_smtp_passwd
- GF_SMTP_PASSWORD__FILE=/run/secrets/gf_smtp_pass
- GF_SMTP_ENABLED
- GF_SMTP_FROM_ADDRESS
- GF_SMTP_SKIP_VERIFY
secrets:
gf_smtp_passwd:
gf_smtp_pass:
external: true
name: ${STACK_NAME}_gf_smtp_passwd_${SECRET_GF_SMTP_PASSWD_VERSION}
name: ${STACK_NAME}_gf_smtp_pass_${SECRET_GF_SMTP_PASS_VERSION}
+4 -4
View File
@@ -2,7 +2,7 @@ version: '3.8'
services:
grafana:
image: grafana/grafana:12.4.0
image: grafana/grafana:13.0.6
volumes:
- grafana-data:/var/lib/grafana:rw
secrets:
@@ -37,14 +37,14 @@ services:
deploy:
labels:
- "traefik.enable=true"
- "traefik.docker.network=proxy"
- "traefik.swarm.network=proxy"
- "traefik.http.services.${STACK_NAME}-grafana.loadbalancer.server.port=3000"
- "traefik.http.routers.${STACK_NAME}-grafana.rule=Host(`${GRAFANA_DOMAIN:-$DOMAIN}`)"
- "traefik.http.routers.${STACK_NAME}-grafana.entrypoints=web-secure"
- "traefik.http.routers.${STACK_NAME}-grafana.tls=true"
- "traefik.http.routers.${STACK_NAME}-grafana.tls.certresolver=${LETS_ENCRYPT_ENV}"
healthcheck:
test: "wget -q http://localhost:3000/ -O/dev/null"
test: "wget -q http://localhost:3000/healthz -O/dev/null"
interval: 5s
timeout: 10s
retries: 3
@@ -75,7 +75,7 @@ configs:
file: grafana-backup-dashboard.json
gf_alerts_node:
template_driver: golang
name: ${STACK_NAME}_gf_alerts_node_${GRAFANA_ALERTS_NODE_VERSION}
name: ${STACK_NAME}_gf_alerts_node_${GF_ALERTS_NODE_VERSION}
file: alerts/node.yml.tmpl
volumes:
+2 -2
View File
@@ -2,7 +2,7 @@ version: '3.8'
services:
loki:
image: grafana/loki:3.6.7
image: grafana/loki:3.7.2
command: -config.file=/etc/loki/local-config.yaml
networks:
- proxy
@@ -27,7 +27,7 @@ services:
condition: on-failure
labels:
- "traefik.enable=true"
- "traefik.docker.network=proxy"
- "traefik.swarm.network=proxy"
- "traefik.http.services.${STACK_NAME}-loki.loadbalancer.server.port=3100"
- "traefik.http.routers.${STACK_NAME}-loki.rule=Host(`loki.${DOMAIN}`)"
- "traefik.http.routers.${STACK_NAME}-loki.entrypoints=web-secure"
+9 -2
View File
@@ -2,7 +2,7 @@ version: '3.8'
services:
prometheus:
image: prom/prometheus:v3.10.0
image: prom/prometheus:v3.12.0
secrets:
- basic_auth
volumes:
@@ -16,15 +16,22 @@ services:
- "--web.console.libraries=/usr/share/prometheus/console_libraries"
- "--web.console.templates=/usr/share/prometheus/consoles"
- "--storage.tsdb.retention.time=${PROMETHEUS_RETENTION_TIME}"
- "--web.enable-remote-write-receiver"
networks:
- proxy
- internal
healthcheck:
test: "wget -q --spider http://localhost:9090/-/healthy || exit 1"
interval: 5s
timeout: 10s
retries: 3
start_period: 30s
deploy:
restart_policy:
condition: on-failure
labels:
- "traefik.enable=true"
- "traefik.docker.network=proxy"
- "traefik.swarm.network=proxy"
- "traefik.http.services.${STACK_NAME}-prometheus.loadbalancer.server.port=9090"
- "traefik.http.routers.${STACK_NAME}-prometheus.rule=Host(`prometheus.${DOMAIN}`)"
- "traefik.http.routers.${STACK_NAME}-prometheus.entrypoints=web-secure"
-25
View File
@@ -1,25 +0,0 @@
version: "3.8"
services:
promtail:
image: grafana/promtail:3.6.7
volumes:
- /var/log:/var/log:ro
- /var/run/docker.sock:/var/run/docker.sock
command: -config.file=/etc/promtail/config.yml
configs:
- source: promtail_yml
target: /etc/promtail/config.yml
networks:
- internal
secrets:
- basic_auth
environment:
- DOMAIN
- LOKI_PUSH_URL
configs:
promtail_yml:
name: ${STACK_NAME}_promtail_yml_${PROMTAIL_YML_VERSION}
file: promtail.yml.tmpl
template_driver: golang
+7 -1
View File
@@ -12,12 +12,18 @@ services:
networks:
- internal
- proxy
healthcheck:
test: "wget -q --spider http://localhost:9191/-/healthy || exit 1"
interval: 5s
timeout: 10s
retries: 3
start_period: 10s
deploy:
restart_policy:
condition: on-failure
labels:
- "traefik.enable=true"
- "traefik.docker.network=proxy"
- "traefik.swarm.network=proxy"
- "traefik.http.services.${STACK_NAME}-pushgateway.loadbalancer.server.port=9191"
- "traefik.http.routers.${STACK_NAME}-pushgateway.rule=Host(`pushgateway.${DOMAIN}`)"
- "traefik.http.routers.${STACK_NAME}-pushgateway.entrypoints=web-secure"
+13
View File
@@ -0,0 +1,13 @@
---
version: "3.8"
services:
app:
secrets:
- source: scrape_auth
target: scrape_auth
secrets:
scrape_auth:
external: true
name: ${STACK_NAME}_scrape_auth_${SECRET_SCRAPE_AUTH_VERSION}
+6
View File
@@ -0,0 +1,6 @@
---
version: "3.8"
services:
app:
ports:
- "514:514"
+52 -72
View File
@@ -3,89 +3,69 @@ version: "3.8"
services:
app:
image: prom/node-exporter:v1.10.2
user: root
environment:
- NODE_ID={{.Node.ID}}
volumes:
- /proc:/host/proc:ro
- /sys:/host/sys:ro
- /:/rootfs:ro
- /etc/hostname:/etc/nodename:ro
command:
- "--path.sysfs=/host/sys"
- "--path.procfs=/host/proc"
- "--path.rootfs=/rootfs"
- "--collector.textfile.directory=/etc/node-exporter/"
- "--collector.filesystem.ignored-mount-points=^/(sys|proc|dev|host|etc)($$|/)"
- "--no-collector.ipvs"
image: grafana/alloy:v1.18.1
hostname: "${DOMAIN}"
configs:
- source: entrypoint
target: /entrypoint.sh
- source: config_alloy
target: /etc/alloy/config.alloy
volumes:
- /:/rootfs:ro
- /var/run/docker.sock:/var/run/docker.sock:ro
- /sys:/sys:ro
- /var/lib/docker:/var/lib/docker:ro
# long-form avoids toolshed/abra#900
- type: bind
source: "${CONTAINERD_SOCKET:-/run/containerd/containerd.sock}"
target: /run/containerd/containerd.sock
read_only: true
- alloy-data:/var/lib/alloy/data
# runs through a shell so ALLOY_HTTP_LISTEN_ADDR (set by
# compose.alloy-webui.yml) is resolved from the container's own
# environment at startup, not by compose at deploy time.
# $$ escapes it from compose's own interpolation
entrypoint: ["/bin/sh", "-c"]
command:
- >-
exec alloy run
--storage.path=/var/lib/alloy/data
--server.http.listen-addr=$${ALLOY_HTTP_LISTEN_ADDR:-127.0.0.1}:12345
/etc/alloy/config.alloy
networks:
- internal
- proxy
entrypoint: [ "/bin/sh", "-e", "/entrypoint.sh" ]
- internal
environment:
- SCRAPE_BASIC_AUTH_USERNAME=${SCRAPE_BASIC_AUTH_USERNAME:-alloy}
- WRITE_BASIC_AUTH_USERNAME=${WRITE_BASIC_AUTH_USERNAME:-admin}
- LIVE_DEBUGGING=${LIVE_DEBUGGING:-false}
- NODE_ID={{.Node.ID}}
secrets:
- basic_auth
# no wget/curl in this image; bash's /dev/tcp is used instead. Works
# against localhost regardless of ALLOY_HTTP_LISTEN_ADDR
healthcheck:
test: ["CMD", "bash", "-c", "exec 3<>/dev/tcp/localhost/12345 && printf 'GET /-/ready HTTP/1.0\r\nHost: localhost\r\n\r\n' >&3 && head -1 <&3 | grep -q 200"]
interval: 5s
timeout: 10s
retries: 3
start_period: 10s
deploy:
restart_policy:
condition: on-failure
labels:
- "backupbot.backup=${ENABLE_BACKUPS:-true}"
- "traefik.enable=true"
- "traefik.docker.network=proxy"
- "traefik.http.services.${STACK_NAME}-node.loadbalancer.server.port=9100"
- "traefik.http.routers.${STACK_NAME}-node.rule=Host(`node.${DOMAIN}`)"
- "traefik.http.routers.${STACK_NAME}-node.entrypoints=web-secure"
- "traefik.http.routers.${STACK_NAME}-node.tls=true"
- "traefik.http.routers.${STACK_NAME}-node.tls.certresolver=${LETS_ENCRYPT_ENV}"
- "traefik.http.routers.${STACK_NAME}-node.middlewares=basicauth@file"
- "coop-cloud.${STACK_NAME}.version=1.6.0+v1.8.1"
- "coop-cloud.${STACK_NAME}.timeout=${TIMEOUT}"
cadvisor:
image: gcr.io/cadvisor/cadvisor:v0.55.1
command:
- "-logtostderr"
- "--enable_metrics=cpu,cpuLoad,disk,diskIO,process,memory,network"
# all possible metrics: advtcp,app,cpu,cpuLoad,cpu_topology,cpuset,disk,diskIO,hugetlb,memory,memory_numa,network,oom_event,percpu,perf_event,process,referenced_memory,resctrl,sched,tcp,udp.
- "--housekeeping_interval=120s"
- "--docker_only=true"
volumes:
- /var/lib/docker/:/var/lib/docker:ro
- /dev/disk/:/dev/disk:ro
- /sys:/sys:ro
- /var/run:/var/run:ro
- /:/rootfs:ro
networks:
- internal
- proxy
deploy:
restart_policy:
condition: on-failure
labels:
- "traefik.enable=true"
- "traefik.docker.network=proxy"
- "traefik.http.services.${STACK_NAME}-cadvisor.loadbalancer.server.port=8080"
- "traefik.http.routers.${STACK_NAME}-cadvisor.rule=Host(`cadvisor.${DOMAIN}`)"
- "traefik.http.routers.${STACK_NAME}-cadvisor.entrypoints=web-secure"
- "traefik.http.routers.${STACK_NAME}-cadvisor.tls=true"
- "traefik.http.routers.${STACK_NAME}-cadvisor.tls.certresolver=${LETS_ENCRYPT_ENV}"
- "traefik.http.routers.${STACK_NAME}-cadvisor.middlewares=basicauth@file"
healthcheck:
test: wget --quiet --tries=1 --spider http://localhost:8080/healthz || exit 1
interval: 15s
timeout: 15s
retries: 5
start_period: 30s
configs:
entrypoint:
name: ${STACK_NAME}_entrypoint_${ENTRYPOINT_VERSION}
file: node-exporter-entrypoint.sh
config_alloy:
template_driver: golang
name: ${STACK_NAME}_config_alloy_${CONFIG_ALLOY_VERSION}
file: config.alloy.tmpl
networks:
proxy:
external: true
internal:
volumes:
alloy-data:
secrets:
basic_auth:
external: true
name: ${STACK_NAME}_basic_auth_${SECRET_BASIC_AUTH_VERSION}
+349
View File
@@ -0,0 +1,349 @@
logging {
level = "info"
format = "logfmt"
}
livedebugging {
enabled = {{ env "LIVE_DEBUGGING" }}
}
discovery.docker "linux" {
host = "unix:///var/run/docker.sock"
}
{{ if ne (env "PROMETHEUS_REMOTE_WRITE_URL") "" }}
prometheus.exporter.cadvisor "docker" {
docker_only = true
enabled_metrics = ["cpu", "cpuLoad", "disk", "diskIO", "memory", "network", "process"]
}
prometheus.exporter.unix "default" {
include_exporter_metrics = true
rootfs_path = "/rootfs"
procfs_path = "/rootfs/proc"
sysfs_path = "/rootfs/sys"
disable_collectors = ["ipvs"]
filesystem {
fs_types_exclude = "^(autofs|binfmt_misc|bpf|cgroup2?|configfs|debugfs|devpts|devtmpfs|tmpfs|fusectl|hugetlbfs|iso9660|mqueue|nsfs|overlay|proc|procfs|pstore|rpc_pipefs|securityfs|selinuxfs|squashfs|sysfs|tracefs)$"
mount_points_exclude = "^/(sys|proc|dev|host|etc)($|/)"
mount_timeout = "5s"
}
netclass { ignored_devices = "^(veth.*)$" }
netdev { device_exclude = "^(veth.*)$" }
}
prometheus.exporter.self "alloy" {}
prometheus.scrape "default" {
scrape_interval = "120s"
targets = array.concat(
prometheus.exporter.self.alloy.targets,
prometheus.exporter.unix.default.targets,
prometheus.exporter.cadvisor.docker.targets,
)
forward_to = [prometheus.remote_write.prometheus.receiver]
}
prometheus.remote_write "prometheus" {
endpoint {
url = "{{ env "PROMETHEUS_REMOTE_WRITE_URL" }}"
basic_auth {
username = "{{ env "WRITE_BASIC_AUTH_USERNAME" }}"
password = "{{ secret "basic_auth" }}"
}
}
}
// Scrape Prometheus metrics from other containers on this host.
// Containers opt in via Docker labels:
// prometheus.io/scrape=true required: enable scraping
// prometheus.io/port=9090 optional: port exposing /metrics (defaults to 80 if not set)
// prometheus.io/path=/metrics optional: path to metrics endpoint (default: /metrics)
// prometheus.io/auth=basic optional: use basic auth with the scrape_auth secret (see compose.scrape-auth.yml)
// prometheus.io/auth=bearer optional: use bearer auth with the scrape_auth secret (see compose.scrape-auth.yml)
discovery.dockerswarm "swarm" {
host = "unix:///var/run/docker.sock"
// "tasks" not "services": dnsrr-mode services (e.g. traefik) have no VIP
// and are invisible to the "services" role
role = "tasks"
}
discovery.relabel "metrics" {
targets = discovery.dockerswarm.swarm.targets
// skip old task history, only scrape currently-running tasks
rule {
source_labels = ["__meta_dockerswarm_task_desired_state"]
regex = "running"
action = "keep"
}
// only scrape hosts running on this host within the swam
rule {
source_labels = ["__meta_dockerswarm_node_id"]
regex = "{{ env "NODE_ID" }}"
action = "keep"
}
rule {
source_labels = ["__meta_dockerswarm_network_name"]
regex = "proxy"
action = "keep"
}
rule {
source_labels = ["__meta_dockerswarm_service_label_prometheus_io_scrape"]
regex = "true"
action = "keep"
}
// default to port 80 when prometheus.io/port isn't set
rule {
source_labels = ["__meta_dockerswarm_service_label_prometheus_io_port"]
regex = "^$"
target_label = "__meta_dockerswarm_service_label_prometheus_io_port"
replacement = "80"
}
// a task with multiple published ports produces one target per port;
// this unifies all of them to the single port above, so duplicates
// collapse at scrape time instead of scraping every port
rule {
source_labels = ["__address__", "__meta_dockerswarm_service_label_prometheus_io_port"]
regex = `(.+):\d+;(\d+)`
target_label = "__address__"
replacement = "$1:$2"
}
rule {
source_labels = ["__meta_dockerswarm_service_label_prometheus_io_path"]
regex = `(.+)`
target_label = "__metrics_path__"
}
rule {
source_labels = ["__meta_dockerswarm_service_name"]
target_label = "job"
}
// task IDs (and the default address-derived instance label) change on
// every redeploy; service+slot is stable across redeploys of the same
// replica, so data stays continuous instead of restarting each deploy
rule {
source_labels = ["__meta_dockerswarm_service_name", "__meta_dockerswarm_task_slot"]
separator = "."
target_label = "instance"
}
rule {
source_labels = ["__meta_dockerswarm_task_slot"]
target_label = "task_slot"
}
// coop-cloud's STACK_NAME is the domain with "." replaced by "_"
// Derive a readable dotted domain label from it. RE2 has no global
// replace, so this is done by chaining multiple replacements. Each
// rule swaps the first remaining _ for a "." until none are left.
rule {
source_labels = ["__meta_dockerswarm_service_label_com_docker_stack_namespace"]
target_label = "domain"
}
rule {
source_labels = ["domain"]
regex = `([^_]*)_(.*)`
target_label = "domain"
replacement = "$1.$2"
}
rule {
source_labels = ["domain"]
regex = `([^_]*)_(.*)`
target_label = "domain"
replacement = "$1.$2"
}
rule {
source_labels = ["domain"]
regex = `([^_]*)_(.*)`
target_label = "domain"
replacement = "$1.$2"
}
rule {
source_labels = ["domain"]
regex = `([^_]*)_(.*)`
target_label = "domain"
replacement = "$1.$2"
}
rule {
source_labels = ["domain"]
regex = `([^_]*)_(.*)`
target_label = "domain"
replacement = "$1.$2"
}
rule {
source_labels = ["domain"]
regex = `([^_]*)_(.*)`
target_label = "domain"
replacement = "$1.$2"
}
}
discovery.relabel "metrics_noauth" {
targets = discovery.relabel.metrics.output
rule {
source_labels = ["__meta_dockerswarm_service_label_prometheus_io_auth"]
regex = "^$"
action = "keep"
}
}
discovery.relabel "metrics_basicauth" {
targets = discovery.relabel.metrics.output
rule {
source_labels = ["__meta_dockerswarm_service_label_prometheus_io_auth"]
regex = "basic"
action = "keep"
}
}
discovery.relabel "metrics_bearerauth" {
targets = discovery.relabel.metrics.output
rule {
source_labels = ["__meta_dockerswarm_service_label_prometheus_io_auth"]
regex = "bearer"
action = "keep"
}
}
prometheus.scrape "containers" {
scrape_interval = "120s"
targets = discovery.relabel.metrics_noauth.output
forward_to = [prometheus.remote_write.prometheus.receiver]
}
{{ if ne (env "SECRET_SCRAPE_AUTH_VERSION") "" }}
prometheus.scrape "containers_basicauth" {
scrape_interval = "120s"
targets = discovery.relabel.metrics_basicauth.output
forward_to = [prometheus.remote_write.prometheus.receiver]
basic_auth {
username = "{{ env "SCRAPE_BASIC_AUTH_USERNAME" }}"
password = "{{ secret "scrape_auth" }}"
}
}
prometheus.scrape "containers_bearerauth" {
scrape_interval = "120s"
targets = discovery.relabel.metrics_bearerauth.output
forward_to = [prometheus.remote_write.prometheus.receiver]
bearer_token = "{{ secret "scrape_auth" }}"
}
{{ end }}
{{ end }}
{{ if ne (env "LOKI_PUSH_URL") "" }}
discovery.relabel "docker" {
targets = discovery.docker.linux.targets
rule {
source_labels = ["__meta_docker_container_name"]
target_label = "container_name"
}
rule {
source_labels = ["__meta_docker_container_id"]
target_label = "container_id"
}
rule {
source_labels = ["__meta_docker_container_label_com_docker_stack_namespace"]
target_label = "stack_namespace"
}
rule {
source_labels = ["__meta_docker_container_label_com_docker_swarm_service_name"]
target_label = "service_name"
}
rule {
source_labels = ["__meta_docker_container_log_stream"]
target_label = "stream"
}
}
loki.source.docker "docker" {
host = "unix:///var/run/docker.sock"
targets = discovery.relabel.docker.output
labels = {"app" = "docker"}
forward_to = [loki.write.loki.receiver]
}
// JOURNALD: reads the systemd journal binary log directly.
// Use on systemd hosts (most modern Linux distros). Requires no syslogd.
{{ if eq (env "JOURNALD") "1" }}
loki.source.journal "journal" {
path = "/rootfs/var/log/journal"
labels = { job = "{{ env "DOMAIN" }}" }
forward_to = [loki.write.loki.receiver]
}
{{ end }}
// SYSLOG_FILES: tails all /var/log/*log files (syslog, auth.log, kern.log, etc.).
// Use on non-systemd hosts where a syslogd writes to /var/log.
{{ if eq (env "SYSLOG_FILES") "1" }}
local.file_match "syslog_files" {
path_targets = [{ __path__ = "/rootfs/var/log/*log" }]
}
loki.source.file "syslog_files" {
targets = local.file_match.syslog_files.targets
forward_to = [loki.process.syslog_files.receiver]
}
loki.process "syslog_files" {
stage.static_labels {
values = { job = "syslog" }
}
forward_to = [loki.write.loki.receiver]
}
{{ end }}
// SYSLOG: opens a network syslog listener on port 514.
// Use when a remote device or a local syslogd configured to
// forward over the network sends logs to this host.
// Requires compose.syslog.yml to publish port 514 to the host.
// This is NOT needed for reading local log files — use SYSLOG_FILES instead.
{{ if eq (env "SYSLOG") "1" }}
loki.relabel "syslog" {
rule {
action = "labelmap"
regex = "__syslog_(.+)"
}
forward_to = []
}
loki.source.syslog "syslog" {
listener {
address = "[::]:514"
label_structured_data = true
labels = { component = "loki.source.syslog" }
}
relabel_rules = loki.relabel.syslog.rules
forward_to = [loki.write.loki.receiver]
}
{{ end }}
loki.write "loki" {
endpoint {
url = "{{ env "LOKI_PUSH_URL" }}"
basic_auth {
username = "{{ env "WRITE_BASIC_AUTH_USERNAME" }}"
password = "{{ secret "basic_auth" }}"
}
}
external_labels = { hostname = "{{ env "DOMAIN" }}" }
}
{{ end }}
+22 -22
View File
@@ -110,7 +110,7 @@
"uid": "PBFA97CFB590B2093"
},
"editorMode": "code",
"expr": "(time() - min(container_start_time_seconds{container_label_com_docker_stack_namespace=~\"$stack\", instance=~\"cadvisor.monitoring.$instance.*\"}))",
"expr": "(time() - min(container_start_time_seconds{container_label_com_docker_stack_namespace=~\"$stack\", instance=~\".*$node_id.*\"}))",
"format": "time_series",
"intervalFactor": 1,
"legendFormat": "",
@@ -215,7 +215,7 @@
"uid": "PBFA97CFB590B2093"
},
"editorMode": "code",
"expr": "sort(\n sum(\n rate(\n container_cpu_usage_seconds_total{\n container_label_com_docker_stack_namespace=~\"$stack\", \n instance=~\"cadvisor.monitoring.$instance.*\"}[$interval]\n )\n ) \n by (container_label_com_docker_swarm_service_name)\n)",
"expr": "sort(\n sum(\n rate(\n container_cpu_usage_seconds_total{\n container_label_com_docker_stack_namespace=~\"$stack\", \n instance=~\".*$node_id.*\"}[$interval]\n )\n ) \n by (container_label_com_docker_swarm_service_name)\n)",
"format": "time_series",
"intervalFactor": 2,
"legendFormat": "{{ container_label_com_docker_swarm_service_name }}",
@@ -283,7 +283,7 @@
"uid": "PBFA97CFB590B2093"
},
"editorMode": "code",
"expr": "count(rate(container_last_seen{container_label_com_docker_stack_namespace=~\"$stack\", instance=~\"cadvisor.monitoring.$instance.*\"}[$interval]))",
"expr": "count(rate(container_last_seen{container_label_com_docker_stack_namespace=~\"$stack\", instance=~\".*$node_id.*\"}[$interval]))",
"format": "time_series",
"intervalFactor": 2,
"range": true,
@@ -386,7 +386,7 @@
"uid": "PBFA97CFB590B2093"
},
"editorMode": "code",
"expr": "sum(\n container_memory_rss{\n container_label_com_docker_stack_namespace=~\"$stack\", \n instance=~\"cadvisor.monitoring.$instance.*\"}\n ) \nby (container_label_com_docker_swarm_service_name)",
"expr": "sum(\n container_memory_rss{\n container_label_com_docker_stack_namespace=~\"$stack\", \n instance=~\".*$node_id.*\"}\n ) \nby (container_label_com_docker_swarm_service_name)",
"format": "time_series",
"hide": false,
"intervalFactor": 2,
@@ -493,7 +493,7 @@
"uid": "PBFA97CFB590B2093"
},
"editorMode": "code",
"expr": "sum(rate(container_network_receive_bytes_total{container_label_com_docker_stack_namespace=~\"$stack\", instance=~\"cadvisor.monitoring.$instance.*\"}[$interval])) by (container_label_com_docker_swarm_service_name)",
"expr": "sum(rate(container_network_receive_bytes_total{container_label_com_docker_stack_namespace=~\"$stack\", instance=~\".*$node_id.*\"}[$interval])) by (container_label_com_docker_swarm_service_name)",
"format": "time_series",
"intervalFactor": 2,
"legendFormat": "RX: {{ container_label_com_docker_swarm_service_name }}",
@@ -507,7 +507,7 @@
"uid": "PBFA97CFB590B2093"
},
"editorMode": "code",
"expr": "sum(rate(container_network_transmit_bytes_total{container_label_com_docker_stack_namespace=~\"$stack\", instance=~\"cadvisor.monitoring.$instance.*\"}[$interval])) by (container_label_com_docker_swarm_service_name) * -1",
"expr": "sum(rate(container_network_transmit_bytes_total{container_label_com_docker_stack_namespace=~\"$stack\", instance=~\".*$node_id.*\"}[$interval])) by (container_label_com_docker_swarm_service_name) * -1",
"hide": false,
"legendFormat": "TX: {{container_label_com_docker_swarm_service_name}}",
"range": true,
@@ -634,7 +634,7 @@
"uid": "PBFA97CFB590B2093"
},
"editorMode": "code",
"expr": "sum(container_fs_usage_bytes{container_label_com_docker_stack_namespace=~\"$stack\", instance=~\"cadvisor.monitoring.$instance.*\"}) by (container_label_com_docker_swarm_service_name)",
"expr": "sum(container_fs_usage_bytes{container_label_com_docker_stack_namespace=~\"$stack\", instance=~\".*$node_id.*\"}) by (container_label_com_docker_swarm_service_name)",
"format": "time_series",
"hide": false,
"intervalFactor": 2,
@@ -742,7 +742,7 @@
"uid": "PBFA97CFB590B2093"
},
"editorMode": "code",
"expr": "sum(rate(container_fs_usage_bytes{container_label_com_docker_stack_namespace=~\"$stack\", instance=~\"cadvisor.monitoring.$instance.*\"}[$interval])) by (container_label_com_docker_swarm_service_name)",
"expr": "sum(rate(container_fs_usage_bytes{container_label_com_docker_stack_namespace=~\"$stack\", instance=~\".*$node_id.*\"}[$interval])) by (container_label_com_docker_swarm_service_name)",
"format": "time_series",
"hide": false,
"intervalFactor": 2,
@@ -865,7 +865,7 @@
"uid": "PBFA97CFB590B2093"
},
"editorMode": "code",
"expr": "sum(rate(container_fs_io_time_seconds_total{container_label_com_docker_stack_namespace=~\"$stack\", instance=~\"cadvisor.monitoring.$instance.*\"}[$interval])) by (container_label_com_docker_swarm_service_name)",
"expr": "sum(rate(container_fs_io_time_seconds_total{container_label_com_docker_stack_namespace=~\"$stack\", instance=~\".*$node_id.*\"}[$interval])) by (container_label_com_docker_swarm_service_name)",
"format": "time_series",
"hide": false,
"intervalFactor": 2,
@@ -973,7 +973,7 @@
"uid": "PBFA97CFB590B2093"
},
"editorMode": "code",
"expr": "container_fs_io_current{container_label_com_docker_stack_namespace=~\"$stack\", instance=~\"cadvisor.monitoring.$instance.*\"}",
"expr": "container_fs_io_current{container_label_com_docker_stack_namespace=~\"$stack\", instance=~\".*$node_id.*\"}",
"format": "time_series",
"hide": false,
"intervalFactor": 2,
@@ -1081,7 +1081,7 @@
"uid": "PBFA97CFB590B2093"
},
"editorMode": "code",
"expr": "sum(rate(container_fs_reads_bytes_total{container_label_com_docker_stack_namespace=~\"$stack\", instance=~\"cadvisor.monitoring.$instance.*\"}[$interval])) by (container_label_com_docker_swarm_service_name)",
"expr": "sum(rate(container_fs_reads_bytes_total{container_label_com_docker_stack_namespace=~\"$stack\", instance=~\".*$node_id.*\"}[$interval])) by (container_label_com_docker_swarm_service_name)",
"format": "time_series",
"intervalFactor": 2,
"legendFormat": "{{ container_label_com_docker_swarm_service_name }}",
@@ -1188,7 +1188,7 @@
"uid": "PBFA97CFB590B2093"
},
"editorMode": "code",
"expr": "sum(rate(container_fs_read_seconds_total{container_label_com_docker_stack_namespace=~\"$stack\", instance=~\"cadvisor.monitoring.$instance.*\"}[$interval])) by (container_label_com_docker_swarm_service_name)",
"expr": "sum(rate(container_fs_read_seconds_total{container_label_com_docker_stack_namespace=~\"$stack\", instance=~\".*$node_id.*\"}[$interval])) by (container_label_com_docker_swarm_service_name)",
"format": "time_series",
"intervalFactor": 2,
"legendFormat": "{{ container_label_com_docker_swarm_service_name }}",
@@ -1295,7 +1295,7 @@
"uid": "PBFA97CFB590B2093"
},
"editorMode": "code",
"expr": "sum(rate(container_fs_writes_bytes_total{container_label_com_docker_stack_namespace=~\"$stack\", instance=~\"cadvisor.monitoring.$instance.*\"}[$interval])) by (container_label_com_docker_swarm_service_name)",
"expr": "sum(rate(container_fs_writes_bytes_total{container_label_com_docker_stack_namespace=~\"$stack\", instance=~\".*$node_id.*\"}[$interval])) by (container_label_com_docker_swarm_service_name)",
"format": "time_series",
"intervalFactor": 2,
"legendFormat": "{{ container_label_com_docker_swarm_service_name }}",
@@ -1402,7 +1402,7 @@
"uid": "PBFA97CFB590B2093"
},
"editorMode": "code",
"expr": "sum(rate(container_fs_write_seconds_total{container_label_com_docker_stack_namespace=~\"$stack\", instance=~\"cadvisor.monitoring.$instance.*\"}[$interval])) by (container_label_com_docker_swarm_service_name)",
"expr": "sum(rate(container_fs_write_seconds_total{container_label_com_docker_stack_namespace=~\"$stack\", instance=~\".*$node_id.*\"}[$interval])) by (container_label_com_docker_swarm_service_name)",
"format": "time_series",
"intervalFactor": 2,
"legendFormat": "{{ container_label_com_docker_swarm_service_name }}",
@@ -1453,7 +1453,7 @@
"query": "label_values(container_label_com_docker_stack_namespace)",
"refId": "PrometheusVariableQueryEditor-VariableQuery"
},
"refresh": 1,
"refresh": 2,
"regex": "",
"skipUrlSync": false,
"sort": 2,
@@ -1567,19 +1567,19 @@
"type": "prometheus",
"uid": "PBFA97CFB590B2093"
},
"definition": "label_values(instance)",
"definition": "label_values(node_uname_info, instance)",
"hide": 0,
"includeAll": true,
"label": "instance",
"label": "Swarm Node",
"multi": true,
"name": "instance",
"name": "node_id",
"options": [],
"query": {
"query": "label_values(instance)",
"query": "label_values(node_uname_info, instance)",
"refId": "PrometheusVariableQueryEditor-VariableQuery"
},
"refresh": 1,
"regex": "/.*cadvisor.monitoring.(?<instance>.*):80/",
"refresh": 2,
"regex": "/^(?:(?:node|cadvisor|monitoring).)*(.*)$/",
"skipUrlSync": false,
"sort": 0,
"type": "query"
@@ -1620,4 +1620,4 @@
"uid": "KdVoGQm7z",
"version": 36,
"weekStart": ""
}
}
+39 -39
View File
@@ -118,7 +118,7 @@
"type": "prometheus",
"uid": "PBFA97CFB590B2093"
},
"expr": "topk(1, sum((node_time_seconds - node_boot_time_seconds) * on(instance) group_left(node_name) node_meta{node_name=~\"$node_id\"}) by (node_name))",
"expr": "topk(1, sum(label_replace(node_time_seconds{instance=~\".*$node_id.*\"} - node_boot_time_seconds{instance=~\".*$node_id.*\"}, \"instance\", \"$1\", \"instance\", \"^(?:(?:node|cadvisor|monitoring).)*(.*)$\")) by (instance))",
"format": "time_series",
"intervalFactor": 2,
"legendFormat": "",
@@ -198,7 +198,7 @@
"type": "prometheus",
"uid": "PBFA97CFB590B2093"
},
"expr": "count(node_meta * on(instance) group_left(node_name) node_meta{node_name=~\"$node_id\"})",
"expr": "count(label_replace(node_uname_info{instance=~\".*$node_id.*\"}, \"instance\", \"$1\", \"instance\", \"^(?:(?:node|cadvisor|monitoring).)*(.*)$\"))",
"format": "time_series",
"intervalFactor": 2,
"legendFormat": "",
@@ -278,7 +278,7 @@
"type": "prometheus",
"uid": "PBFA97CFB590B2093"
},
"expr": "count(node_cpu_seconds_total{mode=\"idle\"} * on(instance) group_left(node_name) node_meta{node_name=~\"$node_id\"})",
"expr": "count(label_replace(node_cpu_seconds_total{mode=\"idle\", instance=~\".*$node_id.*\"}, \"instance\", \"$1\", \"instance\", \"^(?:(?:node|cadvisor|monitoring).)*(.*)$\"))",
"format": "time_series",
"intervalFactor": 2,
"legendFormat": "",
@@ -363,7 +363,7 @@
"uid": "PBFA97CFB590B2093"
},
"editorMode": "code",
"expr": "sum((node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * on(instance) group_left(node_name) node_meta{node_name=~\"$node_id\"} * 100) / count(node_meta * on(instance) group_left(node_name) node_meta{node_name=~\"$node_id\"})",
"expr": "sum(label_replace((node_memory_MemAvailable_bytes{instance=~\".*$node_id.*\"} / node_memory_MemTotal_bytes{instance=~\".*$node_id.*\"}) * 100, \"instance\", \"$1\", \"instance\", \"^(?:(?:node|cadvisor|monitoring).)*(.*)$\")) / count(label_replace(node_uname_info{instance=~\".*$node_id.*\"}, \"instance\", \"$1\", \"instance\", \"^(?:(?:node|cadvisor|monitoring).)*(.*)$\"))",
"format": "time_series",
"intervalFactor": 2,
"legendFormat": "",
@@ -429,11 +429,11 @@
"type": "prometheus",
"uid": "PBFA97CFB590B2093"
},
"expr": "node_load5 * on(instance) group_left(node_name) node_meta{node_name=~\"$node_id\"}",
"expr": "label_replace(node_load5{instance=~\".*$node_id.*\"}, \"instance\", \"$1\", \"instance\", \"^(?:(?:node|cadvisor|monitoring).)*(.*)$\")",
"format": "time_series",
"interval": "",
"intervalFactor": 2,
"legendFormat": "load5 {{node_name}}",
"legendFormat": "load5 {{instance}}",
"refId": "A",
"step": 2
}
@@ -536,7 +536,7 @@
"type": "prometheus",
"uid": "PBFA97CFB590B2093"
},
"expr": "sum(node_memory_MemTotal_bytes * on(instance) group_left(node_name) node_meta{node_name=~\"$node_id\"})",
"expr": "sum(label_replace(node_memory_MemTotal_bytes{instance=~\".*$node_id.*\"}, \"instance\", \"$1\", \"instance\", \"^(?:(?:node|cadvisor|monitoring).)*(.*)$\"))",
"format": "time_series",
"intervalFactor": 2,
"legendFormat": "",
@@ -616,7 +616,7 @@
"uid": "PBFA97CFB590B2093"
},
"exemplar": true,
"expr": "sum(node_filesystem_size_bytes{fstype=\"ext4\",mountpoint=~\"(/$)|(/media.*)\"} * on(instance) group_left(node_name) node_meta{node_name=~\"$node_id\"})",
"expr": "sum(label_replace(node_filesystem_size_bytes{fstype=~\"ext4|xfs\",mountpoint=~\"(/$)|(/media.*)\", instance=~\".*$node_id.*\"}, \"instance\", \"$1\", \"instance\", \"^(?:(?:node|cadvisor|monitoring).)*(.*)$\"))",
"format": "time_series",
"interval": "",
"intervalFactor": 2,
@@ -701,7 +701,7 @@
"type": "prometheus",
"uid": "PBFA97CFB590B2093"
},
"expr": "sum(irate(node_cpu_seconds_total{mode=\"idle\"}[$interval]) * on(instance) group_left(node_name) node_meta{node_name=~\"$node_id\"}) * 100 / count(node_cpu_seconds_total{mode=\"user\"} * on(instance) group_left(node_name) node_meta{node_name=~\"$node_id\"}) ",
"expr": "sum(label_replace(irate(node_cpu_seconds_total{mode=\"idle\", instance=~\".*$node_id.*\"}[$interval]), \"instance\", \"$1\", \"instance\", \"^(?:(?:node|cadvisor|monitoring).)*(.*)$\")) * 100 / count(label_replace(node_cpu_seconds_total{mode=\"user\", instance=~\".*$node_id.*\"}, \"instance\", \"$1\", \"instance\", \"^(?:(?:node|cadvisor|monitoring).)*(.*)$\"))",
"format": "time_series",
"intervalFactor": 2,
"legendFormat": "",
@@ -833,7 +833,7 @@
},
"editorMode": "code",
"exemplar": true,
"expr": "node_filesystem_free_bytes{fstype=\"ext4\"} / node_filesystem_size_bytes{fstype=\"ext4\"} * on(instance) group_left(node_name) node_meta{node_name=~\"$node_id\"} * 100",
"expr": "label_replace(node_filesystem_free_bytes{fstype=~\"ext4|xfs\", instance=~\".*$node_id.*\"}, \"instance\", \"$1\", \"instance\", \"^(?:(?:node|cadvisor|monitoring).)*(.*)$\") / label_replace(node_filesystem_size_bytes{fstype=~\"ext4|xfs\", instance=~\".*$node_id.*\"}, \"instance\", \"$1\", \"instance\", \"^(?:(?:node|cadvisor|monitoring).)*(.*)$\") * 100",
"format": "time_series",
"interval": "",
"intervalFactor": 2,
@@ -900,10 +900,10 @@
"type": "prometheus",
"uid": "PBFA97CFB590B2093"
},
"expr": "100 - (avg(irate(node_cpu_seconds_total{mode=\"idle\"}[$interval]) * on(instance) group_left(node_name) node_meta{node_name=~\"$node_id\"} * 100) by (node_name))",
"expr": "100 - (avg(label_replace(irate(node_cpu_seconds_total{mode=\"idle\", instance=~\".*$node_id.*\"}[$interval]), \"instance\", \"$1\", \"instance\", \"^(?:(?:node|cadvisor|monitoring).)*(.*)$\") * 100) by (instance))",
"format": "time_series",
"intervalFactor": 2,
"legendFormat": "{{node_name}}",
"legendFormat": "{{instance}}",
"refId": "A",
"step": 2
}
@@ -1045,12 +1045,12 @@
"uid": "PBFA97CFB590B2093"
},
"editorMode": "code",
"expr": "sum((node_memory_MemFree_bytes) * on(instance) group_left(node_name) node_meta{node_name=~\"$node_id\"}) by (node_name)",
"expr": "sum(label_replace(node_memory_MemFree_bytes{instance=~\".*$node_id.*\"}, \"instance\", \"$1\", \"instance\", \"^(?:(?:node|cadvisor|monitoring).)*(.*)$\")) by (instance)",
"format": "time_series",
"hide": false,
"interval": "",
"intervalFactor": 2,
"legendFormat": "Free {{node_name}}",
"legendFormat": "Free {{instance}}",
"range": true,
"refId": "free",
"step": 2
@@ -1061,12 +1061,12 @@
"uid": "PBFA97CFB590B2093"
},
"editorMode": "code",
"expr": "sum((node_memory_Cached_bytes + node_memory_Buffers_bytes + node_memory_Slab_bytes) * on(instance) group_left(node_name) node_meta{node_name=~\"$node_id\"}) by (node_name)",
"expr": "sum(label_replace(node_memory_Cached_bytes{instance=~\".*$node_id.*\"} + node_memory_Buffers_bytes{instance=~\".*$node_id.*\"} + node_memory_Slab_bytes{instance=~\".*$node_id.*\"}, \"instance\", \"$1\", \"instance\", \"^(?:(?:node|cadvisor|monitoring).)*(.*)$\")) by (instance)",
"format": "time_series",
"hide": false,
"interval": "",
"intervalFactor": 2,
"legendFormat": "cache,buffer,slab {{node_name}}",
"legendFormat": "cache,buffer,slab {{instance}}",
"range": true,
"refId": "cache,buffer,slab",
"step": 2
@@ -1077,12 +1077,12 @@
"uid": "PBFA97CFB590B2093"
},
"editorMode": "code",
"expr": "sum((node_memory_MemTotal_bytes - node_memory_MemFree_bytes - node_memory_Cached_bytes - node_memory_Buffers_bytes - node_memory_Slab_bytes) * on(instance) group_left(node_name) node_meta{node_name=~\"$node_id\"}) by (node_name)",
"expr": "sum(label_replace(node_memory_MemTotal_bytes{instance=~\".*$node_id.*\"} - node_memory_MemFree_bytes{instance=~\".*$node_id.*\"} - node_memory_Cached_bytes{instance=~\".*$node_id.*\"} - node_memory_Buffers_bytes{instance=~\".*$node_id.*\"} - node_memory_Slab_bytes{instance=~\".*$node_id.*\"}, \"instance\", \"$1\", \"instance\", \"^(?:(?:node|cadvisor|monitoring).)*(.*)$\")) by (instance)",
"format": "time_series",
"hide": false,
"interval": "",
"intervalFactor": 2,
"legendFormat": "Used {{node_name}}",
"legendFormat": "Used {{instance}}",
"range": true,
"refId": "Used",
"step": 2
@@ -1094,13 +1094,13 @@
},
"editorMode": "code",
"exemplar": false,
"expr": "sum((node_memory_MemTotal_bytes) * on(instance) group_left(node_name) node_meta{node_name=~\"$node_id\"}) by (node_name)",
"expr": "sum(label_replace(node_memory_MemTotal_bytes{instance=~\".*$node_id.*\"}, \"instance\", \"$1\", \"instance\", \"^(?:(?:node|cadvisor|monitoring).)*(.*)$\")) by (instance)",
"format": "time_series",
"hide": false,
"instant": false,
"interval": "",
"intervalFactor": 2,
"legendFormat": "Total {{node_name}}",
"legendFormat": "Total {{instance}}",
"range": true,
"refId": "total",
"step": 2
@@ -1161,11 +1161,11 @@
"type": "prometheus",
"uid": "PBFA97CFB590B2093"
},
"expr": "sum(irate(node_disk_read_bytes_total[$interval]) * on(instance) group_left(node_name) node_meta{node_name=~\"$node_id\"}) by (node_name)",
"expr": "sum(label_replace(irate(node_disk_read_bytes_total{instance=~\".*$node_id.*\"}[$interval]), \"instance\", \"$1\", \"instance\", \"^(?:(?:node|cadvisor|monitoring).)*(.*)$\")) by (instance)",
"format": "time_series",
"interval": "",
"intervalFactor": 2,
"legendFormat": "Read {{node_name}}",
"legendFormat": "Read {{instance}}",
"refId": "A",
"step": 2
},
@@ -1174,10 +1174,10 @@
"type": "prometheus",
"uid": "PBFA97CFB590B2093"
},
"expr": "sum(irate(node_disk_written_bytes_total[$interval]) * on(instance) group_left(node_name) node_meta{node_id=~\"$node_id\"}) by (node_name)",
"expr": "sum(label_replace(irate(node_disk_written_bytes_total{instance=~\".*$node_id.*\"}[$interval]), \"instance\", \"$1\", \"instance\", \"^(?:(?:node|cadvisor|monitoring).)*(.*)$\")) by (instance)",
"format": "time_series",
"intervalFactor": 2,
"legendFormat": "Written {{node_name}}",
"legendFormat": "Written {{instance}}",
"refId": "B",
"step": 2
}
@@ -1264,10 +1264,10 @@
"type": "prometheus",
"uid": "PBFA97CFB590B2093"
},
"expr": "sum(irate(node_disk_reads_completed_total[$interval]) * on(instance) group_left(node_name) node_meta{node_name=~\"$node_id\"}) by (node_name)",
"expr": "sum(label_replace(irate(node_disk_reads_completed_total{instance=~\".*$node_id.*\"}[$interval]), \"instance\", \"$1\", \"instance\", \"^(?:(?:node|cadvisor|monitoring).)*(.*)$\")) by (instance)",
"format": "time_series",
"intervalFactor": 2,
"legendFormat": "Reads {{node_name}}",
"legendFormat": "Reads {{instance}}",
"refId": "A",
"step": 2
},
@@ -1276,10 +1276,10 @@
"type": "prometheus",
"uid": "PBFA97CFB590B2093"
},
"expr": "sum(irate(node_disk_writes_completed_total[$interval]) * on(instance) group_left(node_name) node_meta{node_name=~\"$node_id\"}) by (node_name)",
"expr": "sum(label_replace(irate(node_disk_writes_completed_total{instance=~\".*$node_id.*\"}[$interval]), \"instance\", \"$1\", \"instance\", \"^(?:(?:node|cadvisor|monitoring).)*(.*)$\")) by (instance)",
"format": "time_series",
"intervalFactor": 2,
"legendFormat": "Writes {{node_name}}",
"legendFormat": "Writes {{instance}}",
"refId": "B",
"step": 2
}
@@ -1368,10 +1368,10 @@
"type": "prometheus",
"uid": "PBFA97CFB590B2093"
},
"expr": "(avg(irate(node_cpu_seconds_total{mode=\"iowait\"}[$interval]) * on(instance) group_left(node_name) node_meta{node_name=~\"$node_id\"} * 100) by (node_name))",
"expr": "(avg(label_replace(irate(node_cpu_seconds_total{mode=\"iowait\", instance=~\".*$node_id.*\"}[$interval]), \"instance\", \"$1\", \"instance\", \"^(?:(?:node|cadvisor|monitoring).)*(.*)$\") * 100) by (instance))",
"format": "time_series",
"intervalFactor": 2,
"legendFormat": "{{node_name}}",
"legendFormat": "{{instance}}",
"refId": "A",
"step": 2
}
@@ -1464,7 +1464,7 @@
"type": "prometheus",
"uid": "PBFA97CFB590B2093"
},
"expr": "sum(rate(container_last_seen[5m]) * on(container_label_com_docker_swarm_node_id) group_left(node_name) node_meta{node_name=~\"${node_id}\"}) by (container_label_com_docker_swarm_service_name)",
"expr": "sum(rate(container_last_seen{instance=~\".*$node_id.*\"}[5m])) by (container_label_com_docker_swarm_service_name)",
"format": "time_series",
"hide": false,
"intervalFactor": 10,
@@ -1573,7 +1573,7 @@
"uid": "PBFA97CFB590B2093"
},
"exemplar": true,
"expr": "count(rate(container_last_seen[5m]) * on(container_label_com_docker_swarm_node_id) group_left(node_name) node_meta{node_name=~\"$node_id\"})",
"expr": "count(rate(container_last_seen{instance=~\".*$node_id.*\"}[5m]))",
"format": "time_series",
"interval": "",
"intervalFactor": 2,
@@ -1668,11 +1668,11 @@
"type": "prometheus",
"uid": "PBFA97CFB590B2093"
},
"expr": "sum(rate(container_network_receive_bytes_total[$interval]) * on(container_label_com_docker_swarm_node_id) group_left(node_name) node_meta{node_name=~\"${node_id}\"}) by (node_name)",
"expr": "sum(label_replace(rate(container_network_receive_bytes_total{instance=~\".*$node_id.*\"}[$interval]), \"instance\", \"$1\", \"instance\", \"^(?:(?:node|cadvisor|monitoring).)*(.*)$\")) by (instance)",
"format": "time_series",
"interval": "",
"intervalFactor": 2,
"legendFormat": "IN {{node_name}}",
"legendFormat": "IN {{instance}}",
"refId": "A",
"step": 2
},
@@ -1681,11 +1681,11 @@
"type": "prometheus",
"uid": "PBFA97CFB590B2093"
},
"expr": "- sum(rate(container_network_transmit_bytes_total[$interval]) * on(container_label_com_docker_swarm_node_id) group_left(node_name) node_meta{node_name=~\"${node_id}\"}) by (node_name)",
"expr": "- sum(label_replace(rate(container_network_transmit_bytes_total{instance=~\".*$node_id.*\"}[$interval]), \"instance\", \"$1\", \"instance\", \"^(?:(?:node|cadvisor|monitoring).)*(.*)$\")) by (instance)",
"format": "time_series",
"hide": false,
"intervalFactor": 2,
"legendFormat": "OUT {{node_name}}",
"legendFormat": "OUT {{instance}}",
"metric": "",
"refId": "B",
"step": 2
@@ -1762,11 +1762,11 @@
"name": "node_id",
"options": [],
"query": {
"query": "node_meta",
"query": "label_values(node_uname_info, instance)",
"refId": "Prometheus-node_id-Variable-Query"
},
"refresh": 1,
"regex": "/node_name=\"([^\"]+)\"/",
"refresh": 2,
"regex": "/^(?:(?:node|cadvisor|monitoring).)*(.*)$/",
"skipUrlSync": false,
"sort": 0,
"type": "query"
+25 -55
View File
@@ -116,12 +116,12 @@
"uid": "PBFA97CFB590B2093"
},
"exemplar": false,
"expr": "time() - process_start_time_seconds{job=\"$job\"}",
"expr": "time() - label_replace(process_start_time_seconds{instance=~\".*$domain.*\"} or process_start_time_seconds{domain=\"$domain\"}, \"domain\", \"$1\", \"instance\", `^metrics\\.(.+)$`)",
"format": "time_series",
"instant": true,
"interval": "",
"intervalFactor": 2,
"legendFormat": "{{ instance }}",
"legendFormat": "{{ domain }}",
"refId": "A"
}
],
@@ -200,7 +200,7 @@
"uid": "PBFA97CFB590B2093"
},
"exemplar": true,
"expr": "sum(increase(traefik_service_requests_total{code=\"499\", instance=\"$instance\"}[$interval]))",
"expr": "sum(label_replace(increase(traefik_service_requests_total{code=\"499\", instance=~\".*$domain.*\"}[$interval]) or increase(traefik_service_requests_total{code=\"499\", domain=\"$domain\"}[$interval]), \"domain\", \"$1\", \"instance\", `^metrics\\.(.+)$`))",
"interval": "",
"legendFormat": "",
"refId": "A"
@@ -280,7 +280,7 @@
"uid": "PBFA97CFB590B2093"
},
"exemplar": true,
"expr": "sum(increase(traefik_service_requests_total{code=\"404\",protocol=~\"$protocol\", instance=\"$instance\"}[$interval]))",
"expr": "sum(label_replace(increase(traefik_service_requests_total{code=\"404\",protocol=~\"$protocol\", instance=~\".*$domain.*\"}[$interval]) or increase(traefik_service_requests_total{code=\"404\",protocol=~\"$protocol\", domain=\"$domain\"}[$interval]), \"domain\", \"$1\", \"instance\", `^metrics\\.(.+)$`))",
"format": "time_series",
"interval": "",
"intervalFactor": 2,
@@ -365,7 +365,7 @@
"uid": "PBFA97CFB590B2093"
},
"exemplar": true,
"expr": "sum(increase(traefik_service_requests_total{code=\"200\",protocol=~\"$protocol\", instance=\"$instance\"}[$interval]))",
"expr": "sum(label_replace(increase(traefik_service_requests_total{code=\"200\",protocol=~\"$protocol\", instance=~\".*$domain.*\"}[$interval]) or increase(traefik_service_requests_total{code=\"200\",protocol=~\"$protocol\", domain=\"$domain\"}[$interval]), \"domain\", \"$1\", \"instance\", `^metrics\\.(.+)$`))",
"format": "time_series",
"interval": "",
"intervalFactor": 2,
@@ -417,7 +417,7 @@
"uid": "PBFA97CFB590B2093"
},
"exemplar": true,
"expr": "topk(5, sum(traefik_service_requests_total{protocol=~\"$protocol\", instance=\"$instance\"}) by (code))",
"expr": "topk(5, sum(label_replace(traefik_service_requests_total{protocol=~\"$protocol\", instance=~\".*$domain.*\"} or traefik_service_requests_total{protocol=~\"$protocol\", domain=\"$domain\"}, \"domain\", \"$1\", \"instance\", `^metrics\\.(.+)$`)) by (code))",
"format": "time_series",
"interval": "",
"intervalFactor": 2,
@@ -489,7 +489,7 @@
"uid": "PBFA97CFB590B2093"
},
"exemplar": true,
"expr": "sum(increase(traefik_service_requests_total{code=\"499\",method=\"GET\", instance=\"$instance\"}[$interval])) by (service)",
"expr": "sum(label_replace(increase(traefik_service_requests_total{code=\"499\",method=\"GET\", instance=~\".*$domain.*\"}[$interval]) or increase(traefik_service_requests_total{code=\"499\",method=\"GET\", domain=\"$domain\"}[$interval]), \"domain\", \"$1\", \"instance\", `^metrics\\.(.+)$`)) by (service)",
"format": "time_series",
"interval": "",
"intervalFactor": 2,
@@ -590,7 +590,7 @@
"uid": "PBFA97CFB590B2093"
},
"exemplar": true,
"expr": "(sum(traefik_entrypoint_request_duration_seconds_sum{protocol=~\"$protocol\", instance=\"$instance\"}) / sum(traefik_entrypoint_requests_total{protocol=~\"$protocol\", instance=\"$instance\"}) * 1000) - (sum(traefik_service_request_duration_seconds_sum{protocol=~\"$protocol\", instance=\"$instance\"}) / sum(traefik_service_requests_total{protocol=~\"$protocol\", instance=\"$instance\"}) * 1000)",
"expr": "(sum(label_replace(traefik_entrypoint_request_duration_seconds_sum{protocol=~\"$protocol\", instance=~\".*$domain.*\"} or traefik_entrypoint_request_duration_seconds_sum{protocol=~\"$protocol\", domain=\"$domain\"}, \"domain\", \"$1\", \"instance\", `^metrics\\.(.+)$`)) / sum(label_replace(traefik_entrypoint_requests_total{protocol=~\"$protocol\", instance=~\".*$domain.*\"} or traefik_entrypoint_requests_total{protocol=~\"$protocol\", domain=\"$domain\"}, \"domain\", \"$1\", \"instance\", `^metrics\\.(.+)$`)) * 1000) - (sum(label_replace(traefik_service_request_duration_seconds_sum{protocol=~\"$protocol\", instance=~\".*$domain.*\"} or traefik_service_request_duration_seconds_sum{protocol=~\"$protocol\", domain=\"$domain\"}, \"domain\", \"$1\", \"instance\", `^metrics\\.(.+)$`)) / sum(label_replace(traefik_service_requests_total{protocol=~\"$protocol\", instance=~\".*$domain.*\"} or traefik_service_requests_total{protocol=~\"$protocol\", domain=\"$domain\"}, \"domain\", \"$1\", \"instance\", `^metrics\\.(.+)$`)) * 1000)",
"format": "time_series",
"instant": false,
"interval": "",
@@ -709,7 +709,7 @@
"uid": "PBFA97CFB590B2093"
},
"exemplar": true,
"expr": "sum(traefik_entrypoint_request_duration_seconds_sum{instance=\"$instance\"}) / sum(traefik_entrypoint_requests_total{instance=\"$instance\"}) * 1000",
"expr": "sum(label_replace(traefik_entrypoint_request_duration_seconds_sum{instance=~\".*$domain.*\"} or traefik_entrypoint_request_duration_seconds_sum{domain=\"$domain\"}, \"domain\", \"$1\", \"instance\", `^metrics\\.(.+)$`)) / sum(label_replace(traefik_entrypoint_requests_total{instance=~\".*$domain.*\"} or traefik_entrypoint_requests_total{domain=\"$domain\"}, \"domain\", \"$1\", \"instance\", `^metrics\\.(.+)$`)) * 1000",
"format": "time_series",
"interval": "",
"intervalFactor": 2,
@@ -781,7 +781,7 @@
},
"editorMode": "code",
"exemplar": true,
"expr": "sum(delta(traefik_service_requests_total{instance=\"${instance:raw}\"}[$interval]))",
"expr": "sum(label_replace(delta(traefik_service_requests_total{instance=~\".*$domain.*\"}[$interval]) or delta(traefik_service_requests_total{domain=\"$domain\"}[$interval]), \"domain\", \"$1\", \"instance\", `^metrics\\.(.+)$`))",
"format": "time_series",
"interval": "",
"intervalFactor": 2,
@@ -879,7 +879,7 @@
"uid": "PBFA97CFB590B2093"
},
"exemplar": true,
"expr": "sum(rate(traefik_service_requests_total{instance=\"$instance\"}[$interval]))",
"expr": "sum(label_replace(rate(traefik_service_requests_total{instance=~\".*$domain.*\"}[$interval]) or rate(traefik_service_requests_total{domain=\"$domain\"}[$interval]), \"domain\", \"$1\", \"instance\", `^metrics\\.(.+)$`))",
"format": "time_series",
"interval": "",
"intervalFactor": 2,
@@ -984,7 +984,7 @@
},
"editorMode": "code",
"exemplar": true,
"expr": "sum(rate(traefik_service_request_duration_seconds_sum{ instance=\"$instance\" }[5m])) by(service)",
"expr": "sum(label_replace(rate(traefik_service_request_duration_seconds_sum{instance=~\".*$domain.*\"}[5m]) or rate(traefik_service_request_duration_seconds_sum{domain=\"$domain\"}[5m]), \"domain\", \"$1\", \"instance\", `^metrics\\.(.+)$`)) by(service)",
"format": "time_series",
"interval": "",
"intervalFactor": 2,
@@ -1053,11 +1053,11 @@
"uid": "PBFA97CFB590B2093"
},
"exemplar": true,
"expr": "process_open_fds{job=~\"$job\", instance=\"$instance\"}",
"expr": "label_replace(process_open_fds{instance=~\".*$domain.*\"} or process_open_fds{domain=\"$domain\"}, \"domain\", \"$1\", \"instance\", `^metrics\\.(.+)$`)",
"format": "time_series",
"interval": "",
"intervalFactor": 2,
"legendFormat": "{{ instance }}",
"legendFormat": "{{ domain }}",
"refId": "A",
"step": 240
}
@@ -1154,7 +1154,7 @@
"uid": "PBFA97CFB590B2093"
},
"exemplar": true,
"expr": "sum(rate(traefik_service_requests_total{protocol=~\"http|https\", instance=\"$instance\"}[$interval])) by (service)",
"expr": "sum(label_replace(rate(traefik_service_requests_total{protocol=~\"http|https\", instance=~\".*$domain.*\"}[$interval]) or rate(traefik_service_requests_total{protocol=~\"http|https\", domain=\"$domain\"}[$interval]), \"domain\", \"$1\", \"instance\", `^metrics\\.(.+)$`)) by (service)",
"format": "time_series",
"interval": "",
"intervalFactor": 2,
@@ -1255,7 +1255,7 @@
"uid": "PBFA97CFB590B2093"
},
"exemplar": true,
"expr": "sum(traefik_entrypoint_open_connections{instance=\"$instance\"}) by (method)",
"expr": "sum(label_replace(traefik_entrypoint_open_connections{instance=~\".*$domain.*\"} or traefik_entrypoint_open_connections{domain=\"$domain\"}, \"domain\", \"$1\", \"instance\", `^metrics\\.(.+)$`)) by (method)",
"format": "time_series",
"interval": "",
"intervalFactor": 1,
@@ -1355,7 +1355,7 @@
"uid": "PBFA97CFB590B2093"
},
"exemplar": true,
"expr": "sum(traefik_service_open_connections{instance=\"$instance\"}) by (method)",
"expr": "sum(label_replace(traefik_service_open_connections{instance=~\".*$domain.*\"} or traefik_service_open_connections{domain=\"$domain\"}, \"domain\", \"$1\", \"instance\", `^metrics\\.(.+)$`)) by (method)",
"format": "time_series",
"interval": "",
"intervalFactor": 1,
@@ -1459,7 +1459,7 @@
"uid": "PBFA97CFB590B2093"
},
"exemplar": true,
"expr": "sum(increase(traefik_service_requests_total{protocol=~\"$protocol\", instance=\"$instance\"}[$interval])) by (code)",
"expr": "sum(label_replace(increase(traefik_service_requests_total{protocol=~\"$protocol\", instance=~\".*$domain.*\"}[$interval]) or increase(traefik_service_requests_total{protocol=~\"$protocol\", domain=\"$domain\"}[$interval]), \"domain\", \"$1\", \"instance\", `^metrics\\.(.+)$`)) by (code)",
"format": "time_series",
"interval": "",
"intervalFactor": 2,
@@ -1514,36 +1514,6 @@
],
"templating": {
"list": [
{
"current": {
"selected": false,
"text": "default",
"value": "default"
},
"datasource": {
"type": "prometheus",
"uid": "PBFA97CFB590B2093"
},
"definition": "",
"hide": 0,
"includeAll": false,
"label": "Job:",
"multi": false,
"name": "job",
"options": [],
"query": {
"query": "label_values(job)",
"refId": "Prometheus-job-Variable-Query"
},
"refresh": 1,
"regex": "",
"skipUrlSync": false,
"sort": 2,
"tagValuesQuery": "",
"tagsQuery": "",
"type": "query",
"useTags": false
},
{
"current": {
"selected": true,
@@ -1569,7 +1539,7 @@
"query": "label_values(traefik_service_requests_total, protocol)",
"refId": "Prometheus-protocol-Variable-Query"
},
"refresh": 1,
"refresh": 2,
"regex": "",
"skipUrlSync": false,
"sort": 0,
@@ -1663,19 +1633,19 @@
"type": "prometheus",
"uid": "PBFA97CFB590B2093"
},
"definition": "label_values(instance)",
"definition": "query_result(label_replace(traefik_config_reloads_total, \"domain\", \"$1\", \"instance\", `^metrics\\.(.+)$`))",
"hide": 0,
"includeAll": false,
"label": "Instance:",
"label": "Domain:",
"multi": false,
"name": "instance",
"name": "domain",
"options": [],
"query": {
"query": "label_values(instance)",
"query": "query_result(label_replace(traefik_config_reloads_total, \"domain\", \"$1\", \"instance\", `^metrics\\.(.+)$`))",
"refId": "PrometheusVariableQueryEditor-VariableQuery"
},
"refresh": 1,
"regex": ".*8082",
"refresh": 2,
"regex": "/domain=\"([^\"]+)\"/",
"skipUrlSync": false,
"sort": 1,
"tagValuesQuery": "",
-11
View File
@@ -1,11 +0,0 @@
#!/bin/sh -e
NODE_NAME=$(cat /etc/nodename)
mkdir -p /etc/node-exporter
echo "node_meta{node_id=\"$NODE_ID\", container_label_com_docker_swarm_node_id=\"$NODE_ID\", node_name=\"$NODE_NAME\"} 1" > /etc/node-exporter/node-meta.prom
set -- /bin/node_exporter "$@"
exec "$@"
-37
View File
@@ -1,37 +0,0 @@
server:
http_listen_port: 9080
grpc_listen_port: 0
positions:
filename: /tmp/positions.yaml
clients:
- url: {{ env "LOKI_PUSH_URL" }}
basic_auth:
username: admin
password: {{ secret "basic_auth" }}
external_labels:
hostname: {{ env "DOMAIN" }}
scrape_configs:
- job_name: system
static_configs:
- targets:
- localhost
labels:
job: varlogs
__path__: /var/log/*log
- job_name: "docker"
docker_sd_configs:
- host: "unix:///var/run/docker.sock"
refresh_interval: "10s"
relabel_configs:
- source_labels: ['__meta_docker_container_name']
target_label: "container_name"
- source_labels: ['__meta_docker_container_id']
target_label: "container_id"
- source_labels: ['__meta_docker_container_label_com_docker_stack_namespace']
target_label: "stack_namespace"
- source_labels: ['__meta_docker_container_label_com_docker_swarm_service_name']
target_label: "service_name"
+56 -7
View File
@@ -1,12 +1,61 @@
1. OIDC was moved into a seperate compose file. If you have oidc configured you need to add the following line to you .env file:
BREAKING CHANGE
Migration plan for upgrading from 1.6.0+v1.8.1.
COMPOSE_FILE="$COMPOSE_FILE:compose.grafana-oidc.yml"
## 1. Reinsert secrets with shortened names
2. SMTP was moved into a seperate compose file. If you have smtp configured you need to add the following line to you .env file:
Secret and config names were shortened to max 14 characters to prevent going over Docker's 64 character
limit when STACK_NAME and VERSION are added to it.
COMPOSE_FILE="$COMPOSE_FILE:compose.grafana-smtp.yml"
- `abra app secret list <domain>` to see which secrets are missing under their new name
- `abra app cmd --local <domain> migrate_secret_names` to reinsert all of them automatically
(or manually: `abra app secret insert <domain> <secret_name> v1 <value>` per secret)
3. The scrape-config.example.yml file and add_node() command were updated to use a secure endpoint for the traefik metrics instead of http. This requires an updated Traefik recipe that publishes the metrics on https.
## 2. If you use OIDC (moved to seperate compose file)
4. Secret and config names were shortened to max 14 characters to prevent going over Docker's 64 character limit when STACK_NAME and VERSION are added to it.
When upgrading, you need to reinsert the secrets with their shorter names. Run `abra app secret list <domain>` to see which secrets aren't created on the server (because their name was shortened) and run `abra app secret insert <domain> <secret_name> v1 <value>` to reinsert them with the shorter name. Or you can use the migrate_secret_names function in abra.sh to reinsert all existing secrets with their shorter name automatically: `abra app cmd --local <domain> migrate_secret_names`
- Add to your .env: `COMPOSE_FILE="$COMPOSE_FILE:compose.grafana-oidc.yml"`
## 3. If you use SMTP (moved to a seperate compose file)
- Add to your .env: `COMPOSE_FILE="$COMPOSE_FILE:compose.grafana-smtp.yml"`
## 4. node_exporter/cadvisor/promtail replaced by Grafana Alloy
Metrics collection changed from Prometheus scraping endpoints
to Alloy pushing via `remote_write`/`loki push`.
- Remove `compose.promtail.yml`, `compose.expose-ports.yml` and
`compose.basic-auth.yml` from your .env if present. They no
longer exist. `compose.yml` now declares the `basic_auth` secret directly, so
`SECRET_BASIC_AUTH_VERSION` is always required.
- Add `PROMETHEUS_REMOTE_WRITE_URL=https://prometheus.$DOMAIN/api/v1/write`
(`$DOMAIN` if this host also runs `compose.prometheus.yml`, otherwise a remote
Prometheus' URL). Without this, Alloy collects no metrics at all.
- Add `LOKI_PUSH_URL` (existing var, still used) and pick a log source:
`JOURNALD=1` (systemd hosts), `SYSLOG_FILES=1` (non-systemd, tails
`/var/log/*log`), or `SYSLOG=1` + `compose.syslog.yml` (network syslog listener).
- If this host had its own `node.$DOMAIN`/`cadvisor.$DOMAIN` scrape target
configured on a central Prometheus, remove it. Those endpoints are gone.
- `scrape-config.example.yml` and the `add_node`/`add_domain` abra.sh commands
are gone. Replaced by label-based auto-discovery (see README).
- `docker stack deploy` doesn't prune removed services, so old `cadvisor`/
`promtail` containers keep running after a normal `abra app deploy`. Run
`abra app undeploy <domain>` then `abra app deploy <domain>` to clear them out.
- Diff your `.env` against the current `.env.sample`, to verify any other changes.
### New: label-based metrics auto-discovery
Alloy now auto-discovers and scrapes other Docker Swarm services on the same
host/`proxy` network that opt in via `prometheus.io/scrape=true` deploy labels.
See the README's "Auto-discovering metrics from other apps" section.
- If you scrape Traefik metrics: the old `metrics.traefik.$domain` pull-based
endpoint still works if you keep the scrape config in Prometheus and
existing dashboards keep showing its data, but it's recommended to get
Traefik onto the new label-based discovery.
### Dashboards
The Swarm, Stacks and Traefik dashboards were reworked to show old (pull-model)
and new (Alloy push-model) data as one continuous line, so you don't lose history
across the migration.
-4
View File
@@ -1,4 +0,0 @@
- targets:
- 'metrics.traefik.example.org'
- 'node.monitoring.example.org'
- 'cadvisor.monitoring.example.org'