Production Alerting
Origin: data#1925, data#1926, data#1927 (2026-08-17/18). All three incidents were silent. Nothing paged; each was found by hand. This runbook covers what alerting exists, how to activate delivery, what each rule maps onto, and — explicitly — what is still not covered.1. The stack, as it actually is
The single most important fact: because no Alertmanager is installed,
writing a Prometheus rule delivers nothing on its own. Before this work, all
26 existing rules could fire and reach no human. alert-dispatch.sh is the
bridge that makes them real. It is a stopgap for Alertmanager, not a
replacement — no grouping, inhibition, silencing, or escalation.
How an alert reaches a human today
2. Operator activation steps
Nothing below contains a secret. Both webhook values are bearer credentials and must never be committed.Step 1 — Slack webhook on the compute host (REQUIRED)
Without this, every host-side alert is a journal line and nothing more.SLACK_WEBHOOK_URL is NOT set or
Slack webhook POST FAILED in the output means it does not.
Step 2 — GitHub webhook for deploy failures (REQUIRED)
synthetic-auth-probe.yml,
clickhouse-storage-canary.yml and pattern-corpus-canary.yml. If it is
already set, deploy-failure alerting works with no further action.
Prove it without breaking anything:
:test_tube: DRILL message.
Step 3 — Reload Prometheus so the new rules load (REQUIRED)
New rules inmonitoring/prometheus-rules/alerts.yml are inert until
Prometheus re-reads them. The rule files are not copied to the host by CI
— confirm where the host reads them from before reloading.
Step 4 — Verify the timers are armed (after the next deploy)
sudo systemctl enable --now alert-dispatch.timer alert-service-health.timer.
Step 5 (recommended, not done here) — install Alertmanager
alert-dispatch.sh has no silencing. During a long incident it will notify
once per alert and once on resolution, but there is no way to mute a known
issue short of stopping the timer. Installing Alertmanager and pointing
Prometheus at it is the durable fix; alert-dispatch.timer should be stopped
at that point to avoid duplicate notifications.
3. Per-rule signal mapping
Every rule below maps to a signal that is exported today.scripts/check-alert-metrics.py
(wired into data-ci.yml) fails CI if any alert references a metric nothing
emits — a dead alert reads as coverage it does not provide (#1673).
Added by this change
Threshold note for the broker rules: data#1926 logged 118,594 401s in 32h
(~1.03/s). The 0.1/s floor is two orders of magnitude below the observed rate
and far above the zero a healthy credential produces.
Non-Prometheus checks
4. Recovery: services stuck failed (data#1925)
The canonical case. A data-server power loss left Redis loading an 18.4GB RDB
for ~25 minutes; every Go service exited on Redis LOADING, exhausted
StartLimitBurst=10, and stayed failed permanently.
5. GAPS — not covered today
Recorded honestly rather than papered over with rules that cannot fire.5.1 data#1927 (missing-relation query errors) is NOT specifically covered
pattern-detector queried two non-existent tables 27,362 times in 32h. There is no metric for it, and no log-based alerting exists at all — no Loki, no promtail, no Vector, no log shipper of any kind. The journal is the only record.PatternDetectorCycleErrors (pd_cycle_errors > 0) is real coverage for
cycle-level failure but is not a missing-relation detector; those queries
are not known to increment it.
To close this properly, one of:
- A metric at the query site (preferred, ~1 day). A
pd_query_errors_total{table,code}counter incremented where ClickHouse/ FalkorDB errors are handled, then a rule onrate(pd_query_errors_total{code="UNKNOWN_TABLE"}[10m]) > 0. Cheapest, most precise, and passescheck-alert-metrics.pyby construction. - A schema-conformance CI check. Assert every table named in Go query
strings exists in
clickhouse/*.sql. Catches the class before deploy, but not at runtime. - Log-based alerting (largest). Deploy Loki + promtail on compute, add a Grafana/Loki ruler. Also unlocks the whole “repeated error in the journal” class, including data#1926’s 401 log lines.
5.2 Other broker call sites have no auth metrics
Only the options-chain path haste_options_chain_auth_*. Other Alpaca call
sites (order placement, positions, account, market data) have no auth-failure
metrics, so a retired credential on any of those paths is still silent. The
te_options_chain_auth_* pattern should be generalised to a shared
broker_api_errors_total{endpoint,status} counter. Not invented here — the
metric does not exist yet, and a rule on it would be exactly the dead alert
check-alert-metrics.py exists to prevent.
5.3 backfill-api and daily-daemon have no Prometheus coverage
backfill-api exposes no /metrics endpoint at all; daily-daemon’s
scrape job label is not established anywhere in this repo. Both are covered by
alert-service-health.sh (systemd state) but by no Prometheus rule. Adding
a /metrics handler to backfill-api would close half of this.
5.4 The Prometheus config is not in version control
There is noprometheus.yml in this or any repo. Scrape targets, job labels
and rule_files: are host-managed and unreviewable. A target silently dropped
from the scrape config is invisible to code review —
CriticalServiceTargetMissing is a partial mitigation for the four services
whose job labels are known from slo.yml, but the durable fix is to bring the
Prometheus config into the repo and deploy it via CI.
5.5 No silencing, grouping, or escalation
alert-dispatch.sh notifies once per alert and once on resolution. There is no
way to mute a known-ongoing issue except stopping the timer, and no escalation
if Slack is not being read. Alertmanager is the fix (§2 Step 5).
5.6 Delivery is single-channel and unmonitored
Everything lands in one Slack webhook. If the webhook is revoked or the channel is archived,alert-notify.sh logs the failure at user.err — but nothing
alerts on that, and it cannot: the alerting path is the thing that is broken.
A periodic heartbeat to Slack (visible absence rather than an alert) would
close it.
6. Host-side exporter state that CI cannot deploy (2026-09-11, data#2589)
Two exporter changes were made by hand on the data server. Neither unit is in this repo, so nothing re-applies them: a host rebuild loses both, and the loss is silent. This section exists so the next person looking for them finds them.6.1 postgres-exporter now authenticates over the Unix socket
pg_up had been 0 since the 2026-08-17 PSU outage, so PostgresDown fired
continuously for over three weeks while PostgreSQL itself served the API
normally. The exporter was running and reachable; its stored DSN password had
gone stale and every scrape logged password authentication failed for user "sequency".
/etc/default/postgres-exporter (root-owned, mode 600) now holds:
User=sequency and pg_hba.conf carries local all all peer, so a socket connection needs no credential at all. This removes a stored
secret rather than rotating one. The previous file is kept at
/etc/default/postgres-exporter.bak-2026-09-11.
Further improvement, not done: the exporter connects as the application role. A
dedicated role with pg_monitor would be least-privilege, at the cost of a role
to create and a password to store — which is what the socket change just
eliminated, so it is a genuine trade rather than a clear win.
6.2 node-exporter now reports systemd unit state
Until this change nothing alerted when a backup job failed. The four backup units on the data host carry noOnFailure= handler, no live rule covered failed
systemd units, and alert-service-health.sh watches six long-running services on
the other host. A silent backup failure on a single-host data tier that already
lost power once this year is the gap with the most to lose.
/etc/systemd/system/node-exporter.service (root-owned) now runs:
.+ would export every unit on the
box. Ten units produce 50 series, one per unit per state. The previous file is at
/etc/systemd/system/node-exporter.service.bak-2026-09-11.
BackupJobFailed and DataHostUnitFailed in monitoring/prometheus-rules/alerts.yml
read these series. Adding a unit to either alert’s regex requires adding it to
the exporter’s include list too, or the series will not exist and the rule will
match nothing. BackupJobFailed carries an absent() guard for exactly that
mistake; DataHostUnitFailed does not, because those services have independent
coverage through their own scrape targets.
6.3 What is still not covered
A backup that never runs is as bad as one that fails, and neither rule catches it: a oneshot that is never triggered simply staysinactive. pgBackRest exposes
no last-successful-backup metric here, and scripts/pgbackrest-check.sh is not on
a timer. Backup recency therefore remains unmonitored. pg_stat_archiver does
now flow, so continuous WAL archiving failure is covered even though backup
recency is not.
6.4 The API metrics scrape presents the internal secret
/metrics had returned 401 to Prometheus since 2026-06-07, so no API metric
existed and ServiceDown{job="sequency-api"} fired for three months (data#2589).
It was tempting to exempt /metrics from the auth middleware. That would have
been an exposure: /metrics is reachable from the internet through nginx, and
only the 401 was protecting it, so an exemption would have published every
internal gauge and endpoint label. Verified at the time:
http_headers in a scrape config, so the
scraper presents the secret instead and the endpoint’s auth posture is
unchanged. /opt/prometheus/prometheus.yml, sequency-api job:
/opt/prometheus/api-scrape-secret (mode 600, owned by the
Prometheus user) rather than inline, because prometheus.yml is mode 644. Neither
file is in this repo — prometheus.yml is host-managed and the secret must not be
committed — so a host rebuild loses this and the API target silently goes down
again. promtool check config was run before the reload; do the same on any edit,
because an invalid config makes the reload a no-op and leaves stale rules loaded.
The previous config is at /opt/prometheus/prometheus.yml.bak-2026-09-11.
6.5 Two exit-protection alerts are suspended, for different reasons
Suspended inalerts.yml on 2026-09-14 by operator decision (data#2593):
PositionsWithoutExitProtection— paged on a legitimate state. Exit protection is not mandatory for discretionary trading (data#2094). Restore when data#2094 Gap 2 derivesexit_intentfrom the strategy that placed the order.ExitProtectionBrokerStateUnknownwas suspended in the first version of that change and has been kept, for two reasons. CI pins it:test_unknown_broker_state_metric_has_multiprocess_safe_paging_ruleasserts its exact expression,severity: criticalandfor: 2m, so removing it fails the build — a deliberate guard. And its cause is not the discretionary question: it means broker-resident protection could not be read, which is a real failure. Today that is the rejected broker credential, the same one that has 9 of 11 positions auth-suppressed, so its blocker is the BYOK credential substrate in data#2037 rather than data#2094.
ExitProtectionReconcilerStale remains active and is the signal that matters
while those two are off: it fires if the sweep stops running or its gauge
disappears.
Note for anyone tempted to downgrade rather than suspend a noisy alert:
alert-dispatch.sh reads severity only to label the Slack message and does not
filter on it. A warning notifies exactly as loudly as a critical.
6.6 clickhouse-sink is now a scrape target
The sink exportssink_consumer_lag, sink_dead_letters_total,
sink_retry_queue_size and sink_rows_written_total on :8110, and nothing
scraped it until 2026-09-14 (data#2613). Added to /opt/prometheus/prometheus.yml:
clickhouse-sink was added to
CriticalServiceTargetMissing for the same reason: prometheus.yml is
host-managed and not in this repo, so a rebuild drops this target, and losing it
would put the deadline back out of sight.
ServiceDown already covers the sink now that it is scraped — that rule has no
job filter.
Two alerting choices worth recording so nobody “fixes” them:
- No rule on rows-written. Measured 2026-09-15: the four history tables take
over a million rows each per session, then legitimately go silent from the
session boundary to the next open, with Sunday at zero. A
rate()==0rule would page every night and every weekend. - No rule on the LEVEL of
sink_consumer_lag. It sits at a static ~1.93M pending backlog (data#2613), so any absolute threshold either fires forever or is meaningless.SinkConsumerLagGrowingwatches the hourly delta instead; measured drift is ~500/hour against a 50,000 threshold, and a stalled sink would grow at roughly 425k/hour. The metric is recomputed fromXPENDINGeach cycle and was verified stable across a sink restart, so theoffset 1hcomparison is restart-safe.