Files
cls 9dfa06ffee
Docker image / Build (linux/amd64) (push) Has been cancelled
Docker image / Build (linux/arm64) (push) Has been cancelled
Docker image / Merge release multi-arch manifest (push) Has been cancelled
Docker image / Merge debug multi-arch manifest (push) Has been cancelled
Docker image / Build public push gateway (linux/amd64) (push) Has been cancelled
Docker image / Build public push gateway (linux/arm64) (push) Has been cancelled
Docker image / Publish public push gateway image (push) Has been cancelled
Sprig image / Build (linux/amd64) (push) Has been cancelled
Sprig image / Build (linux/arm64) (push) Has been cancelled
Sprig image / Merge multi-arch manifest (push) Has been cancelled
Harbor Buzz Orchestra / Python tests and lint (push) Has been cancelled
CI / Detect Changed Paths (push) Has been cancelled
CI / Rust Lint (push) Has been cancelled
CI / Unit Tests (push) Has been cancelled
CI / Desktop Core (push) Has been cancelled
CI / Desktop Smoke E2E (1) (push) Has been cancelled
CI / Desktop Smoke E2E (2) (push) Has been cancelled
CI / Desktop Smoke E2E (3) (push) Has been cancelled
CI / Desktop Smoke E2E (4) (push) Has been cancelled
CI / Desktop (push) Has been cancelled
CI / Desktop E2E Relay (push) Has been cancelled
CI / Desktop E2E Integration (1/2) (push) Has been cancelled
CI / Desktop E2E Integration (2/2) (push) Has been cancelled
CI / Desktop E2E Integration (push) Has been cancelled
CI / Backend Integration (relay e2e) (push) Has been cancelled
CI / Relay E2E (push) Has been cancelled
CI / Web (push) Has been cancelled
CI / Mobile (push) Has been cancelled
CI / Security (push) Has been cancelled
CI / Dead Token Reference Guard (push) Has been cancelled
CI / Server Cross-Compile (aarch64-unknown-linux-musl) (push) Has been cancelled
CI / Server Cross-Compile (x86_64-unknown-linux-musl) (push) Has been cancelled
CI / Windows Rust (x86_64-pc-windows-msvc) (push) Has been cancelled
CI / Desktop Build (macOS) (push) Has been cancelled
helm chart / lint + unittest + render matrix (push) Has been cancelled
helm chart / install on kind (gated) (push) Has been cancelled
helm chart / publish chart to GHCR (push) Has been cancelled
Mesh Lifecycle / Relay-Driven Mesh Lifecycle Smoke (push) Has been cancelled
Sprig / Build (aarch64-unknown-linux-musl) (push) Has been cancelled
Sprig / Build (x86_64-unknown-linux-musl) (push) Has been cancelled
Sprig / Publish rolling release (push) Has been cancelled
Sprig / Publish tagged release (push) Has been cancelled
feat: import Chinese-localized Buzz source snapshot
Signed-off-by: cls_宁波本机 <908705107@qq.com>
2026-08-13 18:34:25 +08:00

90 lines
4.4 KiB
YAML

{{- if .Values.prometheusRule.enabled }}
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: {{ include "push.name" . }}
labels: {{- include "push.labels" . | nindent 4 }}
{{- with .Values.prometheusRule.labels }}
{{- toYaml . | nindent 4 }}
{{- end }}
spec:
groups:
- name: buzz-push-gateway
rules:
# Sustained configuration faults mean the provider credential/topic is
# unhealthy; no endpoint is being invalidated but nothing is delivering.
- alert: PushGatewayConfigurationFault
expr: |
sum(rate(push_gateway_apns_deliveries_total{outcome="configuration_fault"}[5m])) > 0
for: 10m
labels: { severity: critical }
annotations:
summary: Push gateway APNs configuration faults
description: >-
APNs is returning configuration faults (bad/expired provider token
or topic). Deliveries are failing without invalidating endpoints.
See runbook: check the APNs .p8 key, key id, team id, and topic.
# Authority store unavailable at admission = durable dependency is down.
- alert: PushGatewayAdmissionUnavailable
expr: |
sum(rate(push_gateway_admissions_total{result="unavailable"}[5m])) > 0
for: 5m
labels: { severity: critical }
annotations:
summary: Push gateway authority store unavailable
description: >-
authorize_delivery is returning Unavailable — the PostgreSQL
authority store is unreachable or failing. Check DB connectivity
and the pod's postgres egress NetworkPolicy.
# Readiness failing on the authority cause = the pod will be pulled from
# rotation; alert before all replicas drop out.
- alert: PushGatewayReadinessAuthorityFailing
expr: |
sum(rate(push_gateway_readiness_failures_total{cause="authority"}[5m])) > 0
for: 5m
labels: { severity: warning }
annotations:
summary: Push gateway readiness failing on authority
description: >-
Readiness probes are failing because the authority store check
fails. Replicas will be removed from the Service. Investigate DB
health before capacity drops below the PodDisruptionBudget.
# The retention reaper sweeps expired rows every 5m; a single transient
# failure self-heals on the next tick. Alert on repeated failure —
# at least two sweeps failing within ~30m (six ticks) — which grows the
# bounded crash-before-release window and leaks storage.
- alert: PushGatewayReaperFailing
expr: |
sum(increase(push_gateway_reaper_failures_total[30m])) >= 2
for: 5m
labels: { severity: warning }
annotations:
summary: Push gateway retention reaper failing
description: >-
The retention reaper has failed at least twice within 30m (it runs
every 5m). Expired delivery reservations are not being swept,
growing the bounded-until-expiry window. Check DB write availability.
# High sustained fraction of retryable APNs outcomes indicates APNs
# throttling or degradation. The ratio is a true fraction over the
# window (increase = counts, not per-second rate), gated by a minimum
# sample count so a couple of retries at trivial volume cannot trip it.
- alert: PushGatewayHighApnsRetryRate
expr: |
(
sum(increase(push_gateway_apns_deliveries_total{outcome="retry"}[10m]))
/ sum(increase(push_gateway_apns_deliveries_total[10m]))
> {{ .Values.prometheusRule.apnsRetryRatioThreshold }}
)
and
sum(increase(push_gateway_apns_deliveries_total[10m])) >= {{ .Values.prometheusRule.apnsRetryMinSamples }}
for: 15m
labels: { severity: warning }
annotations:
summary: Push gateway high APNs retry ratio
description: >-
The retryable fraction of APNs attempts over a 10m window
(429/500/503), above a minimum sample count, has exceeded the
configured threshold continuously for 15m. APNs may be throttling
or degraded; deliveries are delayed but not lost.
{{- end }}