Queries, alerts, and objectives
Stage 6 · Mission Operations
Storing time-series samples does not guarantee operational awareness. Operators must turn raw metrics into actionable indicators and alert rules without drowning in notification fatigue.
Core PromQL query patterns
- Rate calculations (handling counter resets):
rate(booking_requests_total{status="200"}[5m])
- Error percentage ratio:
sum(rate(booking_requests_total{status=~"5.."}[5m]))/sum(rate(booking_requests_total[5m])) * 100
- 99th percentile latency:
histogram_quantile(0.99, sum by(le) (rate(booking_request_duration_seconds_bucket[5m])))
Alert rules: routing signals, not performing repairs
An alert rule continuously evaluates a PromQL expression. When the expression evaluates to true for longer than the for duration, Prometheus fires an alert to Alertmanager:
groups:
- name: apollo.booking
rules:
- alert: BookingErrorRateHigh
expr: |
sum(rate(booking_requests_total{status=~"5.."}[5m]))
/
sum(rate(booking_requests_total[5m])) > 0.05
for: 2m
labels:
severity: warning
annotations:
summary: "Booking error rate exceeds 5%"
Operational rules for alerts:
- Alerts route messages: An alert notifies humans (PagerDuty/Slack); it does not automatically scale or restart containers.
- The
forduration dampens flapping: Requiring a 2-minute sustained breach prevents momentary transient spikes from paging engineers at night.
SLIs, SLOs, and Error Budgets
Diagram OB-05 — error budgets balance deployment velocity against system reliability.
- Service Level Indicator (SLI): A quantifiable metric measuring service performance:
- Example: Percentage of booking requests returning HTTP 2xx in under 500ms over a 28-day rolling window.
- Service Level Objective (SLO): The agreed target reliability goal:
- Example: 99.5% success rate.
- Error Budget: The permitted margin of failure (
100% - SLO):- Example: 0.5% failure allowance. If depleted, feature releases pause in favor of reliability engineering.
Evidence and limits
- 1. Active Prometheus alert rules:
kubectl get prometheusrules -n apollo-observability
- 2. Alertmanager status: Check routing and silencing configurations:
kubectl get pods -n apollo-observability -l app=alertmanager