Skip to main content

Correlating a booking

Stage 6 · Mission Operations

When a customer reports that ticket checkout took 8 seconds and returned an error, an on-call engineer must systematically narrow the problem from cluster-wide health down to specific lines of code.


The 4-step triage methodology​

Diagram OB-07 — moving from fleet scope (metric) to distributed bottleneck (trace) to application context (log) to root cause.

  • Step 1: Metric Scope Check:
    • Determine if the issue affected a single user or all passengers.
    • Query Prometheus p99 latency and 5xx error rates around the incident timestamp.
  • Step 2: Distributed Trace Inspection:
    • Retrieve the Trace ID for the affected booking.
    • Inspect span breakdowns in Grafana Tempo.
    • Identify that the flight/CheckSeat span consumed 7.9 seconds of the 8.0-second total duration.
  • Step 3: Target Service Log Search:
    • Query Loki logs in the flight namespace for that exact Trace ID.
    • Locate explicit log event: context deadline exceeded: db query took > 8000ms.
  • Step 4: Infrastructure Verification:
    • Check flight-db resource utilization and database connection pool saturation metrics in Prometheus.

Practical diagnostic commands​

  • Find trace ID from booking reference in logs:
    kubectl logs -n apollo-airlines-apps deploy/booking | grep "AA-2024-001234" | jq .trace_id
  • Inspect span tree via Tempo API:
    curl -s "http://localhost:3100/api/traces/<trace-id>" | jq '.batches[].scopeSpans[].spans[] | {service: .name, duration_ms: (.endTimeUnixNano - .startTimeUnixNano | . / 1000000), status: .status}'
  • Query flight database connection saturation:
    flight_db_pool_connections_used / flight_db_pool_connections_max * 100