Observability Starter Checklist

Technology

Observability Starter Checklist

Ship production-ready metrics, structured logs, and traces with a RED/USE baseline, SLO-first alerts, correlation IDs, and the ops mistakes that create dashboard theater instead of debuggability.

5 min read·Updated June 20, 2026
RD

SRE Lead

Share this guide

Observability is not a dashboard hobby. It is how you answer new questions about production without SSH folklore. If you cannot explain why latency spiked for checkout in the last fifteen minutes—across services—you do not have observability yet. You have charts.

The three signals (and what each is for)

SignalBest forWeak when…
MetricsSLOs, alerting, capacityYou need per-request context
LogsDiscrete events, errors, auditUnstructured text, no IDs
TracesCross-service latency pathsMissing instrumentation at edges
text
User request
   │
   ├─ metrics: rate / errors / duration (aggregate)
   ├─ logs:    structured events + request_id
   └─ traces:  spans across API → service → DB

Metrics: start with RED and USE

For request-driven services, track RED:

  • Rate — requests per second
  • Errors — failed requests (by class)
  • Duration — latency distribution (p50/p95/p99), not only averages

For hosts and stateful dependencies, track USE:

  • Utilization — CPU, connection pools, disk
  • Saturation — queue depth, thread pool wait
  • Errors — disk faults, refused connections

Logs: structured or it did not happen

Emit JSON (or equivalent) with stable fields: timestamp, level, service, request_id, user_id (if safe), error_code. Never log secrets or full card numbers.

text
logger.info("order_created", {
  request_id: ctx.requestId,
  order_id: order.id,
  latency_ms: Date.now() - start,
  outcome: "ok",
});

Traces: spans across boundaries

Propagate context (traceparent / vendor equivalent) through HTTP, queues, and workers. A trace without the database span hides the actual bottleneck half the time.

Minimum SLO set (week-one scope)

Do not SLO every microservice on day one. Pick top user journeys:

  1. Login / session establish
  2. Primary read path (feed, search, dashboard)
  3. Primary write path (checkout, submit, upload)

For each journey define:

  • Availability — successful responses / total valid requests
  • Latency — e.g. p95 under a budget that matches UX
  • Error budget — how much unreliability you can “spend” per window

Warning: Alerting on every CPU spike creates fatigue. Alert on user-impacting SLO burn, then use metrics/logs/traces to diagnose.

Starter checklist (copy into your runbook)

  • RED dashboard per critical service
  • USE panels for DB, cache, queue, and ingress
  • Structured logs with request_id on every hop
  • Trace sampling that still catches errors (tail-based if available)
  • Exemplars or deep links from a metric spike → traces
  • Alert: high burn rate on journey SLOs (not raw pod restarts alone)
  • Runbook link inside the alert page
  • PII scrubbing reviewed once with security

Text diagram: good vs theater

text
THEATER                         USEFUL
────────                       ──────
20 pretty graphs               3 journey SLOs
Alert: CPU > 70%               Alert: 2h error budget 50% burned
Logs: free-text dump           Logs: JSON + request_id
Trace: 1 service only          Trace: edge → app → DB → cache
On-call guesses                On-call follows exemplars

Real ops mistakes (steal these postmortems)

  1. Average latency alerts — hide the p99 pain users feel.
  2. No cardinality discipline — high-cardinality labels (user_id on every metric) melt Prometheus/vendor bills and query speed.
  3. Logging inside hot loops at info — you DDOS your own log pipeline during incidents.
  4. Traces without deployment version — you cannot tell if a bad release caused the cliff.
  5. Separate tools, zero correlation IDs — metrics say “error↑”, logs say nothing joinable.
  6. Alerting on dependencies you do not own without a customer symptom — pager noise, no action.
text
# Example: alert on burn, not vanity
# (pseudoconfig — adapt to your stack)
alert: CheckoutLatencyBurn
expr: slo:checkout_p95_burn_rate_1h > 2
for: 10m
annotations:
  runbook: https://wiki.example/runbooks/checkout-latency
  summary: Checkout p95 burning error budget

Instrumentation sketch (HTTP middleware)

text
app.use(async (req, res, next) => {
  const end = httpRequestDuration.startTimer({
    route: req.route?.path ?? "unknown",
    method: req.method,
  });
  try {
    await next();
    httpRequests.inc({ status: String(res.statusCode) });
  } finally {
    end({ status: String(res.statusCode) });
  }
});

Keep route cardinality bounded—templatize paths (/orders/:id, not /orders/12345).

How to roll this out in five days

DayDeliverable
1List journeys + draft SLOs
2RED metrics on the edge service
3Structured logging + request IDs
4Trace one critical path end-to-end
5Replace two noisy alerts with one burn alert

Incident workflow that uses the three signals

When the pager fires on checkout latency burn:

text
1. Metrics: which dependency saturated? (DB CPU, Redis evictions, queue lag)
2. Traces: which span ate the p95? (app code vs downstream)
3. Logs: which error_code / exception class spiked with that trace id?
4. Change: what deployed, flagged, or scaled in the burn window?

If your tools cannot jump from alert → dashboard → trace → log line in under two minutes, fix navigation before buying another visualization product. Store deployment events and flag changes as first-class annotations on the same timeline as metrics.

Sampling without lying to yourself

Head-based sampling at 1% is fine for latency landscapes but will miss rare errors. Prefer tail-based sampling that keeps error and high-latency traces at much higher rates. Always keep metrics as the complete census for SLOs; traces are the microscope, not the scoreboard.

Closing

Observability is a production contract: when something breaks, a competent engineer can follow signals to a cause without tribal knowledge. Add one RED dashboard per critical service this week, wire request IDs everywhere, and delete one alert that has never driven a useful action. That trio beats any vendor slide deck.

Share this guide

Comments (…)

Share a thought or question about this guide.

Loading comments…