Observability is not a dashboard hobby. It is how you answer new questions about production without SSH folklore. If you cannot explain why latency spiked for checkout in the last fifteen minutes—across services—you do not have observability yet. You have charts.
The three signals (and what each is for)
| Signal | Best for | Weak when… |
|---|---|---|
| Metrics | SLOs, alerting, capacity | You need per-request context |
| Logs | Discrete events, errors, audit | Unstructured text, no IDs |
| Traces | Cross-service latency paths | Missing instrumentation at edges |
User request
│
├─ metrics: rate / errors / duration (aggregate)
├─ logs: structured events + request_id
└─ traces: spans across API → service → DBMetrics: start with RED and USE
For request-driven services, track RED:
- Rate — requests per second
- Errors — failed requests (by class)
- Duration — latency distribution (p50/p95/p99), not only averages
For hosts and stateful dependencies, track USE:
- Utilization — CPU, connection pools, disk
- Saturation — queue depth, thread pool wait
- Errors — disk faults, refused connections
Logs: structured or it did not happen
Emit JSON (or equivalent) with stable fields: timestamp, level, service, request_id, user_id (if safe), error_code. Never log secrets or full card numbers.
logger.info("order_created", {
request_id: ctx.requestId,
order_id: order.id,
latency_ms: Date.now() - start,
outcome: "ok",
});Traces: spans across boundaries
Propagate context (traceparent / vendor equivalent) through HTTP, queues, and workers. A trace without the database span hides the actual bottleneck half the time.
Minimum SLO set (week-one scope)
Do not SLO every microservice on day one. Pick top user journeys:
- Login / session establish
- Primary read path (feed, search, dashboard)
- Primary write path (checkout, submit, upload)
For each journey define:
- Availability — successful responses / total valid requests
- Latency — e.g. p95 under a budget that matches UX
- Error budget — how much unreliability you can “spend” per window
Warning: Alerting on every CPU spike creates fatigue. Alert on user-impacting SLO burn, then use metrics/logs/traces to diagnose.
Starter checklist (copy into your runbook)
- RED dashboard per critical service
- USE panels for DB, cache, queue, and ingress
- Structured logs with
request_idon every hop - Trace sampling that still catches errors (tail-based if available)
- Exemplars or deep links from a metric spike → traces
- Alert: high burn rate on journey SLOs (not raw pod restarts alone)
- Runbook link inside the alert page
- PII scrubbing reviewed once with security
Text diagram: good vs theater
THEATER USEFUL
──────── ──────
20 pretty graphs 3 journey SLOs
Alert: CPU > 70% Alert: 2h error budget 50% burned
Logs: free-text dump Logs: JSON + request_id
Trace: 1 service only Trace: edge → app → DB → cache
On-call guesses On-call follows exemplarsReal ops mistakes (steal these postmortems)
- Average latency alerts — hide the p99 pain users feel.
- No cardinality discipline — high-cardinality labels (
user_idon every metric) melt Prometheus/vendor bills and query speed. - Logging inside hot loops at info — you DDOS your own log pipeline during incidents.
- Traces without deployment version — you cannot tell if a bad release caused the cliff.
- Separate tools, zero correlation IDs — metrics say “error↑”, logs say nothing joinable.
- Alerting on dependencies you do not own without a customer symptom — pager noise, no action.
# Example: alert on burn, not vanity
# (pseudoconfig — adapt to your stack)
alert: CheckoutLatencyBurn
expr: slo:checkout_p95_burn_rate_1h > 2
for: 10m
annotations:
runbook: https://wiki.example/runbooks/checkout-latency
summary: Checkout p95 burning error budgetInstrumentation sketch (HTTP middleware)
app.use(async (req, res, next) => {
const end = httpRequestDuration.startTimer({
route: req.route?.path ?? "unknown",
method: req.method,
});
try {
await next();
httpRequests.inc({ status: String(res.statusCode) });
} finally {
end({ status: String(res.statusCode) });
}
});Keep route cardinality bounded—templatize paths (/orders/:id, not /orders/12345).
How to roll this out in five days
| Day | Deliverable |
|---|---|
| 1 | List journeys + draft SLOs |
| 2 | RED metrics on the edge service |
| 3 | Structured logging + request IDs |
| 4 | Trace one critical path end-to-end |
| 5 | Replace two noisy alerts with one burn alert |
Incident workflow that uses the three signals
When the pager fires on checkout latency burn:
1. Metrics: which dependency saturated? (DB CPU, Redis evictions, queue lag)
2. Traces: which span ate the p95? (app code vs downstream)
3. Logs: which error_code / exception class spiked with that trace id?
4. Change: what deployed, flagged, or scaled in the burn window?If your tools cannot jump from alert → dashboard → trace → log line in under two minutes, fix navigation before buying another visualization product. Store deployment events and flag changes as first-class annotations on the same timeline as metrics.
Sampling without lying to yourself
Head-based sampling at 1% is fine for latency landscapes but will miss rare errors. Prefer tail-based sampling that keeps error and high-latency traces at much higher rates. Always keep metrics as the complete census for SLOs; traces are the microscope, not the scoreboard.
Closing
Observability is a production contract: when something breaks, a competent engineer can follow signals to a cause without tribal knowledge. Add one RED dashboard per critical service this week, wire request IDs everywhere, and delete one alert that has never driven a useful action. That trio beats any vendor slide deck.
