Without observability you learn about outages from a Telegram message. Minimum: JSON logs, golden metrics, alert on 5xx and latency.

Logs
- Structured JSON: traceId, userId, route, duration
- No PII/card data in logs
- Centralize: Loki/ELK/CloudWatch
Metrics
- RED per endpoint
- Business: orders/min, payment success
- Infra: CPU, PG connections
Alerts
Alert on symptoms (error rate), not CPU alone. Runbooks in Notion.

Checklist
- Dashboard per service
- On-call rotation
- Post-incident template
Observability in one day
- Structured JSON logs with traceId
- Dashboard: error rate and p95 on top routes
- One alert on 5xx spike linked to a runbook
- After incidents: update the runbook
