Observability stack for order-tracker

Observability stack for order-tracker

Observability stack for order-tracker

The app stays a single FastAPI process. Telemetry is added only in the container command, so local pytest still talks to the app directly and does not need the collector.

flowchart LR
  app[order-tracker app] -->|OTLP gRPC| collector[OTel Collector]
  collector -->|scrape target :8889| prometheus[Prometheus]
  collector -->|OTLP HTTP| loki[Loki]
  collector -->|OTLP gRPC| tempo[Tempo]
  prometheus --> grafana[Grafana]
  loki --> grafana
  tempo --> grafana

App instrumentation

Add the OpenTelemetry distro, OTLP exporter, FastAPI instrumentation, and logging instrumentation to pyproject.toml, then refresh uv.lock. These are runtime dependencies because Dockerfile installs with uv sync --frozen --no-dev.

Change the image command to wrap uvicorn:

uv run --no-sync opentelemetry-instrument uvicorn app.main:app --host 0.0.0.0 --port 8000

In compose.yaml, point the app at the collector and turn on all three signals:

  • OTEL_SERVICE_NAME=order-tracker
  • OTEL_EXPORTER_OTLP_ENDPOINT=http://otel-collector:4317
  • OTEL_EXPORTER_OTLP_PROTOCOL=grpc
  • OTEL_TRACES_EXPORTER=otlp, OTEL_METRICS_EXPORTER=otlp, OTEL_LOGS_EXPORTER=otlp
  • OTEL_PYTHON_LOGGING_AUTO_INSTRUMENTATION_ENABLED=true
  • OTEL_SEMCONV_STABILITY_OPT_IN=http so HTTP metrics use the stable names

FastAPI auto-instrumentation emits server spans and the http.server.request.duration histogram. The logging bridge ships uvicorn and app logs over OTLP. No application route changes.

Collector and backends

Add config under observability/ and four services plus the collector to compose.yaml. Pin current stable image tags when implementing. Bind published ports to 127.0.0.1, same as the app.

  • Collector (otel/opentelemetry-collector-contrib): OTLP gRPC/HTTP in. Batch traces and logs only. Export metrics with the Prometheus exporter on :8889, logs with otlphttp to http://loki:3100/otlp (the exporter appends /v1/logs), traces with OTLP gRPC to tempo:4317.
  • Prometheus: scrape otel-collector:8889.
  • Loki: filesystem storage, auth_enabled: false, and allow_structured_metadata: true so OTLP ingest is accepted.
  • Tempo: local storage and an OTLP gRPC receiver.
  • Grafana: provision Prometheus, Loki, and Tempo datasources, plus a dashboard file. Default login admin / admin. UI at http://127.0.0.1:3000.

Healthchecks on Prometheus, Loki, Tempo, and Grafana so docker compose up --wait still waits for the backends. The collector image has no shell, so it is started without a Docker healthcheck; the app depends_on it.

Dashboard

Provision a dashboard named Order Tracker requests with two time series over the last 15 minutes, excluding /healthz so the Compose healthcheck does not dominate the charts:

  • Request rate: sum by (http_route) (rate(http_server_request_duration_seconds_count{http_route!="/healthz"}[5m]))
  • Error rate: the same series filtered to http_response_status_code=~"[45]..", broken out by status code

After the stack is up, send a few API calls including a 404, then confirm those series exist in Prometheus. If the installed instrumentation still emits the legacy histogram name, update the dashboard queries to match the live series.

Check

  • uv run --frozen pytest -q still passes.
  • Rebuild with docker compose up --build -d --wait.
  • Open Grafana and confirm both panels show data after the sample requests.