3 MIN READ

Observability for Medusa: Logs, Traces and the Alerts Worth Having

Self-hosting means owning the question 'is checkout working?'. Structured logging, tracing, the four metrics that matter and alerts that do not cry wolf.

BY RAHUL MEHTAUPDATED
Illustration for “Observability for Medusa: Logs, Traces and the Alerts Worth Having” — Observability

On a hosted platform, "is checkout working?" is somebody else's dashboard. Self-hosted, it is yours — and the honest answer is that most teams find out from a customer email.

The instrumentation that prevents that is not elaborate. It is structured logs with a correlation id, traces through the checkout path, four metrics, and alerts on symptoms rather than causes.

Structured logging

src/api/middlewares.tsts
import { defineMiddlewares } from "@medusajs/framework/http"
import { randomUUID } from "node:crypto"

export default defineMiddlewares({
  routes: [
    {
      matcher: "*",
      middlewares: [
        (req: any, res: any, next: any) => {
          // One id per request, echoed to the client so support can quote it.
          const correlationId = req.headers["x-correlation-id"] ?? randomUUID()
          req.correlationId = correlationId
          res.setHeader("x-correlation-id", correlationId)

          const start = Date.now()
          res.on("finish", () => {
            req.scope.resolve("logger").info(
              JSON.stringify({
                correlation_id: correlationId,
                method: req.method,
                path: req.path,
                status: res.statusCode,
                duration_ms: Date.now() - start,
                customer_id: req.auth_context?.actor_id ?? null,
              }),
            )
          })

          next()
        },
      ],
    },
  ],
})

The correlation id is the whole point. Returning it in the response header means a customer support ticket can carry the exact identifier that ties together every log line for that request, across server and worker.

Log JSON. Human-readable logs are pleasant on one instance and useless across six.

Tracing the checkout

Distributed tracing pays for itself on exactly one path: checkout, where a single customer action fans out to a payment provider, a tax engine, a carrier and your database.

src/instrumentation.tsts
import { registerOtel } from "@medusajs/medusa"
import { OTLPTraceExporter } from "@opentelemetry/exporter-trace-otlp-http"

export function register() {
  registerOtel({
    serviceName: "medusa-server",
    exporter: new OTLPTraceExporter({ url: process.env.OTEL_EXPORTER_URL }),
    instrument: {
      http: true,
      workflows: true,
      query: true,
    },
  })
}

With workflow instrumentation on, each workflow step becomes a span. When checkout is slow, the trace tells you whether it was the tax engine or your own query, which is the difference between a fix and a guess.

The metrics worth having

Four technical, four business.

Technical

MetricWhyWarning
p95 latency by routeUsers feel the tail> 1s on store routes
5xx rateSomething is broken> 0.5%
DB connections in useThe first bottleneck> 80% of max
Queue lagBackground work falling behind> 5 min

Business

MetricCatches
Orders per hourEverything, eventually
Checkout completion ratePayment or shipping breakage
Payment failure rate by providerA processor degrading
Cart creation rateStorefront or CDN failure

Business metrics catch failures that technical ones miss entirely. A payment provider silently declining every card returns 200s with a valid error body — your error rate is flat and your orders are zero. Only the orders-per-hour graph tells you.

Alerts that do not cry wolf

Alert on symptoms customers experience:

yaml
- alert: OrdersDroppedToZero
  expr: increase(orders_created_total[30m]) == 0
  for: 30m
  severity: critical
  # Compare against the same window last week, not a constant, so a
  # quiet Tuesday at 4am does not page anyone.

- alert: CheckoutErrorRate
  expr: rate(checkout_errors_total[5m]) / rate(checkout_attempts_total[5m]) > 0.05
  for: 10m
  severity: critical

- alert: PaymentProviderFailing
  expr: rate(payment_failures_total[10m]) / rate(payment_attempts_total[10m]) > 0.2
  for: 10m
  severity: critical

- alert: WorkerQueueLag
  expr: queue_oldest_message_age_seconds > 600
  for: 15m
  severity: warning

Two disciplines that keep alerting useful. Every page must be actionable — if the response is "acknowledge and go back to sleep", delete the alert. And compare against the same window last week, because commerce traffic is deeply seasonal and absolute thresholds page you at 4am on a Tuesday.

Do not alert on CPU. It is a cause, not a symptom, and a Node application under I/O load has low CPU while serving nothing.

Error tracking

Sentry or equivalent, with the correlation id attached:

ts
Sentry.withScope((scope) => {
  scope.setTag("correlation_id", req.correlationId)
  scope.setTag("route", req.path)
  scope.setContext("cart", { id: cartId, item_count: itemCount })
  Sentry.captureException(error)
})

Filter expected errors — declined cards, out-of-stock at completion — into a separate stream. They are customer events, not defects, and leaving them in the exception feed trains everyone to ignore it.

Health checks

/health proves the process is up, which is a low bar. Add a deeper check for your own monitoring:

src/api/admin/health/deep/route.tsts
export async function GET(req: MedusaRequest, res: MedusaResponse) {
  const checks = {
    database: await ping(() => req.scope.resolve("query").graph({
      entity: "region", fields: ["id"], pagination: { take: 1 },
    })),
    redis: await ping(() => req.scope.resolve("cache").set("health", "1", 10)),
  }

  const healthy = Object.values(checks).every(Boolean)
  res.status(healthy ? 200 : 503).json({ healthy, checks })
}

Keep it off the load balancer's check — a slow dependency should not take healthy instances out of rotation. Poll it from monitoring instead.

A synthetic checkout

The single highest-value monitor: a scheduled job that runs a real purchase against production every fifteen minutes — create cart, add item, set address, select shipping, authorise with a test card, complete, then cancel the order.

It exercises every external dependency in the real configuration and tells you checkout is broken before a customer does. In our experience it catches more real incidents than every other alert combined.

Cost of observability

Instrumentation is not free, and the bill is usually logs rather than traces.

LeverEffect
Sample traces at 5–10%Keeps the shape, cuts cost sharply
Trace 100% of errors and slow requestsThe ones you actually need
Log at info in production, debug behind a flagDebug logs are the largest volume
Set log retention to 30 daysDefaults are often forever
Drop health check logsEvery 30 seconds, forever, per instance

That last one is worth checking today. A health check logged on every poll across four instances produces hundreds of thousands of useless lines a month, and it is a one-line filter to remove.

A dashboard worth having

One screen, in this order, because it maps to the order you diagnose:

Row 1 — Business. Orders per hour against last week. Checkout completion rate. Cart creation rate.

Row 2 — Symptoms. p95 latency by route. 5xx rate. Payment failure rate by provider.

Row 3 — Causes. Database connections in use. Slowest queries. Queue depth and oldest message age. Cache hit rate.

Row 4 — Deploys. A deploy marker overlaid on every graph.

The deploy markers are the single highest-value element. Most incidents correlate with a change, and a graph that shows when the change happened turns a twenty-minute investigation into a ten-second one.

Keep it to one screen. A dashboard nobody can take in during an incident is decoration.


Instrumentation is the part of self-hosting people skip and then regret. We set this up as part of every deployment.

Frequently asked questions

How should I log in a Medusa application?

Structured JSON with a correlation id generated per request and echoed in the response header, plus method, path, status, duration and customer id. Text logs become unusable as soon as you run more than one instance.

Does Medusa support OpenTelemetry?

Yes. Register OTel in an instrumentation file and enable HTTP, workflow and query instrumentation, so each workflow step appears as a span. That is what makes checkout latency attributable to a specific external call.

What should I alert on for an ecommerce backend?

Symptoms customers experience: orders dropping to zero, checkout error rate, payment failure rate by provider, and worker queue lag. Compare against the same window last week rather than absolute thresholds, and do not alert on CPU.

Why track business metrics alongside technical ones?

Because some failures produce no technical signal. A payment provider declining every card returns valid responses, so error rates stay flat while orders go to zero. Only the business metric catches it.

What is a synthetic checkout monitor?

A scheduled job that performs a real purchase against production every few minutes using a test card, then cancels the order. It exercises every dependency in the live configuration and is usually the fastest way to learn that checkout is broken.

Should the health check verify the database?

Have two. Keep the load balancer's check shallow so a slow dependency does not remove healthy instances from rotation, and expose a deeper check that verifies database and Redis for your monitoring system to poll.

[ Keep reading ]