On a hosted platform, "is checkout working?" is somebody else's dashboard. Self-hosted, it is yours — and the honest answer is that most teams find out from a customer email.
The instrumentation that prevents that is not elaborate. It is structured logs with a correlation id, traces through the checkout path, four metrics, and alerts on symptoms rather than causes.
Structured logging
import { defineMiddlewares } from "@medusajs/framework/http"
import { randomUUID } from "node:crypto"
export default defineMiddlewares({
routes: [
{
matcher: "*",
middlewares: [
(req: any, res: any, next: any) => {
// One id per request, echoed to the client so support can quote it.
const correlationId = req.headers["x-correlation-id"] ?? randomUUID()
req.correlationId = correlationId
res.setHeader("x-correlation-id", correlationId)
const start = Date.now()
res.on("finish", () => {
req.scope.resolve("logger").info(
JSON.stringify({
correlation_id: correlationId,
method: req.method,
path: req.path,
status: res.statusCode,
duration_ms: Date.now() - start,
customer_id: req.auth_context?.actor_id ?? null,
}),
)
})
next()
},
],
},
],
})The correlation id is the whole point. Returning it in the response header means a customer support ticket can carry the exact identifier that ties together every log line for that request, across server and worker.
Log JSON. Human-readable logs are pleasant on one instance and useless across six.
Tracing the checkout
Distributed tracing pays for itself on exactly one path: checkout, where a single customer action fans out to a payment provider, a tax engine, a carrier and your database.
import { registerOtel } from "@medusajs/medusa"
import { OTLPTraceExporter } from "@opentelemetry/exporter-trace-otlp-http"
export function register() {
registerOtel({
serviceName: "medusa-server",
exporter: new OTLPTraceExporter({ url: process.env.OTEL_EXPORTER_URL }),
instrument: {
http: true,
workflows: true,
query: true,
},
})
}With workflow instrumentation on, each workflow step becomes a span. When checkout is slow, the trace tells you whether it was the tax engine or your own query, which is the difference between a fix and a guess.
The metrics worth having
Four technical, four business.
Technical
| Metric | Why | Warning |
|---|---|---|
| p95 latency by route | Users feel the tail | > 1s on store routes |
| 5xx rate | Something is broken | > 0.5% |
| DB connections in use | The first bottleneck | > 80% of max |
| Queue lag | Background work falling behind | > 5 min |
Business
| Metric | Catches |
|---|---|
| Orders per hour | Everything, eventually |
| Checkout completion rate | Payment or shipping breakage |
| Payment failure rate by provider | A processor degrading |
| Cart creation rate | Storefront or CDN failure |
Business metrics catch failures that technical ones miss entirely. A payment provider silently declining every card returns 200s with a valid error body — your error rate is flat and your orders are zero. Only the orders-per-hour graph tells you.
Alerts that do not cry wolf
Alert on symptoms customers experience:
- alert: OrdersDroppedToZero
expr: increase(orders_created_total[30m]) == 0
for: 30m
severity: critical
# Compare against the same window last week, not a constant, so a
# quiet Tuesday at 4am does not page anyone.
- alert: CheckoutErrorRate
expr: rate(checkout_errors_total[5m]) / rate(checkout_attempts_total[5m]) > 0.05
for: 10m
severity: critical
- alert: PaymentProviderFailing
expr: rate(payment_failures_total[10m]) / rate(payment_attempts_total[10m]) > 0.2
for: 10m
severity: critical
- alert: WorkerQueueLag
expr: queue_oldest_message_age_seconds > 600
for: 15m
severity: warningTwo disciplines that keep alerting useful. Every page must be actionable — if the response is "acknowledge and go back to sleep", delete the alert. And compare against the same window last week, because commerce traffic is deeply seasonal and absolute thresholds page you at 4am on a Tuesday.
Do not alert on CPU. It is a cause, not a symptom, and a Node application under I/O load has low CPU while serving nothing.
Error tracking
Sentry or equivalent, with the correlation id attached:
Sentry.withScope((scope) => {
scope.setTag("correlation_id", req.correlationId)
scope.setTag("route", req.path)
scope.setContext("cart", { id: cartId, item_count: itemCount })
Sentry.captureException(error)
})Filter expected errors — declined cards, out-of-stock at completion — into a separate stream. They are customer events, not defects, and leaving them in the exception feed trains everyone to ignore it.
Health checks
/health proves the process is up, which is a low bar. Add a deeper check for your own monitoring:
export async function GET(req: MedusaRequest, res: MedusaResponse) {
const checks = {
database: await ping(() => req.scope.resolve("query").graph({
entity: "region", fields: ["id"], pagination: { take: 1 },
})),
redis: await ping(() => req.scope.resolve("cache").set("health", "1", 10)),
}
const healthy = Object.values(checks).every(Boolean)
res.status(healthy ? 200 : 503).json({ healthy, checks })
}Keep it off the load balancer's check — a slow dependency should not take healthy instances out of rotation. Poll it from monitoring instead.
A synthetic checkout
The single highest-value monitor: a scheduled job that runs a real purchase against production every fifteen minutes — create cart, add item, set address, select shipping, authorise with a test card, complete, then cancel the order.
It exercises every external dependency in the real configuration and tells you checkout is broken before a customer does. In our experience it catches more real incidents than every other alert combined.
Cost of observability
Instrumentation is not free, and the bill is usually logs rather than traces.
| Lever | Effect |
|---|---|
| Sample traces at 5–10% | Keeps the shape, cuts cost sharply |
| Trace 100% of errors and slow requests | The ones you actually need |
| Log at info in production, debug behind a flag | Debug logs are the largest volume |
| Set log retention to 30 days | Defaults are often forever |
| Drop health check logs | Every 30 seconds, forever, per instance |
That last one is worth checking today. A health check logged on every poll across four instances produces hundreds of thousands of useless lines a month, and it is a one-line filter to remove.
A dashboard worth having
One screen, in this order, because it maps to the order you diagnose:
Row 1 — Business. Orders per hour against last week. Checkout completion rate. Cart creation rate.
Row 2 — Symptoms. p95 latency by route. 5xx rate. Payment failure rate by provider.
Row 3 — Causes. Database connections in use. Slowest queries. Queue depth and oldest message age. Cache hit rate.
Row 4 — Deploys. A deploy marker overlaid on every graph.
The deploy markers are the single highest-value element. Most incidents correlate with a change, and a graph that shows when the change happened turns a twenty-minute investigation into a ten-second one.
Keep it to one screen. A dashboard nobody can take in during an incident is decoration.
Instrumentation is the part of self-hosting people skip and then regret. We set this up as part of every deployment.
Frequently asked questions
How should I log in a Medusa application?
Structured JSON with a correlation id generated per request and echoed in the response header, plus method, path, status, duration and customer id. Text logs become unusable as soon as you run more than one instance.
Does Medusa support OpenTelemetry?
Yes. Register OTel in an instrumentation file and enable HTTP, workflow and query instrumentation, so each workflow step appears as a span. That is what makes checkout latency attributable to a specific external call.
What should I alert on for an ecommerce backend?
Symptoms customers experience: orders dropping to zero, checkout error rate, payment failure rate by provider, and worker queue lag. Compare against the same window last week rather than absolute thresholds, and do not alert on CPU.
Why track business metrics alongside technical ones?
Because some failures produce no technical signal. A payment provider declining every card returns valid responses, so error rates stay flat while orders go to zero. Only the business metric catches it.
What is a synthetic checkout monitor?
A scheduled job that performs a real purchase against production every few minutes using a test card, then cancels the order. It exercises every dependency in the live configuration and is usually the fastest way to learn that checkout is broken.
Should the health check verify the database?
Have two. Keep the load balancer's check shallow so a slow dependency does not remove healthy instances from rotation, and expose a deeper check that verifies database and Redis for your monitoring system to poll.
Testing Medusa: What to Test, What to Skip, and How
Commerce code moves money, so untested checkout logic is a liability. A pragmatic test strategy — what earns its keep, and what is theatre.
Deploying Medusa on AWS: ECS, RDS and the Parts That Bite
A production AWS architecture for Medusa — ECS Fargate, RDS, ElastiCache, S3 — plus the networking and migration details that turn a two-day job into a two-week one.
Deploying Medusa on Railway: The Fastest Production Setup
Railway is the shortest path from a Medusa repository to a production deployment that is actually correct. The full setup, including the worker service everyone forgets.



