Medusa is a stateless Node application in front of Postgres, so it scales the way such applications always do. What is useful is knowing the order in which things break, because it is consistent and it is not what people expect.
Almost nobody's first bottleneck is the application tier.
The order things break
| Load | What gives | Fix |
|---|---|---|
| First | Postgres connections and slow queries | Pooling, indexes, read replicas |
| Second | Worker throughput on events | Partition by queue |
| Third | App instance CPU | Horizontal scaling |
| Fourth | Cache misses on hot reads | Redis with tighter keys |
| Rarely | Redis | Larger instance |
Postgres first
Every Medusa instance holds a connection pool. Ten instances at twenty connections each is two hundred connections, and a small managed Postgres tops out around a hundred.
projectConfig: {
databaseUrl: process.env.DATABASE_URL,
databaseDriverOptions: {
pool: { min: 2, max: 10 },
},
}Beyond a handful of instances, put PgBouncer in transaction mode between the app and the database. It multiplexes many client connections onto few server ones, and it is the difference between scaling to twenty instances and hitting a wall at six.
Then find the slow queries:
SELECT calls, mean_exec_time, total_exec_time, query
FROM pg_stat_statements
ORDER BY total_exec_time DESC
LIMIT 20;On a Medusa store the top of that list is nearly always product listing with price calculation. Which leads to the next point.
Read replicas
Commerce traffic is overwhelmingly reads — browsing, search, product pages — against a small number of writes. Route reads to replicas and keep writes on the primary.
Two caveats worth internalising:
- Replica lag is real. Never read a just-written record from a replica. Post-write reads — the cart you just updated, the order you just placed — go to the primary.
- Route deliberately, not globally. Catalog reads to replicas; anything in the checkout path to the primary.
Application tier
Once the database is comfortable, scale the server tier horizontally. Two things:
Scale on request count per instance, not CPU. Node applications spend their time waiting on I/O, so CPU stays deceptively low while latency climbs. Request count is the honest signal.
Keep instances stateless. Sessions in Redis, files in S3, cache in Redis. Any local state makes an instance non-interchangeable and breaks scaling in ways that only appear under load.
Workers do not scale by addition
Adding worker instances does not linearly increase throughput, because they contend for the same queues. If background processing is the bottleneck, partition instead: dedicate one worker to order events, one to search indexing, one to ERP sync, using separate queues.
That way a slow ERP sync cannot delay order confirmation emails, which is usually the actual problem being solved.
Caching
Two layers, and the outer one matters more.
Storefront. Statically rendered product and category pages served from a CDN. A traffic spike hits the CDN, not your backend, and this single decision does more for peak capacity than anything on the server side. Storefront performance.
Backend. The Redis cache module for repeated computation — price calculation, region resolution, shipping options. Caching strategy.
Cache what is expensive and stable. Do not cache carts.
Search
Product search through the database degrades badly with catalog size and query complexity. Once filtering gets interesting, move search to a purpose-built engine — Algolia or Meilisearch — and take that load off Postgres entirely.
This is often the single largest reduction in database load available on a catalog-heavy store.
Preparing for a peak
A checklist that has held up:
[ ] Load test at 3× expected peak, four weeks out
[ ] Read replicas provisioned and traffic routed
[ ] PgBouncer in front of Postgres
[ ] Storefront fully static with revalidation
[ ] Worker queues partitioned by concern
[ ] Autoscaling limits raised, and verified by actually hitting them
[ ] Rate limiting on auth and cart endpoints
[ ] Dashboards for DB connections, query time, queue depth, error rate
[ ] A rehearsed rollback
[ ] Someone on call who has seen the dashboards beforeThe one that saves the day is the load test in October. A load test in the week of the event only tells you what you cannot fix in time.
Measure before optimising
Every optimisation above costs something. Before applying any of them, look at:
- p95 latency by route
- Database connections in use versus maximum
- Slowest queries by total time
- Queue depth and processing lag
- Cache hit rate
Most "Medusa is slow" reports turn out to be one unindexed query or an over-fetching storefront. Observability covers instrumenting this properly.
A load test that tells you something
Most load tests measure the wrong thing — hammering the product endpoint tells you little, because that is the path your CDN already absorbs.
Model real behaviour:
import http from "k6/http"
import { check, group, sleep } from "k6"
export const options = {
stages: [
{ duration: "2m", target: 100 },
{ duration: "5m", target: 500 },
{ duration: "2m", target: 1500 }, // the spike
{ duration: "5m", target: 1500 },
{ duration: "2m", target: 0 },
],
thresholds: {
http_req_duration: ["p(95)<1000"],
http_req_failed: ["rate<0.01"],
},
}
export default function () {
group("browse", () => {
http.get(`${__ENV.BASE}/store/products?limit=24®ion_id=${__ENV.REGION}`)
sleep(2)
})
// Only a fraction of visitors reach checkout — model the ratio, not the peak.
if (Math.random() < 0.05) {
group("checkout", () => {
const cart = http.post(`${__ENV.BASE}/store/carts`, JSON.stringify({ region_id: __ENV.REGION }))
check(cart, { "cart created": (r) => r.status === 200 })
sleep(3)
})
}
}The write path is what breaks first under real load, and a browse-only test never touches it.
Reading the results
| Symptom under load | Cause | Fix |
|---|---|---|
| Latency climbs, CPU flat | Waiting on the database | Pooling, indexes, replicas |
| Errors at a fixed concurrency | Connection limit reached | PgBouncer |
| Latency fine, orders failing | Write contention or inventory locks | Review the completion workflow |
| Slow first request after idle | Cold instances | Minimum instance count |
| Queue lag growing | Worker throughput | Partition queues |
Run it at 3× expected peak, four weeks out. A test the week before only tells you what you cannot fix in time.
We do load testing and capacity planning ahead of peak seasons. Ask early.
Frequently asked questions
How much traffic can Medusa handle?
As much as the infrastructure behind it allows — it is a stateless Node application, so throughput scales with instances, and the practical ceiling is set by Postgres. Stores handling thousands of concurrent visitors run comfortably on modest hardware with read replicas and a static storefront.
What is the first bottleneck when scaling Medusa?
Postgres, essentially always — either connection exhaustion from too many app instances or slow catalog queries. Add pooling and read replicas before adding application instances.
Should I add read replicas for Medusa?
Once catalog reads dominate, yes. Route browsing and product queries to replicas and keep writes and post-write reads on the primary, since replica lag will otherwise show customers stale data immediately after they change something.
How do I scale Medusa workers?
Partition rather than replicate. Dedicate workers to specific queues — orders, search indexing, third-party sync — so a slow integration cannot delay customer-facing work. Adding identical workers mostly increases contention.
Does Medusa need a CDN?
The storefront does, and it is the highest-leverage capacity decision available. Statically rendered pages served from a CDN absorb traffic spikes without reaching the backend at all.
How do I prepare Medusa for Black Friday?
Load test at three times expected peak at least a month out, provision read replicas and connection pooling, make the storefront fully static, partition worker queues, raise and verify autoscaling limits, and have dashboards and a rehearsed rollback ready.
Caching Medusa: What to Cache, Where, and What Never To
Four cache layers, one rule about carts, and the invalidation strategy that stops a price change taking twelve hours to appear.
Medusa Storefront Performance: Where the Milliseconds Actually Go
Over-fetching, waterfalls and unoptimised images account for most of a slow headless storefront. How to find them and what to do about each.
Medusa Database Migrations: Generating, Reviewing and Deploying Safely
Generated migrations are only as safe as your review of them. How Medusa migrations work, and the expand-and-contract pattern that keeps rolling deploys from failing.



