3 MIN READ

Deploying Medusa on AWS: ECS, RDS and the Parts That Bite

A production AWS architecture for Medusa — ECS Fargate, RDS, ElastiCache, S3 — plus the networking and migration details that turn a two-day job into a two-week one.

BY RAHUL MEHTAUPDATED
Illustration for “Deploying Medusa on AWS: ECS, RDS and the Parts That Bite” — Deployment

AWS is the right home for Medusa when you have compliance requirements, an existing AWS footprint, or a platform team. It is the wrong home when you picked it because it felt more serious than the alternative — you will spend two weeks on networking that Railway gives you in an afternoon.

If you are going, this is the architecture and the parts that actually cost time.

The architecture

LayerServiceNotes
ComputeECS FargateTwo services: server, worker
Load balancingALBHTTPS via ACM, health check on `/health`
DatabaseRDS Postgres 16Multi-AZ in production
Cache and eventsElastiCache RedisSame VPC, private subnets
File storageS3Product images and uploads
CDNCloudFrontIn front of the storefront and S3
RegistryECRContainer images
SecretsSecrets ManagerInjected into task definitions
LogsCloudWatchStructured JSON

Public subnets for the ALB, private subnets for everything else. RDS and ElastiCache should have no route to the internet.

Networking, where the time goes

The single most common failure is a task that starts, cannot reach RDS, and dies with a timeout that names nothing useful.

  • Security groups reference each other, not CIDRs. The RDS security group allows 5432 from the ECS task security group. Same for Redis on 6379.
  • Private subnets need a NAT gateway or VPC endpoints to pull images from ECR and reach Stripe. This is also the line item that surprises people on the first bill.
  • The ALB lives in public subnets and forwards to tasks in private ones.

Get these three right and the rest is configuration.

Task definitions

Two services, one image, distinguished by environment:

server-task-definition.json (abridged)json
{
  "family": "medusa-server",
  "cpu": "1024",
  "memory": "2048",
  "networkMode": "awsvpc",
  "requiresCompatibilities": ["FARGATE"],
  "containerDefinitions": [
    {
      "name": "medusa",
      "image": "<account>.dkr.ecr.<region>.amazonaws.com/medusa:latest",
      "portMappings": [{ "containerPort": 9000 }],
      "environment": [
        { "name": "MEDUSA_WORKER_MODE", "value": "server" },
        { "name": "NODE_ENV", "value": "production" }
      ],
      "secrets": [
        { "name": "DATABASE_URL", "valueFrom": "arn:aws:secretsmanager:...:medusa/database-url" },
        { "name": "JWT_SECRET", "valueFrom": "arn:aws:secretsmanager:...:medusa/jwt-secret" },
        { "name": "COOKIE_SECRET", "valueFrom": "arn:aws:secretsmanager:...:medusa/cookie-secret" }
      ],
      "healthCheck": {
        "command": ["CMD-SHELL", "wget -qO- http://localhost:9000/health || exit 1"],
        "interval": 30,
        "timeout": 5,
        "retries": 3,
        "startPeriod": 60
      },
      "logConfiguration": {
        "logDriver": "awslogs",
        "options": {
          "awslogs-group": "/ecs/medusa",
          "awslogs-region": "<region>",
          "awslogs-stream-prefix": "server"
        }
      }
    }
  ]
}

The worker definition is identical except MEDUSA_WORKER_MODE=worker, no port mapping, and no ALB target group. It also needs no startPeriod grace as generous as the server's, but leave it — Medusa still loads every module on boot.

Secrets go in secrets, never environment. Values in environment are visible in the task definition to anyone with describe permissions.

Migrations as a task

bash
aws ecs run-task \
  --cluster medusa \
  --task-definition medusa-migrate \
  --launch-type FARGATE \
  --network-configuration "awsvpcConfiguration={subnets=[subnet-abc],securityGroups=[sg-tasks]}" \
  --overrides '{"containerOverrides":[{"name":"medusa","command":["npx","medusa","db:migrate"]}]}'

Wire it into the pipeline: build and push the image, run the migration task, wait for exit code zero, then update the services. If the migration fails, the deploy stops and the current revision keeps serving. Migration deployment patterns.

S3 for uploads

Local disk does not survive a Fargate task restart. Configure the S3 file provider:

medusa-config.tsts
modules: [
  {
    resolve: "@medusajs/medusa/file",
    options: {
      providers: [
        {
          resolve: "@medusajs/medusa/file-s3",
          id: "s3",
          options: {
            file_url: process.env.S3_FILE_URL,
            bucket: process.env.S3_BUCKET,
            region: process.env.S3_REGION,
            authentication_method: "iam-role",
          },
        },
      ],
    },
  },
]

Use an IAM task role rather than access keys, and put CloudFront in front of the bucket so file_url points at the distribution.

Autoscaling

Scale the server service on ALB request count per target — it responds faster than CPU for a Node application spending its time on I/O. Two tasks minimum for availability.

Scale the worker on queue depth or not at all. One worker handles a surprising amount, and scaling workers without partitioning simply means more contention.

What it costs

ItemMonthly
Fargate (2 server + 1 worker, small)$70–150
RDS Postgres (db.t4g.medium, Multi-AZ)$120–200
ElastiCache (cache.t4g.micro)$15–30
ALB$20–25
NAT gateway$35–70
S3 + CloudFront$10–50
**Total****~$270–575**

Two to four times a Railway equivalent. You are buying VPC isolation, compliance boundaries and integration with an existing AWS estate — real things, worth paying for when you need them and not before.

Do not use Lambda

Medusa is a long-lived Node process with module loading at boot and background work. On Lambda you get cold starts that make module loading painful, no worker mode, and no durable event processing. ECS Fargate is the right primitive.

Infrastructure as code

Do not click this together in the console. Whatever you build by hand you will rebuild by hand for staging, and then the two will differ in a way nobody discovers until an incident.

infra/ecs.tf (abridged)hcl
resource "aws_ecs_service" "server" {
  name            = "medusa-server"
  cluster         = aws_ecs_cluster.main.id
  task_definition = aws_ecs_task_definition.server.arn
  desired_count   = 2
  launch_type     = "FARGATE"

  network_configuration {
    subnets          = var.private_subnet_ids
    security_groups  = [aws_security_group.tasks.id]
    assign_public_ip = false
  }

  load_balancer {
    target_group_arn = aws_lb_target_group.server.arn
    container_name   = "medusa"
    container_port   = 9000
  }

  deployment_circuit_breaker {
    enable   = true
    rollback = true
  }
}

The circuit breaker is worth calling out: it rolls back automatically when a deployment's tasks fail to stabilise, which turns a bad deploy into a non-event.

Cost control

The AWS bill has three lines that surprise people, all avoidable:

NAT gateway data processing. Every byte from a private subnet to the internet is charged. Add VPC endpoints for ECR, S3, Secrets Manager and CloudWatch Logs, and the NAT traffic drops sharply.

CloudWatch Logs retention. Defaults to never expire. Set retention to 30 days on the log group unless you have a reason not to.

Over-provisioned Fargate. Start at 1 vCPU and 2GB for the server, less for the worker, and size from real metrics rather than a guess. Fargate bills per second, so scaling down when traffic drops is a real saving rather than a rounding error.

Set a billing alarm before the first deploy, not after the first invoice.


We have built this stack more than once and would rather you did not rediscover the NAT gateway bill. Ask us.

Frequently asked questions

What is the best way to deploy Medusa on AWS?

ECS Fargate with two services from a single image — one in server mode behind an ALB, one in worker mode — plus RDS Postgres, ElastiCache Redis, S3 for uploads and CloudFront in front. Migrations run as a standalone ECS task in the deployment pipeline.

Can Medusa run on AWS Lambda?

Not sensibly. Medusa expects a long-lived process with module loading at startup and background job execution, none of which fit Lambda's execution model. Use ECS Fargate.

Why can my ECS task not connect to RDS?

Almost always security groups. The RDS security group must allow port 5432 from the ECS task security group specifically, not a CIDR range, and the task must be in a subnet with a route to the database. The same applies to ElastiCache on 6379.

How should secrets be handled on ECS?

Store them in Secrets Manager or SSM Parameter Store and reference them in the task definition's `secrets` block, which injects them at runtime. Anything placed in `environment` is readable by anyone who can describe the task definition.

How much does running Medusa on AWS cost?

Roughly $270–575 a month for a production setup with Multi-AZ RDS. The NAT gateway and Multi-AZ database are the two largest line items and the ones most often forgotten in early estimates.

Do I need a NAT gateway?

If your tasks run in private subnets and need to reach ECR, Stripe or any external API, yes — or VPC endpoints for the AWS services plus a NAT for the rest. Endpoints reduce the bill but add configuration; most teams start with a single NAT gateway.

[ Keep reading ]