3 MIN READ

Testing Medusa: What to Test, What to Skip, and How

Commerce code moves money, so untested checkout logic is a liability. A pragmatic test strategy — what earns its keep, and what is theatre.

BY RAHUL MEHTAUPDATED
Illustration for “Testing Medusa: What to Test, What to Skip, and How” — Testing

Commerce code moves money and reserves stock, so the cost of a bug is denominated in currency rather than embarrassment. That argues for tests. It does not argue for testing everything — a 90% coverage number on a codebase whose checkout has never been exercised end to end is worse than useless, because it feels like safety.

Here is where tests actually earn their keep on a Medusa build.

Where the value is

CodePriorityType
Workflows with compensationHighestIntegration
Payment and tax calculationHighestUnit + integration
Custom module business rulesHighUnit
Custom API routesHighIntegration
Data importersHighIntegration
SubscribersMediumIntegration
Admin widgetsLowManual
Medusa core behaviourNoneDo not

Setup

jest.config.jsts
const { loadEnv } = require("@medusajs/framework/utils")
loadEnv("test", process.cwd())

module.exports = {
  transform: { "^.+\\.[jt]s$": ["@swc/jest", { jsc: { target: "esnext" } }] },
  testEnvironment: "node",
  moduleFileExtensions: ["js", "ts", "json"],
  testTimeout: 60000,
}

.env.test points at a separate database. Never run tests against development data — the harness truncates tables between runs, and doing that to the catalog you spent a morning building is a memorable way to learn this.

Unit tests for module services

Fast, and the right shape for business rules:

src/modules/loyalty/__tests__/service.spec.tsts
import { moduleIntegrationTestRunner } from "@medusajs/test-utils"
import { LOYALTY_MODULE } from ".."
import LoyaltyModuleService from "../service"

moduleIntegrationTestRunner<LoyaltyModuleService>({
  moduleName: LOYALTY_MODULE,
  testSuite: ({ service }) => {
    describe("tier calculation", () => {
      it("promotes to gold above the threshold", async () => {
        const [account] = await service.createLoyaltyAccounts([
          { customer_id: "cus_1", points: 0 },
        ])

        await service.addPoints(account.id, 5000)
        const updated = await service.retrieveLoyaltyAccount(account.id)

        expect(updated.tier).toBe("gold")
      })

      it("never demotes within a billing period", async () => {
        const [account] = await service.createLoyaltyAccounts([
          { customer_id: "cus_2", points: 5000, tier: "gold" },
        ])

        await service.redeemPoints(account.id, 4900)
        const updated = await service.retrieveLoyaltyAccount(account.id)

        // Business rule, not framework behaviour — exactly what belongs in a test.
        expect(updated.tier).toBe("gold")
      })
    })
  },
})

Test the rules, not the generated CRUD. createLoyaltyAccounts works; your tier logic is the part that will break.

Integration tests for workflows

The highest-value tests you will write:

integration-tests/workflows/renew-subscription.spec.tsts
import { medusaIntegrationTestRunner } from "@medusajs/test-utils"
import { renewSubscriptionWorkflow } from "../../src/workflows/renew-subscription"

medusaIntegrationTestRunner({
  testSuite: ({ getContainer }) => {
    describe("renewSubscriptionWorkflow", () => {
      it("charges, creates an order and advances the period", async () => {
        const container = getContainer()
        const subscription = await seedSubscription(container)

        const { result } = await renewSubscriptionWorkflow(container).run({
          input: { subscription_id: subscription.id },
        })

        expect(result.order).toBeDefined()

        const updated = await container
          .resolve("subscription")
          .retrieveSubscription(subscription.id)

        expect(new Date(updated.next_billing_at).getTime()).toBeGreaterThan(
          new Date(subscription.next_billing_at).getTime(),
        )
      })

      it("refunds the charge and holds the period when order creation fails", async () => {
        const container = getContainer()
        const subscription = await seedSubscription(container)

        jest
          .spyOn(container.resolve("order"), "createOrders")
          .mockRejectedValueOnce(new Error("boom"))

        await expect(
          renewSubscriptionWorkflow(container).run({
            input: { subscription_id: subscription.id },
          }),
        ).rejects.toThrow()

        const updated = await container
          .resolve("subscription")
          .retrieveSubscription(subscription.id)

        // The period must NOT advance, or the customer loses a paid cycle.
        expect(updated.next_billing_at).toEqual(subscription.next_billing_at)
        expect(await getRefunds(container, subscription.id)).toHaveLength(1)
      })
    })
  },
})

That second test is the one that matters. Compensation is code that only runs when something has already gone wrong, which means it is never exercised in development and is exactly where bugs hide. Force a failure at each step and assert the world is unchanged. Workflows and compensation.

API route tests

ts
medusaIntegrationTestRunner({
  testSuite: ({ api }) => {
    describe("GET /store/wishlists", () => {
      it("requires authentication", async () => {
        await expect(api.get("/store/wishlists")).rejects.toMatchObject({
          response: { status: 401 },
        })
      })

      it("returns only the caller's wishlists", async () => {
        const { headers } = await authenticateCustomer(api, "a@example.com")
        const response = await api.get("/store/wishlists", { headers })

        expect(response.status).toBe(200)
        expect(response.data.wishlists.every((w: any) => w.customer_id === "cus_a")).toBe(true)
      })
    })
  },
})

Always test the authorisation case. "Returns only the caller's data" is the assertion that stops a data leak, and it is the one most often left out.

External services

Mock them in tests, contract-test them separately.

In tests, mock the provider so you control failures — declines, timeouts, malformed responses — deterministically.

Contract tests run against the provider's sandbox on a schedule rather than in CI, and tell you when a payment provider changes a response shape. Running them in CI makes your pipeline depend on someone else's uptime, which is a bad trade.

CI

.github/workflows/test.ymlyaml
services:
  postgres:
    image: postgres:16
    env:
      POSTGRES_PASSWORD: postgres
    options: >-
      --health-cmd pg_isready --health-interval 10s --health-retries 5
  redis:
    image: redis:7

steps:
  - uses: actions/checkout@v4
  - uses: actions/setup-node@v4
    with: { node-version: 22, cache: yarn }
  - run: yarn install --immutable
  - run: yarn test:unit
  - run: yarn test:integration
  - run: npx tsc --noEmit

Real Postgres and Redis in CI. Mocked databases pass while the SQL is wrong, which is the opposite of what you want from a test suite.

What not to test

  • Medusa's core. It has its own tests. Testing that createProducts creates a product tells you nothing.
  • Generated CRUD. Same reason.
  • Admin UI components. Manual verification is cheaper and catches more.
  • Configuration. A boot-time validation is better than a test.

The one manual test

Before every release: complete a real purchase on production infrastructure with a real card. Automated tests do not exercise your live payment configuration, your live tax engine, or the carrier account you renewed last week.

It takes four minutes. Do it every time.

Test data

The unglamorous piece that determines whether the suite is pleasant or hated.

Factories over fixtures. A function that creates a subscription with sensible defaults and accepts overrides beats a static JSON file, because tests state only what they care about:

integration-tests/factories/subscription.tsts
export async function createSubscription(
  container: MedusaContainer,
  overrides: Partial<SubscriptionInput> = {},
) {
  const service = container.resolve("subscription")

  const [subscription] = await service.createSubscriptions([
    {
      customer_id: "cus_test",
      interval: "monthly",
      interval_count: 1,
      next_billing_at: new Date("2026-01-01"),
      current_period_start: new Date("2025-12-01"),
      current_period_end: new Date("2026-01-01"),
      payment_provider_id: "pp_test_test",
      ...overrides,
    },
  ])

  return subscription
}

Fixed dates, never new Date(). A test that passes in January and fails in March is worse than no test.

Each test creates what it needs. Shared state between tests produces failures that depend on execution order, which is the least debuggable kind.

Keeping the suite fast

An integration suite that takes fifteen minutes stops being run before it stops being useful.

  • Run unit tests on every commit, integration tests on push. Fast feedback where it matters.
  • Parallelise by file with an isolated database per worker. Jest handles this given separate connection strings.
  • Migrate once, truncate between tests. Re-running migrations per test file is usually the single largest cost.
  • Watch for a suite creeping past ten minutes and treat that as a bug to fix, not a fact to accept.

The goal is a suite people run locally before pushing. Anything slower than that becomes CI's problem, and CI-only tests get ignored when they fail.


We set up test harnesses as part of engagements, because the alternative shows up later. Ask.

Frequently asked questions

What should I test in a Medusa application?

Workflows first, especially their compensation paths, then payment and tax calculation, custom module business rules, custom API routes and data importers. Do not test Medusa's core or its generated CRUD methods.

How do I test Medusa workflows?

With `medusaIntegrationTestRunner` against a real test database. Run the workflow, assert the outcome, then force a failure in a later step and assert that compensation left the system unchanged — that second case is where the real bugs are.

Should Medusa tests use a real database?

Yes. Integration tests against a real Postgres catch SQL, migration and constraint problems that mocks hide entirely. Point them at a separate test database, since the harness truncates tables between runs.

How do I test payment integrations?

Mock the provider in your test suite so you can produce declines, timeouts and malformed responses deterministically. Run contract tests against the provider's sandbox on a schedule, outside CI, so your pipeline does not depend on their uptime.

What test coverage should I aim for?

Coverage is the wrong target. Aim for complete coverage of workflows, payment logic and custom business rules, and accept low coverage elsewhere. A high overall number on a codebase whose checkout has never been tested end to end is actively misleading.

Do I still need manual testing?

Yes — one manual purchase on production infrastructure with a real card before each release. Automated tests do not exercise your live payment configuration, tax engine or carrier accounts.

[ Keep reading ]