Commerce code moves money and reserves stock, so the cost of a bug is denominated in currency rather than embarrassment. That argues for tests. It does not argue for testing everything — a 90% coverage number on a codebase whose checkout has never been exercised end to end is worse than useless, because it feels like safety.
Here is where tests actually earn their keep on a Medusa build.
Where the value is
| Code | Priority | Type |
|---|---|---|
| Workflows with compensation | Highest | Integration |
| Payment and tax calculation | Highest | Unit + integration |
| Custom module business rules | High | Unit |
| Custom API routes | High | Integration |
| Data importers | High | Integration |
| Subscribers | Medium | Integration |
| Admin widgets | Low | Manual |
| Medusa core behaviour | None | Do not |
Setup
const { loadEnv } = require("@medusajs/framework/utils")
loadEnv("test", process.cwd())
module.exports = {
transform: { "^.+\\.[jt]s$": ["@swc/jest", { jsc: { target: "esnext" } }] },
testEnvironment: "node",
moduleFileExtensions: ["js", "ts", "json"],
testTimeout: 60000,
}.env.test points at a separate database. Never run tests against development data — the harness truncates tables between runs, and doing that to the catalog you spent a morning building is a memorable way to learn this.
Unit tests for module services
Fast, and the right shape for business rules:
import { moduleIntegrationTestRunner } from "@medusajs/test-utils"
import { LOYALTY_MODULE } from ".."
import LoyaltyModuleService from "../service"
moduleIntegrationTestRunner<LoyaltyModuleService>({
moduleName: LOYALTY_MODULE,
testSuite: ({ service }) => {
describe("tier calculation", () => {
it("promotes to gold above the threshold", async () => {
const [account] = await service.createLoyaltyAccounts([
{ customer_id: "cus_1", points: 0 },
])
await service.addPoints(account.id, 5000)
const updated = await service.retrieveLoyaltyAccount(account.id)
expect(updated.tier).toBe("gold")
})
it("never demotes within a billing period", async () => {
const [account] = await service.createLoyaltyAccounts([
{ customer_id: "cus_2", points: 5000, tier: "gold" },
])
await service.redeemPoints(account.id, 4900)
const updated = await service.retrieveLoyaltyAccount(account.id)
// Business rule, not framework behaviour — exactly what belongs in a test.
expect(updated.tier).toBe("gold")
})
})
},
})Test the rules, not the generated CRUD. createLoyaltyAccounts works; your tier logic is the part that will break.
Integration tests for workflows
The highest-value tests you will write:
import { medusaIntegrationTestRunner } from "@medusajs/test-utils"
import { renewSubscriptionWorkflow } from "../../src/workflows/renew-subscription"
medusaIntegrationTestRunner({
testSuite: ({ getContainer }) => {
describe("renewSubscriptionWorkflow", () => {
it("charges, creates an order and advances the period", async () => {
const container = getContainer()
const subscription = await seedSubscription(container)
const { result } = await renewSubscriptionWorkflow(container).run({
input: { subscription_id: subscription.id },
})
expect(result.order).toBeDefined()
const updated = await container
.resolve("subscription")
.retrieveSubscription(subscription.id)
expect(new Date(updated.next_billing_at).getTime()).toBeGreaterThan(
new Date(subscription.next_billing_at).getTime(),
)
})
it("refunds the charge and holds the period when order creation fails", async () => {
const container = getContainer()
const subscription = await seedSubscription(container)
jest
.spyOn(container.resolve("order"), "createOrders")
.mockRejectedValueOnce(new Error("boom"))
await expect(
renewSubscriptionWorkflow(container).run({
input: { subscription_id: subscription.id },
}),
).rejects.toThrow()
const updated = await container
.resolve("subscription")
.retrieveSubscription(subscription.id)
// The period must NOT advance, or the customer loses a paid cycle.
expect(updated.next_billing_at).toEqual(subscription.next_billing_at)
expect(await getRefunds(container, subscription.id)).toHaveLength(1)
})
})
},
})That second test is the one that matters. Compensation is code that only runs when something has already gone wrong, which means it is never exercised in development and is exactly where bugs hide. Force a failure at each step and assert the world is unchanged. Workflows and compensation.
API route tests
medusaIntegrationTestRunner({
testSuite: ({ api }) => {
describe("GET /store/wishlists", () => {
it("requires authentication", async () => {
await expect(api.get("/store/wishlists")).rejects.toMatchObject({
response: { status: 401 },
})
})
it("returns only the caller's wishlists", async () => {
const { headers } = await authenticateCustomer(api, "a@example.com")
const response = await api.get("/store/wishlists", { headers })
expect(response.status).toBe(200)
expect(response.data.wishlists.every((w: any) => w.customer_id === "cus_a")).toBe(true)
})
})
},
})Always test the authorisation case. "Returns only the caller's data" is the assertion that stops a data leak, and it is the one most often left out.
External services
Mock them in tests, contract-test them separately.
In tests, mock the provider so you control failures — declines, timeouts, malformed responses — deterministically.
Contract tests run against the provider's sandbox on a schedule rather than in CI, and tell you when a payment provider changes a response shape. Running them in CI makes your pipeline depend on someone else's uptime, which is a bad trade.
CI
services:
postgres:
image: postgres:16
env:
POSTGRES_PASSWORD: postgres
options: >-
--health-cmd pg_isready --health-interval 10s --health-retries 5
redis:
image: redis:7
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with: { node-version: 22, cache: yarn }
- run: yarn install --immutable
- run: yarn test:unit
- run: yarn test:integration
- run: npx tsc --noEmitReal Postgres and Redis in CI. Mocked databases pass while the SQL is wrong, which is the opposite of what you want from a test suite.
What not to test
- Medusa's core. It has its own tests. Testing that
createProductscreates a product tells you nothing. - Generated CRUD. Same reason.
- Admin UI components. Manual verification is cheaper and catches more.
- Configuration. A boot-time validation is better than a test.
The one manual test
Before every release: complete a real purchase on production infrastructure with a real card. Automated tests do not exercise your live payment configuration, your live tax engine, or the carrier account you renewed last week.
It takes four minutes. Do it every time.
Test data
The unglamorous piece that determines whether the suite is pleasant or hated.
Factories over fixtures. A function that creates a subscription with sensible defaults and accepts overrides beats a static JSON file, because tests state only what they care about:
export async function createSubscription(
container: MedusaContainer,
overrides: Partial<SubscriptionInput> = {},
) {
const service = container.resolve("subscription")
const [subscription] = await service.createSubscriptions([
{
customer_id: "cus_test",
interval: "monthly",
interval_count: 1,
next_billing_at: new Date("2026-01-01"),
current_period_start: new Date("2025-12-01"),
current_period_end: new Date("2026-01-01"),
payment_provider_id: "pp_test_test",
...overrides,
},
])
return subscription
}Fixed dates, never new Date(). A test that passes in January and fails in March is worse than no test.
Each test creates what it needs. Shared state between tests produces failures that depend on execution order, which is the least debuggable kind.
Keeping the suite fast
An integration suite that takes fifteen minutes stops being run before it stops being useful.
- Run unit tests on every commit, integration tests on push. Fast feedback where it matters.
- Parallelise by file with an isolated database per worker. Jest handles this given separate connection strings.
- Migrate once, truncate between tests. Re-running migrations per test file is usually the single largest cost.
- Watch for a suite creeping past ten minutes and treat that as a bug to fix, not a fact to accept.
The goal is a suite people run locally before pushing. Anything slower than that becomes CI's problem, and CI-only tests get ignored when they fail.
We set up test harnesses as part of engagements, because the alternative shows up later. Ask.
Frequently asked questions
What should I test in a Medusa application?
Workflows first, especially their compensation paths, then payment and tax calculation, custom module business rules, custom API routes and data importers. Do not test Medusa's core or its generated CRUD methods.
How do I test Medusa workflows?
With `medusaIntegrationTestRunner` against a real test database. Run the workflow, assert the outcome, then force a failure in a later step and assert that compensation left the system unchanged — that second case is where the real bugs are.
Should Medusa tests use a real database?
Yes. Integration tests against a real Postgres catch SQL, migration and constraint problems that mocks hide entirely. Point them at a separate test database, since the harness truncates tables between runs.
How do I test payment integrations?
Mock the provider in your test suite so you can produce declines, timeouts and malformed responses deterministically. Run contract tests against the provider's sandbox on a schedule, outside CI, so your pipeline does not depend on their uptime.
What test coverage should I aim for?
Coverage is the wrong target. Aim for complete coverage of workflows, payment logic and custom business rules, and accept low coverage elsewhere. A high overall number on a codebase whose checkout has never been tested end to end is actively misleading.
Do I still need manual testing?
Yes — one manual purchase on production infrastructure with a real card before each release. Automated tests do not exercise your live payment configuration, tax engine or carrier accounts.
Observability for Medusa: Logs, Traces and the Alerts Worth Having
Self-hosting means owning the question 'is checkout working?'. Structured logging, tracing, the four metrics that matter and alerts that do not cry wolf.
Selling Digital Products with Medusa: Delivery, Licensing and VAT
No shipping, no inventory, and a set of problems physical goods never have: secure delivery, licence enforcement, and tax charged where the customer is.
Building a Marketplace on Medusa: Vendors, Splits and Payouts
Multi-vendor commerce means splitting one customer order into several vendor orders, and moving money to people who are not you. The model and the money mechanics.



