3 MIN READ

Search in Medusa: Algolia, Meilisearch and Keeping the Index Fresh

Database search stops working sooner than you expect. Choosing an engine, building the index, and the sync strategy that survives a bulk import.

BY RAHUL MEHTAUPDATED
Illustration for “Search in Medusa: Algolia, Meilisearch and Keeping the Index Fresh” — Search

Postgres ILIKE is fine for a hundred products. Somewhere between there and a few thousand — sooner if customers expect typo tolerance, facets and instant results — it stops being fine, and it is putting load on the database that scaling work then has to compensate for.

Moving search to a dedicated engine solves both problems at once.

Choosing

AlgoliaMeilisearchPostgres
HostingManagedSelf-host or cloudAlready there
CostPer search and recordServer cost, or cloud tierFree
Typo toleranceExcellentExcellentNone
FacetingExcellentGoodManual
Relevance tuningExtensiveGoodManual
Setup effortLowLow–moderateNone

Algolia when search is a primary discovery path and you would rather buy the problem. Cost scales with usage and becomes noticeable at high volume.

Meilisearch when you want to self-host, keep cost fixed, or keep data in your own infrastructure. Genuinely good, and a small server handles a large catalog.

Postgres for small catalogs where search is a convenience rather than the main way people find things. There is no shame in this; premature search infrastructure is a real cost.

The index document

Flatten and denormalise. Your index is a read model shaped for querying, not a copy of your schema:

src/modules/search/transform.tsts
export function toSearchDocument(product: any) {
  const variants = product.variants ?? []

  return {
    objectID: product.id,
    title: product.title,
    handle: product.handle,
    description: (product.description ?? "").replace(/<[^>]+>/g, "").slice(0, 2000),
    thumbnail: product.thumbnail,

    // Facets — flat arrays index and filter far better than nested objects.
    categories: (product.categories ?? []).map((c: any) => c.name),
    collection: product.collection?.title ?? null,
    tags: (product.tags ?? []).map((t: any) => t.value),
    options: variants.flatMap((v: any) =>
      (v.options ?? []).map((o: any) => `${o.option?.title}:${o.value}`),
    ),

    // Denormalised for sorting and range filters only.
    min_price: Math.min(
      ...variants.map((v: any) => v.calculated_price?.calculated_amount ?? Infinity),
    ),
    in_stock: variants.some((v: any) => (v.inventory_quantity ?? 0) > 0 || !v.manage_inventory),
    created_at_ts: new Date(product.created_at).getTime(),
  }
}

Two habits. Strip HTML from descriptions, or customers search for div and find everything. And treat indexed price and stock as approximations — good enough to filter and sort by, never good enough to display.

Keeping it fresh

A subscriber per relevant event:

src/subscribers/sync-search-index.tsts
import type { SubscriberArgs, SubscriberConfig } from "@medusajs/framework"

export default async function syncSearchIndex({
  event,
  container,
}: SubscriberArgs<{ id: string }>) {
  const search = container.resolve("search")
  const query = container.resolve("query")

  if (event.name === "product.deleted") {
    await search.delete(event.data.id)
    return
  }

  const { data: [product] } = await query.graph({
    entity: "product",
    fields: ["*", "variants.*", "variants.calculated_price", "categories.*", "tags.*"],
    filters: { id: event.data.id },
  })

  // Unpublished products must leave the index, not merely stop being shown.
  if (!product || product.status !== "published") {
    await search.delete(event.data.id)
    return
  }

  await search.upsert(toSearchDocument(product))
}

export const config: SubscriberConfig = {
  event: ["product.created", "product.updated", "product.deleted"],
}

Three things worth building alongside it:

A full reindex script. npx medusa exec ./src/scripts/reindex.ts. You will need it after a schema change, a bad deploy, or a bulk import.

Batching for bulk operations. A 5,000-product import firing 5,000 individual index calls will rate-limit you. Detect bulk context and batch.

A dedicated worker queue. Indexing should not compete with order processing. Worker partitioning.

Searching, then re-fetching

The pattern that keeps stale prices off the page:

lib/data/search.tsts
export async function searchProducts(query: string, filters: SearchFilters) {
  const results = await searchClient.search(query, {
    filters: buildFilterString(filters),
    hitsPerPage: 24,
  })

  if (!results.hits.length) return { products: [], count: 0 }

  // Index gives the ordering; Medusa gives the truth about price and stock.
  const { products } = await sdk.store.product.list({
    id: results.hits.map((h: any) => h.objectID),
    region_id: filters.regionId,
    fields: "id,title,handle,thumbnail,*variants.calculated_price",
  })

  const byId = new Map(products.map((p) => [p.id, p]))
  return {
    products: results.hits.map((h: any) => byId.get(h.objectID)).filter(Boolean),
    count: results.nbHits,
  }
}

The index supplies relevance and ordering; Medusa supplies current price and availability. Rendering prices straight from the index is how a customer sees last week's sale price on a search page.

Facets

Search engines do faceted filtering natively and cheaply — counts per facet value, multi-select, ranges. This is where they earn their place over ILIKE, and it is a materially better browsing experience than category pages with hardcoded filters.

For SEO, be careful which facet combinations become crawlable URLs. Storefront SEO covers where to draw that line.

Practical guidance

Index only published, purchasable products. Draft products in search results are a support ticket.

Log queries with zero results. It is the cheapest merchandising signal available — it tells you what customers want that you do not stock or do not name the way they do.

Set a search timeout with a fallback. A search outage should degrade to database search, not to a broken page.

Reindex after every catalog import. Make it the last step of the import script so nobody has to remember.

Relevance tuning

Default relevance is a starting point. Three adjustments that make the largest difference on a commerce catalog:

Rank searchable attributes deliberately. Title matches should beat description matches. Both engines let you order the attribute list, and the order is the ranking.

Add a business signal as a tiebreaker. Between two equally relevant products, show the one that sells. A popularity field on the index — orders in the last 30 days, refreshed nightly — used as a custom ranking criterion converts noticeably better than pure text relevance.

Add synonyms for how customers actually speak. "Jumper" and "sweater", "trainers" and "sneakers", your SKU prefixes, common misspellings of your brand. Pull the list from your zero-result query log rather than inventing it.

ts
await index.setSettings({
  searchableAttributes: ["title", "categories", "tags", "description"],
  customRanking: ["desc(popularity)", "desc(in_stock)"],
  attributesForFaceting: ["categories", "collection", "tags", "options", "in_stock"],
})

Note in_stock in the custom ranking: out-of-stock products should rank below available ones rather than disappearing, so the page still shows what you sell.

Search analytics

The zero-result log is the highest-value merchandising data a store produces, and almost nobody reads it.

SignalWhat it tells you
Zero-result queriesProducts you lack, or names you do not use
High-volume, low-click queriesRelevance failing on a term that matters
Queries followed by a filter changeYour facets are doing the search's job
Search-then-exitThe single clearest failure to act on

Review it monthly. A recurring zero-result query is either a product opportunity or a synonym you can add in thirty seconds — and either way it is a customer who wanted to buy something and could not find it.


Search is usually the largest single reduction in database load on a catalog-heavy store. Ask us.

Frequently asked questions

Should I use Algolia or Meilisearch with Medusa?

Algolia if you want a managed service with excellent relevance tooling and can accept usage-based pricing. Meilisearch if you want to self-host, keep costs fixed or keep data in your own infrastructure. Both integrate the same way.

How do I keep the search index in sync with Medusa?

Subscribe to product created, updated and deleted events and upsert or delete the corresponding document. Add a full reindex script for recovery, batch during bulk imports, and run indexing on its own worker queue.

Should prices be rendered from the search index?

No. Use the index for relevance, filtering and ordering, then fetch the matching products from Medusa for current prices and stock. Index data is a lagging copy and will occasionally be wrong at exactly the wrong moment.

Can I just use Postgres for product search?

For small catalogs where search is a convenience, yes. Once customers expect typo tolerance, faceted filtering and instant results — or the catalog passes a few thousand products — a dedicated engine is both better and lighter on your database.

How should the index document be structured?

Flat and denormalised: title, handle, cleaned description, arrays of category, tag and option values for faceting, plus minimum price and stock flags for sorting and filtering. Strip HTML from descriptions before indexing.

What should happen if search is unavailable?

Degrade to a database query rather than showing an error. Set a short timeout on the search call and fall back, so an outage at the search provider does not take down product discovery entirely.

[ Keep reading ]