All posts
Forward Development

Medusa 2.21 cart locks: captured payment, no order

In Medusa 2.21, a slow cart completion can outlive its lock and lose it to a webhook. How the race works, how to find orphaned payments, and what to patch.

  • Medusa.js
  • Migrations
  • Operations

A customer emails support with a bank statement showing a charge. There is no order in the admin and no confirmation email, and their cart is still open with the same items in it. The payment provider's dashboard shows the payment as captured. Your Medusa payment record agrees. Nothing connects that payment to an order.

On Medusa 2.21.0 this can be a real race condition and not a one-off glitch. It is documented in medusajs/medusa#17200, and the reporter says it reproduces on every run. As of that report, develop has the same code, so there is no fixed release to upgrade to yet. Until there is, it is worth knowing whether your store is exposed and how to clean up after it.

Two lock defects combine into one bad outcome

Many of Medusa's core workflows guard shared resources with acquireLockStep and releaseLockStep. The issue counts 47 workflows in @medusajs/core-flows 2.21.0 that use them. Two of those workflows matter most here: completeCartWorkflow and processPaymentWorkflow. Both lock the cart with ttl: 120 and pass no ownerId.

That causes two separate problems:

  1. No owner. If you call the steps without an ownerId, every holder gets the provider's default owner, "*". The Redis provider's release() checks the owner before deleting, but that check can't tell two holders apart when both are "*".
  2. No renewal. acquireLockStep calls acquire() with a fixed TTL, and nothing renews it while the workflow runs. locking.execute() behaves differently: it renews the lease and gives each call its own owner (execute:<uuid>). The workflow steps do neither.

The issue points out that either defect alone breaks mutual exclusion. Without renewal, a slow holder and a new holder overlap. With a shared owner, the slow holder deletes the new holder's lock on its way out, and a third caller can get in.

How a slow authorization becomes an orphaned payment

The reproduction uses a separate server process and worker process, workflow-engine-redis, and locking-redis as the default locking provider, on PostgreSQL 16 and Redis 7.4. The sequence:

  1. completeCartWorkflow on the server takes the cart lock and creates the order. Then it blocks inside the payment provider's authorizePayment for more than 120 seconds.
  2. The lease expires. The provider sends a "captured" webhook, and the worker handles it with processPaymentWorkflow. That workflow takes the cart lock (also as "*"), finds no payment, authorizes the session, and starts capturing. It sees the order from step 1, so it won't complete the cart itself.
  3. The original authorizePayment returns. Authorization fails with Payment with payment_session_id: payses_…, already exists. (400), and the completion workflow compensates. The compensation of acquireLockStep releases the cart lock as "*", which deletes the webhook's lock while the webhook is still running. It then deletes the order and reopens the cart, setting completed_at back to null.
  4. The webhook finishes capturing.

The end state: the payment is captured at the provider and marked captured in Medusa, there is no live order, and the cart is open. The customer has paid and has no order.

You can see the root cause without any workflows at all. Acquire cart_1 with no owner and a 2-second expiry, wait 2.5 seconds, acquire it again, then call release("cart_1"). The release returns true and deletes the second holder's lock. A third acquire then succeeds. With distinct owners, the stale release returns false and the third acquire fails with Failed to acquire lock for key "cart_1".

Checking whether your store is exposed

The issue was reproduced with the Redis locking provider and separate server and worker processes, so that is the setup to check first. You need all three of these conditions for the failure in the report:

  • A payment provider whose authorizePayment can take longer than 120 seconds. That might be a 3-D Secure flow, a provider-side queue, a slow fraud check, or simply an HTTP client with no timeout. Look through your provider module's outbound calls for timeouts. If there aren't any, assume the call can hang indefinitely.
  • Webhooks that can arrive while completion is still running. Running webhooks on a worker while the server handles completion makes this easy to hit.
  • A failure path that compensates. In the report, it is the duplicate-payment 400 that triggers the stale release.

If your provider calls reliably finish within a few seconds and you enforce that with a timeout, your practical exposure is much lower. The ownership bug is still present, though, and it affects every workflow that uses these steps. The cart flow is just where the cost is most visible.

Finding orphaned payments

Start from the money. Find every payment that is captured or authorized in Medusa where the related cart has no order, or where the cart has completed_at set to null. In Medusa v2, carts, payment collections and orders are linked through module link tables, not foreign keys. Check the link table names in your own schema before writing the query. The logic is a payment collection with captured payments, linked to a cart, with no order linked to the same payment collection.

Then reconcile against the provider. Export captured charges for the same window from the provider's API and match on the provider's payment reference. Every captured charge without a live order goes into one of three buckets:

  • The customer completed the cart again later. Check for a duplicate charge, and refund the orphan.
  • The customer gave up. Refund, or create the order manually if the goods are still available and the customer wants them.
  • The cart was completed by a different path. Investigate before touching it.

Run this as a scheduled job, not a one-off. A daily reconciliation that flags orphans for a human is cheap, and it also catches the related failure in #17161. In that issue, completeCartWorkflow leaves a confirmed payment without an order after a checkpoint save error. The mechanism is different, but the symptom is the same.

Protections to add until upstream ships a fix

The fix the issue suggests is to give each workflow run its own default owner, derived from the step context's transactionId, and to renew the lease while the lock is held. Until that ships, here is what we would do:

  • Put a hard timeout on provider calls well under 120 seconds. This is the most effective change. If authorizePayment can't outlive the lease, there is no overlap in the first place. Treat a timeout as a failed authorization and let the provider's webhook or your reconciliation job settle the final state.
  • Patch the lock steps locally. Use patch-package or a similar tool to default ownerId to the step context's transactionId in both acquireLockStep and releaseLockStep, including the compensation. This closes the stale-release hole for all 47 workflows. Pin the exact Medusa version so the patch doesn't silently stop applying on upgrade.
  • Pass an explicit ownerId in your own workflows. For locks inside your own code, prefer locking.execute(), which already renews the lease and uses a unique owner.
  • Make webhook handling check order state before capturing. If the webhook finds an open cart with a captured payment, it should alert someone instead of finishing quietly.
  • Alert on the 400. Payment with payment_session_id: ..., already exists. during completion is the signature of this race. Log it as a priority event, not a routine client error.

Also watch #16285. There, releaseAll can delete another owner's lock, which is a separate path to the same broken exclusion.

What to do this week

Pull 90 days of captured charges from your provider and compare them against orders. If you find orphans, refund or fulfil them before the customers find them for you. Then set timeouts on your payment provider calls, patch the lock steps, and add the scheduled reconciliation. Keep the patch until a Medusa release explicitly fixes #17200, and confirm that against the release notes, not the version number.

If you are planning a move to Medusa, this kind of concurrency behavior belongs in the cutover test plan, next to data mapping. Our Magento 2 to Medusa or Vendure migration work treats payment reconciliation as a launch requirement. If you want a second pair of eyes on your checkout's failure modes, get in touch.

Need this done on a real stack?

Magento 2, Adobe Commerce, migrations to Medusa.js or Vendure, enterprise Next.js, WordPress, and AI automation.

Contact us