Skip to content

fix(cluster): handle MOVED and ASK redirects for atomic MULTI - #3395

Merged
nkaradzhov merged 2 commits into
redis:masterfrom
konovalovsergey:fix-cluster-multi-redirects
Aug 11, 2026
Merged

nkaradzhov merged 2 commits into
redis:masterfrom
konovalovsergey:fix-cluster-multi-redirects

Conversation

@konovalovsergey

@konovalovsergey konovalovsergey commented Aug 10, 2026 •

Copy link
Copy Markdown
Contributor

Description

Fixes #3394.

After a coordinated Redis Cluster failover, the previous primary can remain reachable while returning MOVED for slots owned by the promoted replica. Regular cluster commands recover through the redirect loop, but atomic MULTI/EXEC called _executeMulti() directly and could not refresh a stale slot map.

This change routes atomic cluster transactions through the existing redirect loop so queue-time MOVED and ASK replies can refresh routing and retry.

execAsPipeline() remains on its direct non-retrying path because a pipeline can be partially applied. The redirect chainId is passed to _executeMulti() so ASKING and MULTI ... EXEC stay in the same queue chain.

Only explicit Redis routing errors are retried. Socket errors, timeouts, and other ambiguous failures remain non-retried.

No public API or configuration is changed.

Testing

  • Added a unit test for retrying MULTI after MOVED.
  • Added a unit test for sharing the queue chain between ASKING and redirected MULTI.
  • Added a unit test confirming that non-atomic pipelines are not retried.
  • Added a coordinated failover integration test with the previous primary still reachable.
  • Build, lint, type checks, and 25 related unit tests pass locally.

Checklist

  • Does npm test pass with this change (including linting)?
  • Is the new or changed code fully tested?
  • Is a documentation update included (if this change modifies existing APIs, or introduces new ones)? Not applicable: no public API or configuration changes.

Note

Medium Risk
Changes cluster transaction routing and retry behavior on redirect errors; scope is limited to atomic MULTI and internal APIs, with explicit tests for MOVED/ASK and failover.

Overview
Fixes stale-slot failures after failover when the old primary still answers with MOVED by routing cluster multi().exec() through _execute instead of calling _executeMulti on a one-shot client pick.

_executeMulti accepts an optional chainId from the redirect path so ASKING and the retried MULTI … EXEC share one queue chain. execAsPipeline() is unchanged and still does not retry on MOVED.

Docs in extractCommandsForSlots note that atomic MULTI can retry via the redirect loop while non-atomic pipelines do not. Tests cover MOVED retry, ASK chain sharing, pipeline no-retry, and a live failover with a reachable stale master.

Reviewed by Cursor Bugbot for commit 8b06219. Bugbot is set up for automated code reviews on this repo. Configure here.

Route atomic cluster transactions through the existing redirect loop so
queue-time MOVED and ASK replies can refresh routing and retry.

Keep execAsPipeline on the direct path because it can be partially applied.
Pass the redirect chain ID through _executeMulti so ASKING and MULTI/EXEC
stay in the same queue chain.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 12feb64d8c

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread packages/client/lib/cluster/index.ts
Comment thread packages/client/lib/cluster/multi-redirect.spec.ts

@nkaradzhov nkaradzhov left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you @konovalovsergey, this looks good to go! Appreciate the careful and clean report and fix, its rare these days :D

@nkaradzhov
nkaradzhov merged commit ba1b90b into redis:master Aug 11, 2026
15 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Redis Cluster MULTI/EXEC does not recover from MOVED after master failover

2 participants