Description
Summary
After a coordinated CLUSTER FAILOVER, the previous primary remains reachable but starts returning MOVED for slots owned by the promoted replica.
Regular cluster commands recover by refreshing the slot map and retrying. Atomic MULTI/EXEC does not: it selects a node from the cached topology once and executes directly, bypassing the cluster redirect loop. A transaction-only workload therefore keeps failing until another cluster-routed command, topology refresh, or restart updates the slot map.
Reproduction
Use a cluster with at least two primaries and a replica for the key's primary:
await cluster.set(key, '0');
// Promote the replica using CLUSTER FAILOVER and wait for the role swap.
// Keep the previous primary online and do not issue another command via cluster.
await cluster.multi()
.incr(key)
.get(key)
.exec();
Before the transaction:
- the previous primary still responds to
PING;
- a direct
GET on it returns MOVED;
- the client's cached slot map still points to that previous primary.
Expected:
Actual with node-redis 5.12.1 and 6.2.0:
MOVED <slot> <new-primary>
Every subsequent transaction invocation fails the same way. An ordinary cluster.get(key) refreshes the topology, after which the transaction succeeds.
Root cause
Cluster MULTI calls _executeMulti() directly on the node selected from the cached slot map. Unlike ordinary commands, it does not pass through _execute(), which handles MOVED, ASK, topology rediscovery, and maxCommandRedirections.
Observed versions:
- 4.7.0:
EXECABORT can mask the queued-command MOVED;
- 5.12.1 and 6.2.0: raw
MOVED, without rediscovery.
A retry is safe for this queue-time routing rejection because Redis refuses to execute the dirty transaction. Non-atomic execAsPipeline() must not be retried because it may have been partially applied.
This appears closely related to #3256. PRs #3366/#3377 address batch slot metadata but do not add redirect handling for atomic MULTI.
I have a candidate fix and regression tests ready.
Node.js Version
24.15.0
Redis Server Version
7.4.2
Node Redis Version
6.2.0 (also reproduced with 5.12.1)
Platform
Windows host, Redis in Linux Docker containers
Logs
Description
Summary
After a coordinated
CLUSTER FAILOVER, the previous primary remains reachable but starts returningMOVEDfor slots owned by the promoted replica.Regular cluster commands recover by refreshing the slot map and retrying. Atomic
MULTI/EXECdoes not: it selects a node from the cached topology once and executes directly, bypassing the cluster redirect loop. A transaction-only workload therefore keeps failing until another cluster-routed command, topology refresh, or restart updates the slot map.Reproduction
Use a cluster with at least two primaries and a replica for the key's primary:
Before the transaction:
PING;GETon it returnsMOVED;Expected:
Actual with node-redis 5.12.1 and 6.2.0:
Every subsequent transaction invocation fails the same way. An ordinary
cluster.get(key)refreshes the topology, after which the transaction succeeds.Root cause
Cluster
MULTIcalls_executeMulti()directly on the node selected from the cached slot map. Unlike ordinary commands, it does not pass through_execute(), which handlesMOVED,ASK, topology rediscovery, andmaxCommandRedirections.Observed versions:
EXECABORTcan mask the queued-commandMOVED;MOVED, without rediscovery.A retry is safe for this queue-time routing rejection because Redis refuses to execute the dirty transaction. Non-atomic
execAsPipeline()must not be retried because it may have been partially applied.This appears closely related to #3256. PRs #3366/#3377 address batch slot metadata but do not add redirect handling for atomic
MULTI.I have a candidate fix and regression tests ready.
Node.js Version
24.15.0
Redis Server Version
7.4.2
Node Redis Version
6.2.0 (also reproduced with 5.12.1)
Platform
Windows host, Redis in Linux Docker containers
Logs