Skip to content

getMessage() returns { conversation: '' } on miss — Baileys relays an empty message and burns the retry, leaving "Waiting for this message" forever #2705

Description

@brHIZ

What happened

getMessage() returns { conversation: '' } when the message is not found. Baileys treats any
truthy return as "message found", so it relays an empty message and consumes one of the retry
attempts
. On the recipient's phone the message becomes a permanent
"Waiting for this message. This may take a moment." placeholder — the retry that was supposed to
recover it delivered an empty envelope instead.

src/api/integrations/channel/whatsapp/whatsapp.baileys.service.ts (still on main today):

private async getMessage(key: proto.IMessageKey, full = false) {
  try {
    const webMessageInfo = (await this.prismaRepository.$queryRaw`
      SELECT * FROM "Message"
      WHERE "instanceId" = ${this.instanceId} AND "key"->>'id' = ${key.id}
    `) as proto.IWebMessageInfo[];
    if (full) return webMessageInfo[0];
    ...
    return webMessageInfo[0].message;
  } catch {
    return { conversation: '' };   // <-- row not found => webMessageInfo[0] is undefined
  }                                //     => `.message` throws => empty text is returned
}

Baileys, Socket/messages-recv.js:

if (msg && (await willSendMessageAgain(ids[i], participant))) {
    updateSendMessageAgainCount(ids[i], participant);   // burns 1 of maxMsgRetryCount (4)
    await relayMessage(key.remoteJid, msg, msgRelayOpts);
} else {
    logger.debug({ jid: key.remoteJid, id: ids[i] }, 'recv retry request, but message not available');
}

Returning undefined instead makes Baileys take the else branch: no attempt is consumed and
nothing is sent
, so the peer can ask again and gets the real message once the row is committed.

Why the row can be missing

sendMessageWithTyping persists the row after the send, after the Chatwoot integration
round-trip and, for media, after writing the media file — while retryRequestDelayMs is 350 ms.
Baileys' in-memory recent-message cache (512 entries) usually covers this, but it is lost on every
socket restart, so after a reconnect or a re-pair the DB path is the only one left. That is exactly
when a burst of retries happens (a re-pair invalidates every peer's session), so the failure
concentrates precisely where it hurts most.

We saw this in production right after a QR re-pair: group media sent 4 minutes later reached every
participant as the "Waiting for this message" placeholder, permanently.

Suggested fix

  } catch (error) {
    this.logger.warn(`getMessage failed for ${key.id}: ${error}`);
    return undefined;
  }

All 7 call sites of this.getMessage( already handle a falsy return — and four of them get
strictly better, because today they silently receive a fake message and carry on:

call site current handling
Baileys getMessage option if (msg && ...) — relays the empty message
messages.edit f?.id
poll updates h && (...) — aggregates votes against a fake message
quoted message S && (c = S) — quotes an empty message
getBase64FromMediaMessage if (!n) throw 'Message not found' — never reached today
formatUpdateMessage t?.messageType
updateMessage if (!i) throw new BadRequestException('Message not found') — never reached today

Also worth noting: the retry lookup filters on "key"->>'id', and there is no index for it — it is
a sequential scan of Message on every retry. Adding
("instanceId", (("key"->>'id'))) took the query from 9.76 ms to 0.158 ms on a 15.8k-row table
here, and it only gets worse as the table grows.

Version

v2.3.6, self-hosted, Postgres, Chatwoot integration enabled.

Activity

  1. augustoadsa commented on Aug 24, 2026

    @augustoadsa

    Independent confirmation on v2.3.7, plus a consequence I think raises the severity of this
    issue considerably: in group chats the retry you are burning is the only mechanism that can
    ever repair a broken sender-key distribution.

    The code is still there in 2.3.7. From the running bundle:

    async getMessage(key, full = false) {
      try {
        const rows = await this.prismaRepository.$queryRaw`
          SELECT * FROM "Message" WHERE "instanceId" = ${this.instanceId} AND "key"->>'id' = ${key.id}`
        if (full) return rows[0]
        if (rows[0].message?.pollCreationMessage) { /* ... */ }
        return rows[0].message
      } catch { return { conversation: '' } }
    }

    Worth noting explicitly: on a miss the query returns [], so rows[0].message?.… throws a
    TypeError and lands in the same catch. The "not found" path and the "database error" path
    therefore converge on the same empty envelope — the miss is swallowed as if it were an error.

    Why this is worse in groups. In baileys@7.0.0-rc.9 — the version this project pins on main
    and on the 2.4.0-rc2 tag — the group sender-key bookkeeping map (sender-key-memory-<jid>@g.us)
    is written to only with true and is never invalidated by participant or identity changes; the
    comment // on participant change in group, we should do sender memory manipulation is still
    unimplemented, in rc.9 and in current master alike. The single place in the whole library that
    clears that map is sendMessagesAgain in src/Socket/messages-recv.ts:

    if (isJidGroup(remoteJid)) {
      await authState.keys.set({ 'sender-key-memory': { [remoteJid]: null } })
    }

    So the retry receipt is not merely one recovery path among several — for groups it is the only
    one
    . Feeding it an empty envelope consumes the recipient's limited retries without ever delivering
    real content, and once they are exhausted no further retry arrives. A transient decryption failure
    becomes permanent.

    What we observe, on a fully LID-addressed group (Evolution 2.3.7, baileys 7.0.0-rc.9, Node
    20.20.2, Redis-backed auth state, ~14 participants, all @lid):

    • Every send is accepted with no error, and a subset of members — all of them iPhone — see only
      "Waiting for this message", indefinitely (> 1 week). Android members read everything. Messages
      posted by humans in the same group are read by everyone, including the affected iPhones.
    • The group's sender-key-memory map is always 100% true — 29, 22, 30 and 22 device entries
      observed on four separate occasions, always zero false. It never self-clears, which is itself
      evidence that the retry path is not doing its job in practice.
    • Deleting that one Redis field restores delivery, confirmed on the affected handsets, 3 out of 3
      times
      (2026-08-03, 08-08, 08-19), and it regresses within days: 4 events in 21 days.
    • One affected iPhone user created a new WhatsApp Business registration and immediately started
      receiving the bot's group messages, with no change on our side — consistent with the map being the
      gate, since a fresh registration is a device not yet marked true.

    We now clear the map hourly from cron as a stop-gap. That is obviously not a fix, and it should not
    be necessary.

    Supporting the proposed change: returning undefined instead of { conversation: '' } is the
    right call. Two suggestions on top of it:

    1. Distinguish the miss from the error — if (!rows.length) return undefined before touching
      rows[0] — so a genuine database failure can be logged instead of silently masquerading as a
      cache miss.
    2. Log at warn when getMessage misses. Today this failure mode is completely silent on the
      sender side: the socket reports success, Message.status never leaves PENDING for group sends,
      and the only signal that anything is wrong is a human saying "I never got it".

    Happy to test a patched build against the live group.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions