Skip to content

fix(workspace): index the unread count so a poll doesn't scan every message (#3064) - #3076

Merged
dolho merged 1 commit into
devfrom
fix/3064-portal-unread-index
Sep 29, 2026
Merged

dolho merged 1 commit into
devfrom
fix/3064-portal-unread-index

Conversation

@dolho

@dolho dolho commented Sep 29, 2026

Copy link
Copy Markdown
Contributor

Summary

  • count_unread_by_session runs on every 20s Workspace poll, per open tab. Its message arm filters enterprise_portal_messages by the viewer, role = 'assistant', the thread and created_at past a read cursor, and no index led on client_email. The plan read every message in the install.
  • New index: idx_portal_messages_unread ON enterprise_portal_messages(client_email, role, session_id, created_at). Equality columns come first, then the GROUP BY key, then the range column. Index-only: no query change, no behaviour change.

Dual track (Rule #9)

  • db/schema.py INDEXES, for fresh installs on both engines.
  • SQLite: versioned migration portal_messages_unread_index in db/migrations.py (CREATE INDEX IF NOT EXISTS, idempotent).
  • PostgreSQL: Alembic 0081_portal_messages_unread_idx, down_revision = "0080_agent_skill_sets". check_alembic_heads.py reports 82 revisions and 1 head, and check_alembic_parity.py origin/dev HEAD passes.
  • db/tables.py is unchanged: the portal indexes are declared in schema.py DDL, not as Index objects, matching idx_portal_messages_convo and idx_portal_messages_session.

Query plans (the live query, captured at the cursor)

SQLite

before: SCAN m USING INDEX idx_portal_messages_session
after:  SEARCH m USING COVERING INDEX idx_portal_messages_unread (client_email=? AND role=?)

PostgreSQL 16: scratch container, 120k messages, 400 viewers, 20 agents, ANALYZEd, EXPLAIN (ANALYZE, BUFFERS).

before: Seq Scan on enterprise_portal_messages m
          Rows Removed by Filter: 119859   Buffers: shared hit=2269   Execution Time: 6.451 ms
after:  Bitmap Heap Scan on enterprise_portal_messages m
          Index Cond: ((client_email = …) AND (role = 'assistant'))
          Buffers: shared hit=143          Execution Time: 0.868 ms

With a single agent in the table, PostgreSQL already used idx_portal_messages_convo (a non-leading client_email condition, so it still walks the whole index). The seq scan appears once messages span several agents, which is the production shape.

Test Plan

  • tests/unit/test_3064_portal_unread_index.py (new; red before the change):
    • EXPLAINs the SQL count_unread_by_session actually sends, captured via before_cursor_execute, on a sqlite built from the real schema.py DDL, and asserts the message table is not scanned and the new index is used.
    • A negative control: drop only the new index and the same query scans again.
    • The upgrade track: the SQLite migration adds the index to an install that predates it, and running it twice is a no-op.
    • The Alembic revision carries the same DDL and a downgrade.
  • 177 passed on seeds 12345 and 99999: the new file plus schema_parity, migrations, alembic_parity_guard, alembic_revision_id_length, and the unread suites ent359_portal_chat_state and ent557_unread_never_opened_chat. Unread counts are unchanged.
  • pg-migrations CI runs the real alembic upgrade head.

Merge-order note: open PRs #2984, #3021 and #3035 also add an 0081_* revision on 0080_agent_skill_sets. Whichever merges second re-chains onto the live head; alembic-head-watch will flag it.

Fixes #3064

🤖 Generated with Claude Code

…essage (#3064)

count_unread_by_session runs on every 20s Workspace poll, per open tab. Its
message arm filters enterprise_portal_messages by the viewer, role =
'assistant', the thread and created_at past a read cursor. No index led on
client_email, so the plan read every message in the install:
- SQLite: `SCAN m USING INDEX idx_portal_messages_session`.
- PostgreSQL: a Seq Scan once rows span more than one agent.

New index idx_portal_messages_unread ON enterprise_portal_messages(client_email,
role, session_id, created_at): the equality columns first, then the GROUP BY
key, then the range. Index-only; no query or behaviour change.

Dual track (Rule #9):
- db/schema.py INDEXES, for fresh installs on both engines;
- the versioned SQLite migration `portal_messages_unread_index`;
- Alembic `0081_portal_messages_unread_idx` on `0080_agent_skill_sets`
  (single head).

Plans measured on the live query:
- SQLite: SCAN m → SEARCH m USING COVERING INDEX idx_portal_messages_unread
  (client_email=? AND role=?).
- PostgreSQL 16 (120k rows, 400 viewers, 20 agents, ANALYZEd): Seq Scan with
  119,859 rows filtered, 2,269 buffers, 6.45 ms → Bitmap Index Scan with
  Index Cond (client_email, role), 143 buffers, 0.87 ms.

Test: test_3064 EXPLAINs the SQL count_unread_by_session actually sends, which
is captured at the cursor rather than copied. It asserts no scan of the message
table, with a negative control proving the assertion can fail. It also covers
the SQLite migration on an install that predates the index (idempotent) and the
Alembic revision. It was red before the change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@obasilakis obasilakis left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Validated. Index matches the count_unread_by_session predicate (seek on client_email + role, covering session_id/created_at), both migration tracks present, single Alembic head, pg-migrations green. Mutation check: reverting the fix turns 3/4 of test_3064 red (the 4th is the negative control). Merge-order note: #2984/#3021/#3035 also add 0081 on 0080 — disjoint tables, mechanical re-parent for whichever lands later.

@dolho
dolho merged commit e7b239a into dev Sep 29, 2026
24 checks passed
dolho added a commit that referenced this pull request Sep 29, 2026
…/ent703-pull-heartbeat

dev gained 0081_portal_messages_unread_idx (#3076). Resolve:
- db/migrations.py: keep dev's portal_messages_unread_index entry, then pull_sync.
- Alembic: 0081_pull_sync -> 0082_pull_sync, down_revision
  0081_portal_messages_unread_idx, so the version line keeps a single head.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
dolho added a commit that referenced this pull request Sep 29, 2026
dev gained 0081_portal_messages_unread_idx (#3076). Resolve:
- db/migrations.py: dev's portal_messages_unread_index entry, then
  seat_ask_class_state_table.
- Alembic: 0081_seat_ask_class_state -> 0082_seat_ask_class_state, chained
  off 0081_portal_messages_unread_idx (single head); references updated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
dolho added a commit that referenced this pull request Sep 29, 2026
…on to 0082

dev added 0081_portal_messages_unread_idx (#3076) off the same parent.
Rename the Alembic revision to 0082_ent720_email_identity, parented on
0081_portal_messages_unread_idx, and order the SQLite entry after dev's.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants