Skip to content

fix(flags): bound persons db work with a request deadline - #108017

Draft
haacked wants to merge 1 commit into
haacked/flags-db-timeout-budgetfrom
haacked/flags-persons-db-deadline
Draft

haacked wants to merge 1 commit into
haacked/flags-db-timeout-budgetfrom
haacked/flags-persons-db-deadline

Conversation

@haacked

@haacked haacked commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Problem

  • When the persons database stops answering, SDKs calling /flags get a 503 for every flag after the 4.5s request timeout, including flags that never read persons data.
  • Postgres statement_timeout cannot cancel a query on a database that has stopped answering.
  • Each persons call is bounded only by pool acquire plus statement_timeout, and the calls run in sequence.
  • The degrade path that errors only the persons flags already exists, but it runs only when a call returns an error.

Changes

  • SDKs now get a 200 with errorsWhileComputingFlags: true within the deadline. Only flags that need persons data fail, with reason.code set to timeout:persons_db_deadline.
  • PERSONS_DB_DEADLINE_MS (default 2500, 0 disables) sets one deadline for all persons DB calls in a /flags evaluation.
  • It covers the hash key override check, write, and read, and the person, cohort, and group properties fetch.
  • The group type mapping lookup stays outside the deadline. A plain timeout there cancels a fetch that other requests share, so fix(flags): apply the persons db deadline to the group type lookup #108772 adds it with a cache fix.
  • The calls share one Instant, so sequential calls cannot add up past it. The default leaves 2s of the 4.5s request timeout for the rest of the request.
  • A call that would start after the deadline fails without taking a connection. timeout_at alone would start the call and then abandon it.
  • The internal batch evaluation endpoint does not apply the deadline. Django drops a person whose evaluation errors from the static cohort and still reports success, so a slow query must finish there.
  • Each stopped call increments the existing flags_database_error_total counter with timeout_type="persons_db_deadline" and its operation.
  • Other properties fetch errors, such as a statement timeout, now also reach flags_database_error_total under operation="fetch_properties". Before, only the hash key calls reported there.
  • The canonical log field persons_db_deadline_exceeded records the first stopped call.
  • before_persons_db_deadline boxes each persons DB future. Without the box, the wrapper layers grow the /flags handler future past the 2 MiB debug test stack.
  • Look first at before_deadline in utils/deadline.rs and before_persons_db_deadline in flag_matching.rs. The rest is wiring and tests.

Before:

flowchart LR
  A["POST /flags"] --> B[hash key override read]
  B --> C[properties fetch]
  C --> D[evaluate flags]
  B -. database frozen .-> T[4.5s request timeout]
  C -. database frozen .-> T
  T --> E["503 for every flag"]
  classDef phBlue fill:#1d4aff,stroke:#1d4aff,color:#fff;
  classDef phRed fill:#f54e00,stroke:#f54e00,color:#fff;
  classDef phYellow fill:#f9bd2b,stroke:#f9bd2b,color:#000;
  class A,E phYellow;
  class B,C,D phBlue;
  class T phRed;
Loading

After:

flowchart LR
  A["POST /flags"] --> B[hash key override read]
  B --> C[properties fetch]
  C --> D[evaluate flags]
  B -. database frozen .-> X[2.5s persons DB deadline]
  C -. deadline passed .-> X
  X --> D
  D --> E["200, persons flags errored"]
  classDef phBlue fill:#1d4aff,stroke:#1d4aff,color:#fff;
  classDef phRed fill:#f54e00,stroke:#f54e00,color:#fff;
  classDef phYellow fill:#f9bd2b,stroke:#f9bd2b,color:#000;
  class A,E phYellow;
  class B,C,D phBlue;
  class X phRed;
Loading

Note

This PR stacks on #108014 and must merge after it. #108014 sets the persons reader statement timeout to 1000ms, the writer statement timeout to 2000ms, and the acquire timeout to 1s. Each sits below the 2.5s deadline, so Postgres cancels an ordinary slow query before the deadline drops it. Without #108014, the deadline would drop 2.5 to 3s queries that finish today, and each dropped query keeps its connection until sqlx's release ping finishes.

Known limits:

  • The deadline starts when the matcher is built. Time spent waiting for a concurrency permit counts against the 4.5s request timeout but not against the deadline, so a saturated pod can still return 503s.
  • When the pool has no free connection, the 1s acquire timeout can fire first, and the reason is timeout:pool_timeout. Dashboards should count both codes.
  • On a frozen database, a request that misses the group type cache still waits for the request timeout and returns a 503. fix(flags): apply the persons db deadline to the group type lookup #108772 fixes this.
  • A slow group type lookup still counts against the deadline. The properties fetch after it can then fail even when person queries are fast.
  • A flag that depends on a persons flag through flag_evaluates_to returns a normal-looking value, with no failed: true, when the persons flag errors. That behavior is unchanged, but a stall now reaches it through a 200.

How did you test this code?

  • test_stalled_persons_db_degrades_within_one_deadline runs with and without $anon_distinct_id, on paused time, against a fake database that never answers. It fails if a call loses its deadline (hang), if each call gets its own timeout, or if a call starts after the deadline.
  • The same test asserts the persons_db_deadline_exceeded log field. It fails if the log records a later call than the first one stopped.
  • it_degrades_person_flags_when_the_persons_db_stops_answering runs the real server with the persons read URL pointed at a socket that never replies. It fails if the config does not reach the matcher.
  • Each test was confirmed to fail by temporarily reintroducing the regression it targets.
  • Stack headroom: in a local debug build, it_rejects_invalid_token passes at a 1.75 MiB stack and overflows at 1.6875 MiB, the same as the merge base. CI on Linux is the real check.
  • Manual run, before the rebase onto fix(flags): fit db timeouts inside the request timeout #108014, so with the 3s persons statement timeout and the 2s acquire timeout. It used a local debug build against the dev Postgres, with the persons URLs behind a TCP proxy that can stop forwarding.
    • A lock on posthog_person returns a 200 with person flags at timeout:persons_db_deadline. flags_database_error_total and the canonical log field name the stopped call.
    • With PERSONS_DB_DEADLINE_MS=0, the same lock returns timeout:query_canceled at the 3s statement timeout.
    • A proxy frozen while a query is in flight: the request still returns in 2.7s. The dropped connection stays open until the thaw, then closes.
    • A proxy frozen while the pool is idle: the first acquire hits the 2s pool timeout. The retry stops at the deadline, and flags report timeout:persons_db_deadline.
    • The batch endpoint waits out a 2.8s lock and returns its match with errors_count: 0.
    • Responses ran about 0.5s past the deadline, which is the work before the matcher starts (see known limits).
  • Not rerun after the rebase. With fix(flags): fit db timeouts inside the request timeout #108014's 1s reader statement timeout, Postgres cancels the lock case before the deadline. The frozen proxy cases do not depend on the statement timeout.

👉 Stay up-to-date with PostHog coding conventions for a smoother review.

Release status

  • No feature flag controls this change
  • This change is behind a feature flag and is not available to users
  • This change makes a previously flagged feature available to everyone

Automatic notifications

  • Publish to changelog?

Docs update

  • Adds a persons DB deadline section to docs/internal/feature-flags/database-interaction-patterns.md. No public docs change.

🤖 Agent context

Autonomy: Human-driven (agent-assisted)

Agent: Claude Code, Claude Opus 5.5 (claude-opus-5-5)

  • Implemented from the assignee's plan, which set the scope, the default, and the test approach. Skills invoked: /writing-tests, /writing-code-comments, /reviewing-with-coderabbit, /writing-pr-descriptions, /review-code, /simplify, /stacking-prs.
  • A separate agent review found a comment that claimed sqlx closes a connection dropped mid-query. On a frozen database sqlx keeps it checked out. Fixed.
  • The batch endpoint first applied the deadline, as the plan asked. The assignee chose to remove it after the Django side turned out to drop errored persons from the cohort without failing the run.
  • CodeRabbit CLI pass with --deep, three findings. Two duplicates asked to lower the persons statement timeouts below the deadline, which fix(flags): fit db timeouts inside the request timeout #108014 does. The third concerned the group type cache and moved with that code to fix(flags): apply the persons db deadline to the group type lookup #108772.
  • PR review round: Greptile's finding concerned the group type cache and moved with that code to fix(flags): apply the persons db deadline to the group type lookup #108772. ReviewHog raised the deadline start point (left as a known limit) and two pre-existing silent-false cases for dependent flags and group flags, which are tracked as follow-ups.
  • A /review-code --fix pass found that the earlier pushes overflowed the stack in 47 feature-flags integration tests. Fixed by boxing the persons DB future in before_persons_db_deadline. The same pass extended the stall test.
  • Scope review: the group type cache change went past the plan's scope, so the assignee moved it to fix(flags): apply the persons db deadline to the group type lookup #108772. This PR now stacks on fix(flags): fit db timeouts inside the request timeout #108014, which must merge first.
  • /simplify replaced a dedicated flags_persons_db_deadline_exceeded_total counter with flags_database_error_total. Deadline stops now share one counter and one set of operation names with other persons DB errors.
  • No duplicate: searches for open PRs on persons deadlines and flags persons timeouts found none.
  • Public artifact: the work started from an internal investigation of a persons database stall. No numbers, hostnames, or log content from it appear in the code, tests, or this description, and the test data is invented.
  • Patch coverage is not checked yet. It waits on CI.

@haacked haacked self-assigned this Sep 28, 2026
@trunk-io

trunk-io Bot commented Sep 28, 2026

Copy link
Copy Markdown

Merging to master in this repository is managed by Trunk.

  • To merge this pull request, check the box to the left or comment /trunk merge below.

After your PR is submitted to the merge queue, this comment will be automatically updated with its status. If the PR fails, failure details will also be posted here

@github-actions

github-actions Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

🤖 CI report

✅ Trunk lane — non-backend lane (rust:crate:feature-flags)

This PR is assigned to the non-backend lane (rust:crate:feature-flags). It does not run backend Python tests and may merge in parallel with PRs in other lanes.

🚨 Comment density — 10% of added code lines are comments (32 of 336)

This section warns when comments are more than 3% of the code lines a PR adds, and alerts above 6%. Before agent-assisted PRs, the typical share was about 2%. Only full-line comments count. Docstrings, generated files, snapshots, migrations, and workflow files are left out.

Comments that restate the code, record how the change came about, or narrate the next line add noise for the next reader. Keep the comments that explain a reason the code cannot show, and remove the rest. See .agents/skills/writing-code-comments/SKILL.md for the house rules.

Files with the most added comment lines:

File Comment lines Added lines
rust/feature-flags/src/flags/flag_matching.rs 9 97
rust/feature-flags/src/config.rs 7 14
rust/feature-flags/src/utils/deadline.rs 4 21
rust/feature-flags/src/api/batch_flag_evaluation.rs 3 3
rust/feature-flags/src/flags/test_flag_matching.rs 2 96
rust/feature-flags/src/handler/canonical_log.rs 2 5
rust/feature-flags/tests/test_flags.rs 2 55
rust/feature-flags/src/api/errors.rs 1 8

This check does not block merging. It updates on every push and clears when the share drops.

@haacked haacked added the reviewhog ($$$) Reviews pull requests before humans do label Sep 28, 2026
@haacked
haacked requested a lite review from Copilot September 28, 2026 21:05
@posthog

posthog Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

🦔 PostHog Review reviewed this pull request

Found 0 must fix, 2 should fix, 1 consider.

Published 3 findings (view the review).

Resolved comments: 3 left for you

@coderabbitai

coderabbitai Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository: PostHog/posthog/.coderabbit.yaml

Review profile: QUIET

Plan: Enterprise

Run ID: bfe02d2b-bccf-4d79-8697-51fa5278d9a4

📥 Commits

Reviewing files that changed from the base of the PR and between 513d6f9 and c6a8c40.

📒 Files selected for processing (7)
  • docs/internal/feature-flags/database-interaction-patterns.md
  • rust/feature-flags/src/config.rs
  • rust/feature-flags/src/flags/flag_matching.rs
  • rust/feature-flags/src/flags/flag_matching_utils.rs
  • rust/feature-flags/src/flags/test_flag_matching.rs
  • rust/feature-flags/src/handler/canonical_log.rs
  • rust/feature-flags/src/utils/test_utils.rs

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 9 remain after this review.


📝 Walkthrough

Walkthrough

The changes add a configurable shared deadline for persons-database calls during live flag evaluation. The matcher applies it to hash-key override operations and property fetches, and reports deadline errors by operation. Expired calls produce partial flag results, while batch matching does not use this deadline. Tests cover stalled database calls and the resulting flag and error responses.

Priority: ⬇️ Low

Merge Risk: 🟡 Moderate · up to c6a8c

The deadline preserves partial results during database stalls, but flags depending on a timed-out flag can report misleading successful values. Resolve dependency-error propagation before merging, or explicitly accept this bounded correctness risk.

Security Architecture Review

Security architecture risk: 🟡 Moderate · up to c6a8c

The deadline improves outage isolation, but a flag can return a successful-looking result after its prerequisite times out. Database cancellation also has unresolved recovery behavior. No authorization bypass or cross-tenant access expansion was established.

Retained concerns

  • Medium · reliability · inferred: A persons-data preparation timeout errors the prerequisite flag but does not propagate that failure to dependency-only consumers. An absent prerequisite value becomes a condition non-match, so a dependent can return a successful false result or match a later catch-all. The response still carries errorsWhileComputingFlags, but its dependent per-flag results can appear authoritative. This behavior predates the PR for independently returned database errors; the new deadline expands its reachability by producing a partial response before the outer request timeout.
Security review details

Security Blast Radius

  • inferred — The evidenced result-level exposure is to flags evaluated for the selected team, including downstream flag dependencies. Connection recovery affects other requests sharing the same service-instance pools. Pool aliasing can broaden that availability coupling, but the inspected comparison does not establish newly introduced cross-tenant data access or privilege gain.

Security Findings and Attack Paths

  • inferred — The supported failure path is deadline expiry, prerequisite preparation error, absent dependency value, then non-match or fallback evaluation. This establishes a failure-containment concern, not a verified authorization exploit. The evidence does not establish an attacker’s ability to stall PostgreSQL or a downstream authorization sink.

Trust Boundaries and Controls

  • observed — The deadline comes from server configuration rather than request input. Direct property evaluation retains fail-closed handling for pending data, and preparation failures set the aggregate error bit. Static missing-dependency handling covers metadata defects, but is distinct from runtime prerequisite timeout propagation.

Resilience and Maintainability Implications

  • inferred — Transactions and conflict suppression mitigate partial insertion and repeated-write hazards on completed operations. However, a deadline error does not establish that an in-flight COMMIT failed or that connection cleanup completed. The inspected stalled-client test checks response behavior and connection-request count, not persisted write outcomes or real pool recovery. Those recovery states remain unverified rather than observed failures.

Hardening Proposals

  • proposed — Keep runtime prerequisite failure distinct from a legitimate false value and define how that failure propagates to dependent flags. Separately specify cancellation recovery for commit ambiguity and pool release before treating the response deadline as a guarantee that database work has stopped.
🚥 Pre-merge checks | ✅ 1
✅ Passed checks (1 passed)
Check name Status Explanation
Description check ✅ Passed The description is complete and standalone. It covers the problem, user-visible changes, deadline behavior, known limits, tests, release status, documentation, agent context, and stacked-PR dependency…
✨ Finishing Touches
📝 Generate docstrings
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR

Comment @coderabbitai help to get the list of available commands.

@trunk-io

trunk-io Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Static Badge   Static Badge   Static Badge

Failed Test Failure Summary Logs
flags::feature_flag_list::tests::test_fetch_flags_from_redis Logs ↗︎

View Full Report ↗︎ ⋅ Docs

@posthog

posthog Bot commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

PostHog Review alpha 🦔 If you find any issues helpful - please reply "valid", "invalid", etc., for evaluation purposes 🙏

@posthog posthog Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

PostHog Review

Found 2 should fix, 1 consider.

Comment thread rust/feature-flags/src/flags/flag_matching.rs
Comment thread rust/feature-flags/src/flags/flag_matching.rs
Comment thread rust/feature-flags/src/flags/flag_matching.rs Outdated
@posthog posthog Bot removed the reviewhog ($$$) Reviews pull requests before humans do label Sep 28, 2026
@greptile-apps

greptile-apps Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Retrigger

[High risk] Adds request-scoped timeout to persons database operations.

The PR should not merge until coalesced group-mapping fetches stop timing out requests before their own deadlines.

Reviews (1) · Last reviewed commit: "fix(flags): bound persons db work with a..."

Comment thread rust/feature-flags/src/flags/flag_group_type_mapping.rs Outdated

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot encountered an error and was unable to review this pull request. You can try again by re-requesting a review.

@haacked
haacked force-pushed the haacked/flags-persons-db-deadline branch 3 times, most recently from b6b47dd to 096445d Compare September 29, 2026 15:59
Add PERSONS_DB_DEADLINE_MS (default 2500ms, 0 disables). All persons DB
calls in one flag evaluation share this deadline: the hash key override
check, write, and read, and the person, cohort, and group properties
fetch. When it passes, the call fails with timeout:persons_db_deadline
and the existing degrade path errors only the flags that need persons
data. A persons database that stops answering now yields a partial 200
instead of a 503 at the request timeout.

A call that would start after the deadline fails without taking a
connection. The deadline wrapper boxes each persons DB future, so the
/flags handler future stays within a 2 MiB stack in debug builds.

Expiries increment flags_database_error_total with
timeout_type="persons_db_deadline" and the call's operation, and set
persons_db_deadline_exceeded on the canonical log line. Other properties
fetch errors also increment flags_database_error_total, under operation
fetch_properties.

The internal batch evaluation endpoint does not apply the deadline,
because Django drops a person whose evaluation errors from the static
cohort without failing the run.
@haacked
haacked force-pushed the haacked/flags-persons-db-deadline branch from 096445d to c6a8c40 Compare September 29, 2026 21:11
@haacked
haacked changed the base branch from master to haacked/flags-db-timeout-budget September 29, 2026 21:11
@haacked
haacked added this pull request to stack #108773 September 29, 2026 21:22

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

2 participants