Skip to content

feat(server): split liveness and readiness probes - #512

Closed
Yumeio wants to merge 1 commit into
FlashML-org:mainfrom
Yumeio:codex/server-health
Closed

Yumeio wants to merge 1 commit into
FlashML-org:mainfrom
Yumeio:codex/server-health

Conversation

@Yumeio

@Yumeio Yumeio commented Sep 18, 2026 •

Copy link
Copy Markdown

Add /healthz for HTTP liveness and /readyz for admission readiness. Readiness returns 200 only while serving with no fatal worker error; loading, rebuilding, failed and stopping return 503 with the existing lifecycle body. The desktop /health endpoint retains its existing 200 status and body in every phase.

Validation on Windows x86-64 / Python 3.13.5 (CPU HTTP tests, no model or GPU): 8 tests passed in tests/server/test_health.py. Tests pin legacy response bodies, startup progress, phase transitions, backend-independent liveness and the worker-death/serving race. The existing unmodified server baseline also passed 535 tests.

Reproduce in an installed development environment: PYTHONPATH=python pytest -q tests/server/test_health.py. This host used the source tree plus isolated missing test dependencies; no engine launch, checkpoint inference or GPU gate was run. An inherited Ruff configuration reports an existing Callable-import style finding in control_api.py; that unrelated import was preserved.

Design comparison: docs/server-hardening/02-health.md. Independent of the auth PR. No breaking changes and no linked issue was supplied.

AI-assisted draft requested by the repository consumer. Human review and a real-server smoke test remain required before merge under CONTRIBUTING.md; this draft does not claim that review has happened.

Combined review evidence

All seven independent changes were also merged on the integration branch. The combined server and targeted scheduler run passed 779 tests; the branch records the configuration/error-handler conflict resolutions and detailed validation scope.

Validation host: Windows/Python 3.13.5 with existing PyTorch 2.13.0+cu126, outside the project's pinned 2.11 range. Results are scoped CPU server/state-machine evidence; they do not establish compatibility with the pinned production environment or real GPU/model execution.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant