Skip to content

fix(server)!: answer /health with 503 until the server accepts requests - #613

Draft
YevheniiKotyrlo wants to merge 1 commit into
FlashML-org:mainfrom
YevheniiKotyrlo:fix-health-readiness
Draft

YevheniiKotyrlo wants to merge 1 commit into
FlashML-org:mainfrom
YevheniiKotyrlo:fix-health-readiness

Conversation

@YevheniiKotyrlo

Copy link
Copy Markdown
Contributor

Closes #537

Problem

/health answers 200 in every state, including while the model loads, while a cache rebuild runs and after a backend failure; only its body says so. Readiness probes read the status code alone. llama-swap's health check (doStart) passes on the first 200 and forwards the queued request, which the generation routes then refuse. With llama-swap v262 in front of main, the first request came back 503 {"error":"model is still loading"} 6.3 s after it was sent, before the model was serving.

Solution

/health answers 200 only when its document says maintenance: serving, the one state in which _maintenance_gate and new_user admit a request; loading, rebuilding, stopping and failed answer 503. The JSON body is unchanged in every state.

The in-repo consumers read that body from the 503, by the rule shell/client.py already applies (an error response whose JSON body carries a status is an answer, not a transport failure):

  • daemon/proxy.py: /engine/health keeps reporting the serve's document instead of {"status": "error", "httpStatus": 503}.
  • control_cli.py: ft ctl health prints the document. The same rule makes ft ctl cache print a refused rebuild's document, which its non-ok branch already expected but never received.
  • benchmarks/bench_decode_moe.py: wait_ready still stops on a startup error.

This is a breaking change for any client that reads a 200 from /health as liveness: it now sees 503 for the length of a load. The desktop app polls /health; if its client rejects a non-2xx answer, it needs the same one-line rule. I could not check, since its source is not public.

Tests

  • tests/server/test_rebuild_maintenance.py::test_health_is_200_only_while_a_request_would_be_admitted covers loading, serving, rebuilding, stopping, failed and a latched fatal error. In each state, /health answers 200 exactly when openai_api._maintenance_gate admits the request, and its body equals build_health's document. The expectation comes from the gate, not a table, so the two cannot drift apart. On main it fails at loading (200, expected 503).
  • tests/daemon/test_daemon_proxy.py: a health document answered with 503 reads exactly like the same document answered with 200. An error without a status document (plain text, or FastAPI's {"detail": ...}) still reads as {"status": "error", "httpStatus": ...}. On main the first case fails.

Each of these breakages fails a test: /health always 503, /health 503 only while loading, the proxy accepting any JSON object, the proxy dropping the document.

Verification

pytest tests/server tests/daemon: 596 passed. tests/e2e/test_cache_rebuild.py with Qwen/Qwen3-0.6B: 1 passed. Both on WSL2 Ubuntu 22.04, Python 3.10, RTX 3090 Ti (driver 617.14), Ryzen 9 9950X3D.

Against a real server, python -m freetoken --model-path Qwen3-0.6B --host 127.0.0.1 --port <port> --num-pages 4000:

  • /health answered 503 loading to all 181 samples of a 24.5 s load and 200 once serving. It answered 503 rebuilding during POST /v1/cache/rebuild, and 503 error after a failed load until the serve stopped itself.
  • llama-swap v262, the same command with no extra flag: the first request returned 200 after the load, with the health check passing 0.7 s after the ready line. A failed load came back as a 500 once the serve exited.
  • ft ctl health during the load printed status=loading ... and exited 0. ft daemon's /engine/health showed 74 loading documents and no error. ShellClient.wait_until_ready reported 38 progress documents and returned ok. bench_decode_moe.wait_ready returned at ready, and reported the startup error of a failed load.
  • Windows 11, Python 3.13, over fix(windows): build the extensions and JIT kernels with MSVC #575, fix(windows): serve over loopback ZMQ and a selector event loop #580 and fix(windows): read checkpoints without O_DIRECT, posix_fadvise or PROT_READ #581: the same 503-then-200 load, 503 during a rebuild, and 503 until a failed serve stopped.

Related

This replaces the --bind-when-ready branch I offered on #537: with /health answering 503, llama-swap needs no flag.

Readiness probes such as llama-swap's read only the status code, so /health now
answers 503 while the server loads, rebuilds its cache, stops or has failed, with
the same JSON body. ft ctl, the daemon proxy and the decode benchmark read that
body from the 503, as the shell client already did. A client that treats a 200
from /health as liveness sees 503 during a load.
KarrAcaRn pushed a commit to KarrAcaRn/FreeToken-ByAI that referenced this pull request Oct 6, 2026
FlashML-org#618 deferred with FlashML-org#605 (pre-sm70 only); drafts FlashML-org#582, FlashML-org#583, FlashML-org#586, FlashML-org#613 deferred.

Assisted-by: Claude Opus 5.5

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

/health endpoint returns status 200 when model is not yet loaded

1 participant