Repository navigation
fix(server)!: answer /health with 503 until the server accepts requests - #613
Draft
YevheniiKotyrlo wants to merge 1 commit into
Draft
YevheniiKotyrlo wants to merge 1 commit into
YevheniiKotyrlo wants to merge 1 commit into
Conversation
Readiness probes such as llama-swap's read only the status code, so /health now answers 503 while the server loads, rebuilds its cache, stops or has failed, with the same JSON body. ft ctl, the daemon proxy and the decode benchmark read that body from the 503, as the shell client already did. A client that treats a 200 from /health as liveness sees 503 during a load.
KarrAcaRn
pushed a commit
to KarrAcaRn/FreeToken-ByAI
that referenced
this pull request
Oct 6, 2026
FlashML-org#618 deferred with FlashML-org#605 (pre-sm70 only); drafts FlashML-org#582, FlashML-org#583, FlashML-org#586, FlashML-org#613 deferred. Assisted-by: Claude Opus 5.5
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #537
Problem
/healthanswers 200 in every state, including while the model loads, while a cache rebuild runs and after a backend failure; only its body says so. Readiness probes read the status code alone. llama-swap's health check (doStart) passes on the first 200 and forwards the queued request, which the generation routes then refuse. With llama-swap v262 in front ofmain, the first request came back 503{"error":"model is still loading"}6.3 s after it was sent, before the model was serving.Solution
/healthanswers 200 only when its document saysmaintenance: serving, the one state in which_maintenance_gateandnew_useradmit a request; loading, rebuilding, stopping and failed answer 503. The JSON body is unchanged in every state.The in-repo consumers read that body from the 503, by the rule
shell/client.pyalready applies (an error response whose JSON body carries astatusis an answer, not a transport failure):daemon/proxy.py:/engine/healthkeeps reporting the serve's document instead of{"status": "error", "httpStatus": 503}.control_cli.py:ft ctl healthprints the document. The same rule makesft ctl cacheprint a refused rebuild's document, which its non-ok branch already expected but never received.benchmarks/bench_decode_moe.py:wait_readystill stops on a startup error.This is a breaking change for any client that reads a 200 from
/healthas liveness: it now sees 503 for the length of a load. The desktop app polls/health; if its client rejects a non-2xx answer, it needs the same one-line rule. I could not check, since its source is not public.Tests
tests/server/test_rebuild_maintenance.py::test_health_is_200_only_while_a_request_would_be_admittedcovers loading, serving, rebuilding, stopping, failed and a latched fatal error. In each state,/healthanswers 200 exactly whenopenai_api._maintenance_gateadmits the request, and its body equalsbuild_health's document. The expectation comes from the gate, not a table, so the two cannot drift apart. Onmainit fails atloading(200, expected 503).tests/daemon/test_daemon_proxy.py: a health document answered with 503 reads exactly like the same document answered with 200. An error without a status document (plain text, or FastAPI's{"detail": ...}) still reads as{"status": "error", "httpStatus": ...}. Onmainthe first case fails.Each of these breakages fails a test:
/healthalways 503,/health503 only while loading, the proxy accepting any JSON object, the proxy dropping the document.Verification
pytest tests/server tests/daemon: 596 passed.tests/e2e/test_cache_rebuild.pywithQwen/Qwen3-0.6B: 1 passed. Both on WSL2 Ubuntu 22.04, Python 3.10, RTX 3090 Ti (driver 617.14), Ryzen 9 9950X3D.Against a real server,
python -m freetoken --model-path Qwen3-0.6B --host 127.0.0.1 --port <port> --num-pages 4000:/healthanswered 503loadingto all 181 samples of a 24.5 s load and 200 once serving. It answered 503rebuildingduringPOST /v1/cache/rebuild, and 503errorafter a failed load until the serve stopped itself.ft ctl healthduring the load printedstatus=loading ...and exited 0.ft daemon's/engine/healthshowed 74loadingdocuments and noerror.ShellClient.wait_until_readyreported 38 progress documents and returnedok.bench_decode_moe.wait_readyreturned at ready, and reported the startup error of a failed load.Related
This replaces the
--bind-when-readybranch I offered on #537: with/healthanswering 503, llama-swap needs no flag.