Repository navigation
/health endpoint returns status 200 when model is not yet loaded #537
Description
Activity
This is a workaround, not a fix. The bug described here still stands and still needs fixing upstream — the wrapper below just lets anyone blocked by it get moving in the meantime.
Confirming the report independently on v0.1.3 (release build, not source) with
Qwen3.6-35B-A3B-FP8on Linux / CUDA 13 / RTX 4090, so it isn't specific to a model or to building frommain.There is no alternative endpoint to point a proxy at
I checked every GET endpoint the server exposes while the model was loading:
endpoint status during load /health200 — body {"status":"loading","phase":"expert_banks","progress":{…}}/v1/models200 /v1/stats200 /v1/cache/status200 /metrics,/ping404 (not implemented) POST /v1/chat/completionsdoes correctly return 503{"error":"model is still loading"}, and/healthflips to{"status":"ok"}once ready. But since no GET returns a non-success code while loading, there's nothing to redirect a proxy'scheckEndpointto — which rules out the obvious workaround and is why something more involved is needed.Why raising the health-check timeout doesn't help
llama-swap decides readiness purely from the HTTP status of
checkEndpoint— no body matching, no retry on an upstream error. The sequence is:- llama-swap starts
ft serve /healthreturns 200 within ~2 seconds- llama-swap marks the model ready and forwards the queued request
- the caller gets 503
model is still loading— in my case ~90 seconds before the model was usable
healthCheckTimeoutnever comes into play, because the check passes immediately rather than timing out. SettingcheckEndpoint: noneis strictly worse.Workaround: hold the public port closed until the server is actually ready
Run
ft serveon a private port, and only open the port the proxy was told to use once/healthreportsok. While the public port is closed the proxy's health check gets connection-refused, which it treats as "not up yet" and retries — so the proxy's existing wait loop does the waiting, bounded by its ownhealthCheckTimeout. No patch to FreeToken, no patch to llama-swap.#!/usr/bin/env bash # ft-gated PUBLIC_PORT [ft serve args...] set -uo pipefail PUBLIC_PORT="${1:?usage: ft-gated PUBLIC_PORT [ft serve args...]}"; shift INTERNAL_PORT=$((PUBLIC_PORT + 1000)) # must not collide with the proxy's port range SRV_PID=""; FWD_PID="" cleanup() { [[ -n "$FWD_PID" ]] && kill "$FWD_PID" 2>/dev/null [[ -n "$SRV_PID" ]] && kill "$SRV_PID" 2>/dev/null [[ -n "$SRV_PID" ]] && wait "$SRV_PID" 2>/dev/null # let it release VRAM return 0 } trap cleanup EXIT INT TERM ft serve --host 127.0.0.1 --port "$INTERNAL_PORT" "$@" & SRV_PID=$! while true; do kill -0 "$SRV_PID" 2>/dev/null || { echo "ft serve exited during startup" >&2; exit 1; } status=$(curl -sf --max-time 5 "http://127.0.0.1:$INTERNAL_PORT/health" 2>/dev/null \ | python3 -c 'import sys,json; print(json.load(sys.stdin).get("status",""))' 2>/dev/null) [[ "$status" == "ok" ]] && break sleep 2 done socat "TCP-LISTEN:$PUBLIC_PORT,fork,reuseaddr,bind=127.0.0.1" "TCP:127.0.0.1:$INTERNAL_PORT" & FWD_PID=$! wait -n "$SRV_PID" "$FWD_PID" # exit if either dies, so the proxy sees the model stop
llama-swap then calls the wrapper instead of the server directly:
models: "Qwen3.6-35B-A3B-FP8": cmd: | /path/to/ft-gated ${PORT} --model /path/to/Qwen3.6-35B-A3B-FP8
Verified end to end: cold start took 98s and the first request returned normally instead of 503.
Two things that will bite you if you adapt this:
- Don't
exec socat. That replaces the shell and discards the trap, soft serveis orphaned when the proxy sends SIGTERM and its VRAM is never released — which then breaks the next model load. Background the forwarder andwait -non both PIDs. - Bail out if the server dies during startup. Otherwise a crash-on-load is indistinguishable from slow loading and just hangs until the proxy's timeout.
Costs one extra TCP hop for all traffic including token streaming, which is negligible next to generation latency. It can be deleted outright once
/healthreports readiness properly.On the actual fix
Agreed with the proposal here — gating
/healthonmaintenance_state != "serving"would match what llama.cpp, vLLM, TGI and the others do, and is what proxies already expect. The one thing worth deciding deliberately is whether anyone is currently treating that 200 as a liveness signal rather than a readiness one, since for them this would be a breaking change. If that's a concern, a separate/readyendpoint (or/health?ready=1) would give proxies what they need without changing existing behavior. Either way the current state — 200 while the API layer is returning 503 — seems clearly wrong.- llama-swap starts
A supervisor that treats a listening port as readiness has the same problem as a
/healthprobe: the port is bound long before the model is loaded.I added an opt-in
--bind-when-ready: the API port is bound only once the backend is serving, and if the load fails the server exits with the reason instead of binding at all, so a listening port means a serving model. It is off by default, which keeps today's early bind and the load progress on/health. It does not change/health's status codes, which is what this issue asks for; it covers the supervisors that watch the port.Branch, with tests in
tests/server/test_bind_when_ready.py: https://github.com/YevheniiKotyrlo/FreeToken/tree/feat-bind-when-readyI can open it as a PR if you want it, or fold it into whatever fix lands here.
I would like to add my support for this issue. Currently, the fact that the /health endpoint returns a 200 status while the model is still loading creates significant integration issues with proxies like llama-swap, which rely on HTTP status codes to determine service readiness.
While the current behavior might be considered "cosmetic" by some, it deviates from the standard established by other major inference engines such as llama.cpp and vLLM. Aligning FreeToken with these industry standards by returning a 503 error during the loading phase would resolve these compatibility issues directly at the source. Implementing this would be much cleaner than requiring users to rely on complex custom workarounds or wrapper scripts to manage service startup.
I traced the current /health implementation and its internal consumers. Changing /health to return 503 whenever the engine is not serving would directly address the proxy/readiness problem described here, but it also has some compatibility impact inside FreeToken.
shell/client.py already handles an HTTP error response by parsing the JSON body and can still observe lifecycle states such as loading. However, daemon/proxy.py currently turns any HTTP error from the serve /health endpoint into {"status": "error", ...}, which would discard the existing loading phase/progress information. ft ctl health also currently treats a 503 response as an HTTP error rather than a lifecycle document.
One option would therefore be to update those consumers together with /health, preserving the lifecycle body while using 200 only when the engine is actually ready. Another option is to keep /health as the lifecycle/progress endpoint and add a dedicated readiness endpoint that returns 200 only when maintenance_state == "serving" and no fatal error is latched.
The latter is similar to the approach explored in #512 (/readyz), although that PR was a draft closed by its author and did not receive maintainer review.
Since this is an API contract / compatibility decision rather than just a status-code change, would maintainers prefer changing the existing /health semantics, or keeping it backward-compatible and adding a dedicated readiness endpoint?I went with the first option instead of my
--bind-when-readybranch: makes/healthanswer 503 while the server loads, rebuilds its cache, stops or has failed, and 200 only while it accepts requests, with the same JSON body in every state. It updates the consumers you traced:ft ctl healthprints the document from the 503, and the daemon's/engine/healthkeeps showing the load instead oferror;shell/client.pyalready read it. With llama-swap v262 and no extra flag, the first request now returns 200 after the load; onmainthe same config answered 503model is still loading6.3 s in.
Before you start
mainwhen building from source.What happened
With most other inference engines ( llama.cpp, tabbyapi, ninfer, ds4, colibri, ik_llama.cpp, and vllm) a GET request to /health will return an error status such as 404 or 503 if the model is not yet loaded and 200 only when ready, which is a behavior other tools such as llama-swap rely on a non-successful error code to deterimine the health of the engine.
FreeToken deviates from every other implementation i could find by always returning status 200 from /health even when the model is not ready, which makes it look like a bug.
Behavior can easily be seen via running curl while model is loading:
note: status is "loading" but status is "200 OK", where instead something like 503 would be expected instead.
I believe we should set the error code in the /health response until model is actually loaded.
How did you install FreeToken
Built from source
FreeToken version
cab110e
OS
Fedora
OS details
44
GPU and driver
RTX 5090 615.71.09
CPU and system RAM
i7-12600k DDR4
Checkpoint
nvidia/Qwen3.8-Flash-Next-NVFP4
Command
Full log
Anything else
No response