Repository navigation
Conversation
Adds a new automatic allocation backend that submits Slurm allocations through a FirecREST v2 API (https://eth-cscs.github.io/firecrest-v2/) instead of executing sbatch locally. This allows the HyperQueue server to run outside of the target cluster (e.g. on a VM), which is needed on clusters with login-node wall-time limits, such as CSCS Alps. - New ManagerType::Firecrest and FirecrestQueueParams (cluster-side paths for the hq binary, worker access file and allocation workdir; OAuth2 client credentials, with the secret read from an environment variable of the server and never persisted) - New queue handler performing submit/status/cancel via HTTPS with a cached OAuth2 client-credentials token; reuses the Slurm submit script builder, and spawned workers use --manager slurm since FirecREST allocations are ordinary Slurm jobs - reqwest (rustls) dependency gated behind a new `firecrest` cargo feature, enabled by default; servers built without it report a clear error for FirecREST queues Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…rnal restore - The OAuth2 token cache is now guarded by an async mutex held across the fetch, so concurrent requests reuse one fetch instead of each fetching their own token. Any API call rejected with HTTP 401 (e.g. a token revoked before its expiration time) invalidates the cached token and is retried once with a freshly fetched one. - Allocation status refresh now uses a single job-list request per queue instead of one request per allocation, keeping API (rate-limit) pressure independent of the number of active allocations. Allocations that already dropped out of the list are still queried individually as a fallback. - Restoring the server from a journal no longer aborts the remaining state restoration when one allocation queue cannot be recreated (for example when the FirecREST client-secret environment variable is not set after a restart); the queue is skipped with an error in the log instead. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C8RVhZJp3g2FjfhN9sikWj
The firecrest backend talks to an HTTP API instead of scheduler binaries, so the tests use an in-process mock of the token endpoint + FirecREST v2 API (tests/autoalloc/mock/firecrest.py) running on an ephemeral localhost port, with request counters and fault injection. Covered: submit script contents, dry-run success and token/submit errors, allocation state lifecycle (queued/running/finished/failed), cancellation on queue removal, batched status queries, transparent retry with a fresh token on HTTP 401, journal restore with and without the secret environment variable, and that the client secret is never persisted to the journal. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C8RVhZJp3g2FjfhN9sikWj
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C8RVhZJp3g2FjfhN9sikWj
- Check the fallback timestamp in finished_without_end_time_falls_back_to_now - Pass the queue time limit explicitly in test_firecrest_submit_script - Document that an off-cluster server needs a public address advertised to workers Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017QUYDeeb8vu1fxjgf4Tdkv
|
I'm no longer very involved in the maintenance of HyperQueue, but I think that the right way forward is to define a general API for autoalloc backends, and let HQ use it with an externally provided backend. Most likely it would make sense to define a JSON REST API, as performance is not super critical here, and that is as general as it gets. That would allow anyone to implement their own autoalloc backend and use it with HyperQueue, rather than it being maintained in-tree. You could still implement it e.g. in Rust, and in this case it would essentially just do a translation between HQ's autoalloc REST API and the FirecREST API. It would require running the wrapper binary together with HyperQueue on some TCP/IP port, but that shouldn't be a big issue, it can just run on a single node togethe with the HQ server. |
|
Disclaimer: I had an offline conversation about this with @rjanalik few weeks ago. I did not have time to dig in the code yet, but as far as I understand, firecrest is basically a kind of API over SLURM. We do not need to invent our own. I am ok to maintaining it directly in HQ, I see it as an alternative to two other backends that we already have. Of course, if there would be like 3-5 other competing APIs in the future, than it is probably time to propose a unifying one for HQ. But in the current situation, having Firecrest directly in HQ makes sense to me. This brings an out-of-the-box solution for people who does not run server on login node and wants to use autoallocator. |
|
Fair enough :) |
Motivation
The PBS and Slurm autoalloc backends run
qsub/sbatchlocally, so the HQ server has to run on thecluster's login node. On clusters with wall-time limits on login nodes (e.g. CSCS Alps), a long-running
server is not possible there. FirecREST is a REST API in front
of Slurm, so submitting allocations through it lets the server run anywhere with HTTPS access to the API
(a VM, a K8s pod, ...). Workers still connect back to the server over TCP as usual.
What this adds
hq alloc add firecrest/hq alloc dry-run firecrest. It takes the usual shared queueoptions plus the FirecREST connection settings:
--api-url,--system,--token-urland--client-id.--client-secret-envnames an environment variable of the server that holds thesecret (default
HQ_FIRECREST_CLIENT_SECRET). The secret itself is never stored in the queuedefinition or the journal.
--remote-hq-path,--remote-server-dirand--remote-workdir. They are usedverbatim on the cluster, because the server makes no assumptions about the cluster's filesystem.
an OAuth2 client-credentials token.
Slurm jobs, so workers start with
--manager slurm.fetches a fresh token and retries once.
keeps API rate-limit pressure low. Allocations missing from the list are queried individually.
isn't set after a restart), restore no longer aborts. That queue is skipped with an error in the log,
and the rest of the state is restored.
firecrestcargo feature, on by default. It pulls inreqwestwith rustls. A server builtwithout the feature reports a clear error for FirecREST queues.
docs/deployment/allocation.md, including the network requirements (the server must bereachable from compute nodes), and a CHANGELOG entry.
Testing
tests/autoalloc/test_firecrest.py) against an in-process mock of the tokenendpoint and the FirecREST v2 API (
tests/autoalloc/mock/firecrest.py), with fault injection. They cover:daint): dry-run, thenalloc add→ a Slurm job submitted via FirecREST → a worker on a compute node connected and ran thetask →
alloc remove --forcecancelled the allocation via the API.