Skip to content

Add FirecREST autoalloc backend (hq alloc add firecrest) - #1142

Draft
rjanalik wants to merge 5 commits into
It4innovations:mainfrom
rjanalik:firecrest-autoalloc
Draft

rjanalik wants to merge 5 commits into
It4innovations:mainfrom
rjanalik:firecrest-autoalloc

Conversation

@rjanalik

@rjanalik rjanalik commented Oct 6, 2026

Copy link
Copy Markdown

Motivation

The PBS and Slurm autoalloc backends run qsub/sbatch locally, so the HQ server has to run on the
cluster's login node. On clusters with wall-time limits on login nodes (e.g. CSCS Alps), a long-running
server is not possible there. FirecREST is a REST API in front
of Slurm, so submitting allocations through it lets the server run anywhere with HTTPS access to the API
(a VM, a K8s pod, ...). Workers still connect back to the server over TCP as usual.

What this adds

  • New backend hq alloc add firecrest / hq alloc dry-run firecrest. It takes the usual shared queue
    options plus the FirecREST connection settings: --api-url, --system, --token-url and --client-id.
    • Client secret: --client-secret-env names an environment variable of the server that holds the
      secret (default HQ_FIRECREST_CLIENT_SECRET). The secret itself is never stored in the queue
      definition or the journal.
    • Cluster-side paths: --remote-hq-path, --remote-server-dir and --remote-workdir. They are used
      verbatim on the cluster, because the server makes no assumptions about the cluster's filesystem.
  • Queue handler. It submits, checks status and cancels allocations through the FirecREST v2 API, using
    an OAuth2 client-credentials token.
    • The submit script comes from the existing Slurm script builder. FirecREST allocations are ordinary
      Slurm jobs, so workers start with --manager slurm.
    • The token is cached and shared between concurrent requests. If a request gets HTTP 401, the handler
      fetches a fresh token and retries once.
    • Status is polled with one job-list request per queue rather than one request per allocation, which
      keeps API rate-limit pressure low. Allocations missing from the list are queried individually.
  • Journal restore. If one allocation queue can't be recreated on restore (e.g. the secret variable
    isn't set after a restart), restore no longer aborts. That queue is skipped with an error in the log,
    and the rest of the state is restored.
  • New firecrest cargo feature, on by default. It pulls in reqwest with rustls. A server built
    without the feature reports a clear error for FirecREST queues.
  • Docs in docs/deployment/allocation.md, including the network requirements (the server must be
    reachable from compute nodes), and a CHANGELOG entry.

Testing

  • Rust unit tests for job state mapping and response parsing.
  • Python integration tests (tests/autoalloc/test_firecrest.py) against an in-process mock of the token
    endpoint and the FirecREST v2 API (tests/autoalloc/mock/firecrest.py), with fault injection. They cover:
    • the submit script, and dry-run success, token errors and submit errors
    • the allocation lifecycle (queued, running, finished, failed) and cancellation on queue removal
    • batched status queries and the retry after a 401
    • journal restore with and without the secret variable, and that the secret never reaches the journal
  • Manual end-to-end test against the real CSCS FirecREST API on Alps (daint): dry-run, then
    alloc add → a Slurm job submitted via FirecREST → a worker on a compute node connected and ran the
    task → alloc remove --force cancelled the allocation via the API.

rjanalik and others added 5 commits October 6, 2026 11:26
Adds a new automatic allocation backend that submits Slurm allocations
through a FirecREST v2 API (https://eth-cscs.github.io/firecrest-v2/)
instead of executing sbatch locally. This allows the HyperQueue server
to run outside of the target cluster (e.g. on a VM), which is needed on
clusters with login-node wall-time limits, such as CSCS Alps.

- New ManagerType::Firecrest and FirecrestQueueParams (cluster-side
  paths for the hq binary, worker access file and allocation workdir;
  OAuth2 client credentials, with the secret read from an environment
  variable of the server and never persisted)
- New queue handler performing submit/status/cancel via HTTPS with a
  cached OAuth2 client-credentials token; reuses the Slurm submit
  script builder, and spawned workers use --manager slurm since
  FirecREST allocations are ordinary Slurm jobs
- reqwest (rustls) dependency gated behind a new `firecrest` cargo
  feature, enabled by default; servers built without it report a clear
  error for FirecREST queues

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…rnal restore

- The OAuth2 token cache is now guarded by an async mutex held across the
  fetch, so concurrent requests reuse one fetch instead of each fetching
  their own token. Any API call rejected with HTTP 401 (e.g. a token revoked
  before its expiration time) invalidates the cached token and is retried
  once with a freshly fetched one.
- Allocation status refresh now uses a single job-list request per queue
  instead of one request per allocation, keeping API (rate-limit) pressure
  independent of the number of active allocations. Allocations that already
  dropped out of the list are still queried individually as a fallback.
- Restoring the server from a journal no longer aborts the remaining state
  restoration when one allocation queue cannot be recreated (for example
  when the FirecREST client-secret environment variable is not set after a
  restart); the queue is skipped with an error in the log instead.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C8RVhZJp3g2FjfhN9sikWj
The firecrest backend talks to an HTTP API instead of scheduler binaries,
so the tests use an in-process mock of the token endpoint + FirecREST v2
API (tests/autoalloc/mock/firecrest.py) running on an ephemeral localhost
port, with request counters and fault injection.

Covered: submit script contents, dry-run success and token/submit errors,
allocation state lifecycle (queued/running/finished/failed), cancellation
on queue removal, batched status queries, transparent retry with a fresh
token on HTTP 401, journal restore with and without the secret environment
variable, and that the client secret is never persisted to the journal.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C8RVhZJp3g2FjfhN9sikWj
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C8RVhZJp3g2FjfhN9sikWj
- Check the fallback timestamp in finished_without_end_time_falls_back_to_now
- Pass the queue time limit explicitly in test_firecrest_submit_script
- Document that an off-cluster server needs a public address advertised to workers

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017QUYDeeb8vu1fxjgf4Tdkv
@Kobzol

Kobzol commented Oct 6, 2026

Copy link
Copy Markdown
Member

I'm no longer very involved in the maintenance of HyperQueue, but I think that the right way forward is to define a general API for autoalloc backends, and let HQ use it with an externally provided backend. Most likely it would make sense to define a JSON REST API, as performance is not super critical here, and that is as general as it gets.

That would allow anyone to implement their own autoalloc backend and use it with HyperQueue, rather than it being maintained in-tree. You could still implement it e.g. in Rust, and in this case it would essentially just do a translation between HQ's autoalloc REST API and the FirecREST API. It would require running the wrapper binary together with HyperQueue on some TCP/IP port, but that shouldn't be a big issue, it can just run on a single node togethe with the HQ server.

@spirali

spirali commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator

Disclaimer: I had an offline conversation about this with @rjanalik few weeks ago.

I did not have time to dig in the code yet, but as far as I understand, firecrest is basically a kind of API over SLURM. We do not need to invent our own.

I am ok to maintaining it directly in HQ, I see it as an alternative to two other backends that we already have.

Of course, if there would be like 3-5 other competing APIs in the future, than it is probably time to propose a unifying one for HQ. But in the current situation, having Firecrest directly in HQ makes sense to me.

This brings an out-of-the-box solution for people who does not run server on login node and wants to use autoallocator.

@Kobzol

Kobzol commented Oct 6, 2026

Copy link
Copy Markdown
Member

Fair enough :)

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants