Skip to content

About

A GPU node pool as its own Cluster API release: MachinePool, KubeadmConfig and KarpenterMachinePool with their own worker bootstrap (Agent Platform, bumblebee-plans#46)

Resources

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

CircleCI OpenSSF Scorecard

gpu-node-pool

A GPU node pool on an existing Giant Swarm cluster is today a nodePools entry in the cluster's values: a pull request against the cluster's definition, re-rendered by the cluster's release on every upgrade, and out of reach for the Agent Platform — App CRs are deprecated, no one may edit a cluster's App CR or HelmRelease, and the platform's Models pages have nothing to add GPU capacity with. Depending on the shared cluster chart or copying the cluster's own worker pool both bind the pool to one cluster release: the first renders the release's containerd Secret under the same name and hook Jobs that act on the cluster's own objects, the second has to be re-derived on every cluster upgrade and never gets a lifecycle of its own.

This chart makes a GPU node pool its own Cluster API release with its own bootstrap and lifecycle: one Flux HelmRelease per pool in org-<org>, rendering for an existing Cluster API cluster named in its values a MachinePool, a hash-named KubeadmConfig carrying the chart's own worker bootstrap (the rendered CAPA Flatcar Karpenter worker spec with its files inline, no chart-owned Secrets), and a KarpenterMachinePool in the GPU shape (instance families from a curated accelerator list, the nvidia.com/gpu taint, scale to zero). The pool pins its own Kubernetes version and machine image — never newer than the control plane — so a cluster upgrade never touches it, and it is deleted with the cluster through Cluster API's owner references. Consumers are cluster-manager (create_node_pool, delete_node_pool) and the Dev Portal's Add GPU node pool dialog. A separate, cluster-owned pool means creating and removing GPU capacity never touches the cluster's other pools.

Status

The repository carries the chart's skeleton and CI. The templates, the values contract and the bootstrap oracle (make verify against the newest cluster-aws) follow in giantswarm/giantswarm#37713.

Installing

The chart is released to the Giant Swarm catalog (oci://gsoci.azurecr.io/charts/giantswarm/gpu-node-pool) and installed as a Flux HelmRelease with an exactly pinned chart version — a bootstrap change rolls GPU nodes, so bumps are explicit.

About

A GPU node pool as its own Cluster API release: MachinePool, KubeadmConfig and KarpenterMachinePool with their own worker bootstrap (Agent Platform, bumblebee-plans#46)

Resources

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages