supported_gpus.yml
The supported_gpus.yml inventory file decouples the GPU configuration from the actual code implementation of Hydra. It is the single source of truth for three things:
- Which GPU hardware profiles we offer.
- How we address them internally.
- How each compute provider names them.
The file lives once per environment in the inventory directories: infra/inventory/<env>/supported_gpus.yml. Hydra reads the file from its active inventory directory at startup. If the file cannot be read or parsed, the load and the whole service fails.
Top-level structure
The file contains a single key, supported_gpus, holding a list of GPU profiles. Each entry's type becomes the key by which the profile is looked up, so type values must be unique.
supported_gpus:
- type: a100
labeling:
nvidia_smi: A100
ansible: gpu.a100
instance_counts:
max: 7
allowed_gpus_per_instance: [1, 2, 4, 8]
hardware_specs:
memory_gib:
max: 120
providers:
- name: prime-intellect
api_names:
- A100_80GB
instance_types:
- on-demand
- spot
- name: verda
api_names:
- "A100 80GB"
Fields
GPU profile
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
type |
string | yes | — | Internal identifier and lookup key for the hardware profile, e.g. a100, h100.185gb. This is how we address the GPU everywhere in our code. |
labeling |
object | yes | — | How the GPU is detected and labelled on a machine. See Labeling. |
instance_counts |
object | no | unlimited | Limits on how many instances of this type may run as an execlet (no impact on bare VMs). See Instance counts. |
inference_worker_instance_counts |
object | no | not offered | Node group bounds for inference workers rented by the Kubernetes node autoscaler. Omit to keep the type out of the autoscaler. See Instance counts. |
inference_worker_resources_per_gpu |
object | no | 8 vCPU, 30 GiB per node | vCPUs, memory and disk an inference-worker node provides per GPU. See Inference worker resources. |
allowed_gpus_per_instance |
list of ints | no | [1] |
The GPU counts a single instance may carry, e.g. [1, 2, 4, 8]. A request for any other count is rejected. |
hardware_specs |
object | no | no constraints | Optional vCPU and memory constraints for the underlying machine. See Hardware specs. |
providers |
list | yes | — | The compute providers that offer this GPU and how they name it. See Providers. |
Labeling
Labelling is mainly required for running execlets on this GPU type.
| Field | Type | Required | Description |
|---|---|---|---|
nvidia_smi |
string | yes | Pattern matched against the nvidia-smi output to detect this GPU on a machine. |
ansible |
string | yes | Ansible label applied to the host, e.g. gpu.a100. |
Instance counts
| Field | Type | Required | Description |
|---|---|---|---|
min |
int | no | Minimum number of instances the node autoscaler keeps for this GPU type. Ignored under instance_counts. |
max |
int | no | Maximum number of instances of this GPU type that may exist at once, counted separately for execlets (instance_counts) and inference workers (inference_worker_instance_counts). Neither applies to bare VMs. |
Hardware specs
hardware_specs is optional. When omitted, no vCPU or memory constraints are applied. It may contain vcpus and/or memory_gib, each a constraint object:
| Field | Type | Required | Description |
|---|---|---|---|
vcpus |
constraint | no | Constraint on the number of vCPUs of the machine. |
memory_gib |
constraint | no | Constraint on the machine's RAM in GiB. |
Each constraint object accepts min and/or max (non-negative integers):
hardware_specs:
memory_gib:
min: 185
max: 256
This is what distinguishes profiles such as h100 and h100.185gb, which share the same physical GPU but differ in the required machine memory.
Inference worker resources
inference_worker_resources_per_gpu sets the node shape the Kubernetes node autoscaler plans with. cluster-autoscaler checks pending pods against a template node per node group before it rents one; the template gets vcpus and memory_gib times the group's GPU count, and Hydra then rents only machines with at least that much. Without the key, the template node has 8 vCPUs and 30 GiB, so a pod that requests more CPU or memory never makes the autoscaler rent that type. disk_gb times the GPU count sets the smallest disk Hydra rents, since replicas keep their model weights on the node's disk; without it Hydra asks for 500 GB. Atlas books by GPU type and count only and ignores all three.
| Field | Type | Required | Description |
|---|---|---|---|
vcpus |
int | yes | vCPUs per GPU. |
memory_gib |
int | yes | Machine RAM in GiB per GPU. |
disk_gb |
int | no | Machine disk in GB per GPU. |
inference_worker_resources_per_gpu:
vcpus: 16
memory_gib: 128
disk_gb: 250
Providers
providers lists every compute provider that can supply the GPU. Each provider entry maps our internal type to the names the provider uses.
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
name |
string | yes | — | Provider identifier, e.g. prime-intellect, verda, atlas. |
api_names |
list of strings | yes | — | The names the provider uses for this GPU. Multiple entries are allowed when a provider exposes several SKUs that all map to our type (e.g. H200_96GB and H200_141GB). |
instance_types |
list | no | [on-demand] |
Pricing/availability variants we allow for this provider/GPU combination. |
Valid instance_types values are:
on-demand: a regular pay-as-you-go instance the provider will not preempt.spot: a preemptible/spot instance the provider may terminate at any time.
Unknown values cause the file to fail parsing.
Example
To create an l40s on our end, Hydra tells Prime Intellect to create an L40S_48GB or Verda to create an L40S. The api_names entries are exactly those provider-side names.