Skip to content

Hydra -- autoscaler

For spinning up computing instances with different providers, new instances go through the following statuses, which involves going through our own provisioning. As the newly spun-up instances live at several sites, we put some emphasis on having a secure connection between our servers.

Status

Inside Hydra, instances created with the actual compute provider through a backend (e.g. Prime Intellect) have different statuses:

  • pending: Instance is created but not yet set up by the provider
  • running: Provider's instance setup is ready and we can access the instance
  • waiting_for_storage: Instance is up, but its provisioning is held back until its block storage has reached it. The provider reports such an instance as running.
  • stopped_for_storage: Instance is shut down so that a volume can be attached to it. It keeps its hardware, and the provider still bills it. Hydra does not bill the customer for it.
  • provisioning: Our provisioning is running on the instance (usually through Ansible)
  • ready: Instance is ready for customer usage
  • terminated: Instance is terminated, also on provider's side
  • error: Could be reported either by the provider (very unlikely) or Hydra itself, having issues with this instance, most likely a failed provisioning attempt

Our provisioning of new machines

After a provider is declared as running by the provider, Hydra rans an own provisioning flow in form of an Ansible playbook. The used playbook varies based on whether Hydra needs to prepare an execlet or a bare VM.

When a playbook execution does not succeed, we attempt rerunning it based on on our config. When a run succeeded, we persist these information to Hydra's own database so that these information survive Hydra service restarts. This means that restarting Hydra will not automatically rerun the Ansible provisioning if it already succeeded.

Instance managers

Each provider gets one or more pools, usually one for bare VMs and one for execlets. As this means an InstanceManagerBackend can have several pools, each instance is assigned to a pool based on their name prefix.

Example pools

Examples pools are verda-vm-pool, prime-intellect-vm-pool and prime-intellect-execlet-pool.

On a slightly higher level than the backends, we have InstManagers, which are able to direct actions to specific pools. For each pool, we create one SinglePoolInstManager, which inherently is able tell its backend to perform actions for its pool. MultiPoolInstManager lets Hydra interact with a set of SinglePoolInstManagers. It also implements the InstManager interface, and routes actions for a pool to the relevant SinglePoolInstManager.

uml diagram

Providers

Online services supplying GPU compute capacity are integrated as providers. Each provider interacts outside of Hydra through a custom client. To abstract the client's functionalities, each provider gets an InstanceManagerBackend and Hydra knows every provider inform of this backend.

uml diagram

Tip

Create an unnamed variable for letting the compiler check that a struct implements an interface.

var _ as.InstanceManagerBackend = (*InstanceManagerBackend)(nil)

Supported providers

The following already integrated providers have each some relevant unique details. Hydra's interaction with these providers happens via the account <provider>@lyceum.technology. The password is in the SOPS inventory files. You can also use your own account to interact with the compute instances. Ask to be added to each team.

Whenever we want to spin up an instance with any of our providers, Hydra looks through the different providers and starts the cheapest option.

Atlas

Atlas is our internal provider and its instances will always be preferred because of an artifially cheap price.

Mithrill

  • Quite modern provider which is not very stable, offers only US compute instances in spot and reserved modes (no on-demand)
  • One project with own SSH key per env. However, API secret is able to access all projects. Acceptable because we also specify to which project to connect to.
  • Interaction through REST API

Currently disabled

The providers code is already in production but disabled. It needs more testing and adjustment to the providers model. See the related Linear ticket ⧉ to see more details about the paused integration.

Mithril ⧉ is used as a provider for spot instances only (no on-demand). The provider itself functions through a Vickrey auction, i.e. a sealed 2nd price auction. This means that we have to bid to then potentially receive an instance. Bid and instance would therefore have to be considered together to handle the instance further. We only create bids for single instances. Whenever an instance is preempted we cancel the bid automatically, because we do not want another instance. Our current sales model demands this.

We determine the price we want to bid by first querying the spot instance prices which we would have to pay with other providers. Afterwards, we bid a cent less with Mithril. However, we currently only place a Mithril bid if we think we can immediately win the auction to get an instance. To decide this we look at the highest price that recently won a bid. We only place a bid if our price is a certain percentage higher.

The current problem seems to be that our availability prediction is not good enough. We place a bid, it does not win, and the customer waits a long time only to be told it did not work.

Prime Intellect

  • app.primeintellect.ai ⧉, quite flaky provider wrapping other providers (also Verda)
  • Offers access to providers at many different global location (also a lot in the US)
  • Uses same way of accessing for every env (also same SSH key)
  • Interaction happens through REST API ⧉ (might changes fast as company seems to be young)

UpCloud

  • upcloud.com ⧉, European provider offering fixed GPU server plans (H100, B200, L40S, L4)
  • Servers are created from fixed plans (e.g. GPU-12xCPU-240GB-1xH100) instead of arbitrary instance types. The api_names in supported_gpus.yml match the plans' gpu_model attribute (e.g. NVIDIA H100 80GB HBM3, NVIDIA B200)
  • Spot capacity exists as separate GPU-SPOT- plans and the backend supports it, but we currently only enable on-demand in supported_gpus.yml
  • Authentication via an API subaccount (username + password). The OS disk is cloned from UpCloud's AI/ML-ready GPU Ubuntu template which ships with NVIDIA drivers preinstalled
  • Interaction with provider happens through UpCloud's own Go SDK ⧉

Verda

  • verda.com ⧉, very reliable provider in Finland offering currently EU-only hardware
  • Different cloud credentials and SSH keys used for the different environments
  • Each environment could easily modify instances in other environments
  • Interaction with provider happens through Verda's own Go SDK ⧉
  • Offers block storage. Verda attaches a volume only at instance creation or to a stopped instance, so Hydra can stop and start Verda instances.

Billing metrics

Hydra exposes these Prometheus metrics for the events it sends to Croesus:

Metric Type Meaning
hydra_billing_events_recorded_total counter VM runtime intervals Croesus accepted, repeats included
hydra_billing_events_failed_total counter VM runtime intervals Croesus refused, retried later
hydra_billing_lag_seconds gauge age of the oldest unbilled VM runtime
hydra_billing_volume_events_recorded_total counter block storage intervals Croesus accepted, repeats included
hydra_billing_volume_events_failed_total counter block storage intervals Croesus refused, retried later
hydra_billing_volume_lag_seconds gauge age of the oldest unbilled block storage

Supported GPUs

To decouple the configuration and the actual code implementation of Hydra, we use an inventory file called supported_gpus.yml. There we define the type, which is how we address this GPU hardware profile internally. We also define how many GPUs of this type we allow per instance.

Relevant for the providers is how they address these hardware profiles on their end. And this is what is listed under api_names of each provider, how a provider calls the hardware profile which we call as mentioned under type.

Example

To create an l40s on our end, we tell Prime Intellect to create an L40S_48GB. For Atlas we ask for an NVIDIA L40S.

See the supported_gpus.yml reference for the full schema of this file.