Skip to content

Block Storage

Block storage lets a customer create a persistent volume for VMs that outlives any single VM. The volume is created first, attached to a VM at boot, and survives that VM's termination so it can be attached to a new one. Like VMs, volumes are provider resources managed by Hydra and billed through Croesus.

Design proposal

Only the first bit of this feat is implemented. This page describes the intended design so that it can be reviewed before the rest of the code exists; it will become the explanation of the feat as the phases land.

Goals

  • Create one or more volumes, then start a VM with those volumes already attached.
  • Terminating the VM leaves the volumes intact and reattachable.
  • Volumes are billed for as long as they exist, attached or not.

Attaching storage to already running VMs

Attaching storage to an already-running VM is explicitly deferred.

NFS

The initial design uses only block storage which can be attached to a single VM at a time. Attaching shared storage in the form of NFS is explicitly deferred.

Provider Support

Only two of our providers offer block storage, the design must not assume every backend can do it.

Verda UpCloud Inferred restriction
SDK verdacloud-sdk-go, VolumeService upcloud-go-api/v8, storage service
Create parameters size, type, name, location_code size, tier, title, zone, encrypted Fix as many as possible (opinionated vs. flexible)
Types / tiers NVMe (HDD deprecated), NVMe_Shared for SFS maxiops, standard, hdd (Archive) NVMe/maxiops selected as default
Size limits per volume type 1–4096 GB per device 1 to 4096 GB
Devices per VM not documented 16 devices, 64 TB total 16 block volumes max
Attach at instance create yes, existing_volumes: []string yes, storage device with action: attach Attach only at create for now
Attach to running instance no, the instance must be shut down first hot-attach, avoid IDE addressing Not for now
Detach keyed on instance ID device address (e.g. virtio:1), not storage ID ID and address
Resize grow only grow only Offer no storage device resizing for now
Guest-side help returns Target, MountCommand, fstab command none, we derive the device path ourselves Not need because attached at startup only for now
Pricing /volume-types → price_per_month_per_gb per GB per hour, per tier Decide on common block storage price for these and future providers

Two constraints are common to both and shape everything else:

  • A volume lives in one location and can only attach to an instance in that same location.
  • A volume attaches to one instance at a time and must be detached before it can be attached elsewhere.

Note

UpCloud's detach API takes the device address rather than the storage UUID. We therefore have to persist the address we were given at attach time; it cannot be recomputed later.

Placement

Today no user ever chooses a location. CheapestAcrossPools() picks a pool, the backend picks the cheapest instance (in a location) inside it, and that choice is after creation discarded (md.Instance carries no location field).

A volume breaks that, because it permanently pins every future VM that wants to use it to one (backend, location) pair. Two consequences:

  1. Instance location has to become persisted state.
  2. Creating a VM with volumes cannot use CheapestAcrossPools(). The volumes decide the pool and the location; the VM request is pinned to them.

For the user-facing side we keep the "we pick the hardware" promise and ask for the intended workload rather than a region:

The user says "500 GB for my H100 work". Hydra has to be smart in resolving which backend and location actually has H100 capacity, create the volume there, and return the resolved provider and location on the response.

Data We Need

From the user, to create a volume

Field Required Notes
size_gb yes 1–4096, the tighter of the two providers' limits
display_name no user-editable label, same treatment as user_vms.display_name
placement no hardware_profile to place near, or an explicit provider / location

From the user, to create a VM with volumes

CreateVMRequest grows volume_ids: list[str]. Before calling Hydra the Gateway validates that every volume belongs to the active org, is currently detached, and that all of them share one (backend, location), and that the count fits the provider's device limits.

From the VM

Datum Where it comes from Why
provider instance ID, backend already on md.Instance attach/detach calls
location / zone missing today. the same-location constraint
device address attach response UpCloud detach is keyed on it

From ourselves

A is_formatted flag on the volume record. A brand-new volume needs mkfs; every later attach must not, or we destroy the customer's data. This is durable state we own, not something to infer in the guest at attach time.

Domain Model

Hydra gets a volumes_block table alongside instances, following the same shape and the same conventions.

id, name, pool_name, backend, org_id, user_id,
size_gb, volume_type, location, status,
attached_instance_id, device_address, device_path, is_formatted,
created_at, updated_at, billed_until

For Supabase, we will have to check which information would actually be needed and if everything (i.e. storage AND VMs) can go via Hydra instead.

Interfaces

Volume support is a separate optional interface rather than new methods on InstanceManagerBackend, because four of our six backends cannot implement it.

type VolumeManagerBackend interface {
    ListVolumes(ctx context.Context, poolID string) ([]md.Volume, error)
    CreateVolume(ctx context.Context, poolID string, spec md.VolumeSpec) (md.Volume, error)
    DeleteVolume(ctx context.Context, poolID, volumeID string) error
    AttachVolume(ctx context.Context, poolID, volumeID, instanceID string) (md.VolumeAttachment, error)
    DetachVolume(ctx context.Context, poolID, volumeID, instanceID string) error
    // ResizeVolume(ctx context.Context, poolID, volumeID string, newSizeGB uint32) error
    VolumeLocations(ctx context.Context, poolID string) ([]string, error)
    VolumePricePerGBHour(ctx context.Context, poolID, volumeType, location string) (float64, bool, error)
}

CreateInstance needs two more inputs: the pinned location and the volume IDs to attach. It already takes five positional parameters, so it should collapse into a CreateInstanceParams struct rather than grow to seven.

Hydra's HTTP API mirrors the instance handlers: /api/v1/volumes and /api/v1/pools/{poolID}/volumes. The Gateway exposes POST|GET /volumes, GET|PATCH|DELETE /volumes/{id}, and later on POST /vms/{vm_id}/volumes/{volume_id}:attach and :detach.

Storage in availability

GPU counts of one hardware profile can sit in different datacentres, and only some datacentres hold block storage. Availability therefore carries which counts can be started with a volume:

  • Availability says where each GPU count can be started. See md.AvailableGpu.LocationsByInstanceType. Every location names its provider, so this says whose capacity it is as well as where it is.
  • Hydra reports block_storage_counts_by_instance_type per GPU type in GET /api/v1/instances/real-availability. A count is in it when some location with capacity for it also holds block storage. It is always a subset of counts_by_instance_type.

A backend that cannot name its locations reports no storage-capable counts. Nothing is promised on capacity whose place is unknown.

Billing

A new meter in croesus/config.yaml ⧉:

meters:
    -   slug: block_storage
        eventType: block_storage
        valueProperty: $.gb_seconds
        groupBy:
            volume_id: $.volume_id
            storage_tier: $.storage_tier
        aggregation: SUM

meter_billing_policy:
    block_storage:
        pricing_keys: [ storage_tier ]
        transaction_keys: [ volume_id ]

Hydra's Biller generalises to volumes almost unchanged: the same billed_until watermark, and the same deterministic event ID (<volume-id>:<interval-start-unix>) as for instances so that a replayed interval is recorded once.

The relevant difference is that a VM is only billable while it is ready, gated on ReadySince. A volume is billable from creation to deletion regardless of whether anything is attached to it; the billing gate is simply "exists and has an owner".

A forgotten volume bills forever

This is the main product risk of the feature: A 4 TB volume nobody has attached in months still costs money, and unlike a VM, there is no running workload making that visible. Before launch, we need the running costs surfaced in a Grafana dashboard , and a policy for what happens when an org's credits run out. A grace period followed by deletion is a realistic option.

Reconciliation

The periodic reconciler learns to list provider volumes carrying our hyd name prefix and to reconcile status and attachment, exactly as it does for instances.

Both directions of drift matter, and asymmetrically:

  • A volume in our database that no longer exists at the provider should be marked deleted (stop billing).
  • A volume at the provider with our prefix that is not in our database must be a loud warning because it is a cash leak.

Steps

0. Unblock

Neither item is user-visible, and everything else depends on both.

  • Do not delete customer storage on VM teardown. internal/upcloud/client.go terminated with DeleteServerAndStorages(Backups: delete), which destroys every attached storage. It should now detach all non-OS disks first and only then delete the server with its OS disk, identified by the <hostname>-os title CreateServer gives it. A failed detach aborts the teardown rather than proceeding to destroy the volume it was protecting.
  • Persist instance location. best.zone and best.Locations[0] were chosen inside CreateInstance and then thrown away. md.Instance should now carry a Location, stored in a new instances.location column and reported on the API. Verda and UpCloud should populate it on both the create and the list path; the other backends should leave it empty, which is enough because they cannot offer block storage anyway.

1. MVP

Volume CRUD end to end (Hydra backends, repository, HTTP API, Gateway endpoints, Supabase table, Croesus meter), and volume_ids on VM creation using attach-at-create on both providers. Volumes survive termination and can be attached to a later VM.

2. Block Storage Migration

Avoid exposing the confusing location choice by simply offering migrating storage to the other location.

Storage in FIN01 -->|create instance| Instance in FIN02 -->|migration allowed| Storage migrated to FIN02

3. Live attachment

Attach and detach on running VMs, resize, and format-on-first-attach automation driven by is_formatted. UpCloud supports hot-attach directly. Verda does not: the instance has to be shut down first, which their Go SDK (VolumeService.AttachVolume), their Python SDK ("if the instance isn't shutdown an exception would be raised") and their docs ("shutting down temporarily pauses it so technical processes can occur, such as attaching or detaching volumes") all say.

4. Shared filesystems/NFS

NFS is a different resource, not a flag on a volume: Verda SFS (NVMe_Shared) is multi-attach, network-mounted and has its own endpoint rather than a block device. It belongs in a sibling shared_filesystems resource, so that the single-attach invariant that keeps block volumes safe is not weakened to accommodate it. Also investigate UpCloud to make it behave the same way.

Open Questions

  • Do users pick a region explicitly, or is it inferred from the intended hardware_profile? This page assumes inferred with an override.
  • What happens to an idle volume when an org runs out of credits, and after how long?
  • Should there be a per-org quota on total provisioned GB?