Skip to content

Deployment and replica status lifecycle

This document explains the status values and state transitions for deployments and replicas in the Iris inference system.


Deployment status

Deployment status is stored in dedicated_deployments.status and tracks the overall lifecycle of a deployment.

Status Description Transitions To Triggered By
created Active deployment. Replicas may be pending, provisioning, or running paused, failed, stopped Gateway on /inference/create after submitting initial replicas
paused Scaled to zero (no replicas) created DeploymentScaler when idle timeout reached
failed Permanent failure (all replicas failed fatally) (terminal) Gateway callback when last non-terminal replica fails with fatal error (image pull failure, container init failure)
stopped User-terminated deployment (terminal) User via Gateway /inference/stop. All replicas aborted

Source:

Notes:

  • failed is terminal. The user must create a new deployment with corrected configuration
  • stopped is terminal. The deployment cannot be restarted
  • paused is resumable. InferenceProxy triggers resume on the first inference request

Replica status and state machine

Replica status is stored in deployment_replicas.status and tracks the lifecycle of each individual replica (VM + docker-compose job).

Status values

Status Description Terminal?
pending Replica record created, job submitted to Streamer No
pulling_docker_image Execlet is downloading Docker images No
initing_docker_container Execlet is starting containers (before healthy) No
running Containers healthy and accepting requests No
failed Replica failed (see Failure Reasons below) Yes
stopped User-terminated or aborted by Gateway Yes

Source: shared_models/src/lyceum/shared_models/inference.py ⧉

State machine

                    STARTING              STARTING               RUNNING
                  "Pulling image"    "Model server..."         "healthy"
                        │                   │                      │
  pending ─────────────►│                   │                      │
                        ▼                   │                      │
             pulling_docker_image ──────────┼──────────────────────┤
                                            ▼                      │
                              initing_docker_container ────────────┘
                                                                   │
                                                                   ▼
                                                                running
                                                                   │
                                                      (crash, OOM, SYSTEM_FAILURE)
                                                                   │
                                                                   ▼
                                                                failed


  Any non-terminal state can also transition to:
    - stopped  (user calls /inference/stop; Gateway writes status then aborts)
    - failed   (see Failure Transitions table for specific conditions)

Triggers:

  • pending → pulling_docker_image: Execlet progress callback with message "Pulling image..."
  • pulling_docker_image → initing_docker_container: Execlet progress callback with message "Model server is starting..."
  • initing_docker_container → running: Execlet progress callback with "healthy" in message
  • Any → stopped: User calls Gateway /inference/stop → Gateway writes status='stopped' → Gateway aborts job
  • Any → failed: See Failure Transitions below

Failure transitions and fatality

Replicas can fail for different reasons, and some failures are fatal (indicate permanent misconfiguration) while others are non-fatal (transient infrastructure issues).

From State Execlet Status Fail Reason Fatal? Description
pulling_docker_image FAILED image_pull_failed Yes Bad HuggingFace token, private model without access, or wrong image tag
initing_docker_container FAILED container_init_failed Yes Application crashes during startup, OOM during init, or bad configuration
running SYSTEM_FAILURE container_crashed No Container exited unexpectedly after becoming healthy (crash, OOM, etc.)
Any VM_CRASH container_crashed No Execlet process died (VM crash, network partition). Transient infrastructure failure

Fatal Failures:

  • If a replica fails with a fatal reason and no other replicas remain in non-terminal state, the deployment is automatically marked as status='failed'
  • This prevents DeploymentScaler from spawning new replicas that will fail the same way
  • User must create a new deployment with corrected configuration (valid HF token, accessible model, etc.)

Non-Fatal Failures:

  • DeploymentScaler will eventually spawn replacement replicas when current_replicas < min_replicas
  • These are treated as transient issues that auto-scaling can recover from

Current Replicas Count:

  • Incremented by +1 when replica transitions to running
  • Decremented by -1 when replica transitions from running → failed (crashed)
  • Not decremented when replica fails before reaching running (never counted)

Source:


Billing

A dedicated replica is billed for the time it is running, whether or not it serves requests. Tokens sent through a dedicated deployment are not billed. The price comes from the replica's hardware profile, dedicated.<gpu_type>.<n>x (for example dedicated.h100.1x).

Running windows

Each stretch a dedicated replica spends in running is one row in replica_running_windows. A trigger on deployment_replicas opens the row when the replica enters running and sets ended_at when it leaves. A container restart (running → restarting → running) closes one window and opens a new one, and both are billed. Serverless replicas get no windows.

Billing clock

Once a minute, the gateway's billing clock reads every window that is not yet billed to its end and sends Croesus one dedicated_inference_running event per minute of it. Events line up with clock minutes, so the first event of a window and the last event of a closed window can be shorter than a minute. An open window is billed only up to the start of the current minute.

billed_until on each window is the watermark: everything before it has been billed. The clock moves it forward after Croesus accepts an event, and the next tick starts from it. If the gateway or Croesus is down, billed_until stays at the last accepted minute, and the clock bills the missed time once both are back. Each event id is the replica id plus the interval start, so Croesus drops a resend of a minute it already has. A window sends at most 60 events per tick, so after a long outage it catches up an hour of running time per tick.

Heartbeats

While a replica is running, its health reports refresh deployment_replicas.last_health_check. A window is billed only up to the last heartbeat plus 90 s:

  • If the heartbeats stop while the replica stays running, billing pauses. When they come back, the paused time is billed.
  • If the replica leaves running, its window keeps the heartbeat it had at that moment, and the time after that heartbeat plus 90 s is never billed. This covers a node that dies before the replica's status changes.

For example, a replica that runs from 12:00, sends its last heartbeat at 12:30 and is marked failed at 13:30 is billed from 12:00 to 12:31:30. The clock then sets billed_until to 13:30, the window's ended_at, so the window counts as done and the time in between is not billed.

Switches and dry-run

DEDICATED_BILLING_ENABLED starts the clock and makes deployment creation refuse a hardware profile that Croesus has no price for. It needs CROESUS_HOST, and the gateway refuses to start without it. DEDICATED_BILLING_DRY_RUN requires DEDICATED_BILLING_ENABLED and does nothing without it. With both on, the clock logs the next event of each window instead of sending it, and logs a suspended org instead of stopping its deployments. It leaves billed_until where it is, as Hydra's dry-run does, so when dry-run is turned off the clock bills the dry-run period as well.

Missing prices and suspended orgs

A window whose hardware profile has no price is skipped and keeps its billed_until, because one unpriced reading fails the Croesus tally for every product. Once the price row exists, the clock bills the window from where it stopped. When Croesus reports an org as suspended, the clock stops that org's running deployments through the normal stop path.

Source: