Deployment and replica status lifecycle
This document explains the status values and state transitions for deployments and replicas in the Iris inference system.
Deployment status
Deployment status is stored in dedicated_deployments.status and tracks the overall lifecycle of a deployment.
| Status | Description | Transitions To | Triggered By |
|---|---|---|---|
created |
Active deployment. Replicas may be pending, provisioning, or running | paused, failed, stopped |
Gateway on /inference/create after submitting initial replicas |
paused |
Scaled to zero (no replicas) | created |
DeploymentScaler when idle timeout reached |
failed |
Permanent failure (all replicas failed fatally) | (terminal) | Gateway callback when last non-terminal replica fails with fatal error (image pull failure, container init failure) |
stopped |
User-terminated deployment | (terminal) | User via Gateway /inference/stop. All replicas aborted |
Source:
- Schema:
db/migrations/iris/000002_add_dedicated_deployments_tables.up.sql⧉ - Lifecycle:
app/src/app/api/v2_streaming/external/compute/inference/lifecycle/dedicated.py⧉
Notes:
failedis terminal. The user must create a new deployment with corrected configurationstoppedis terminal. The deployment cannot be restartedpausedis resumable. InferenceProxy triggers resume on the first inference request
Replica status and state machine
Replica status is stored in deployment_replicas.status and tracks the lifecycle of each individual replica (VM + docker-compose job).
Status values
| Status | Description | Terminal? |
|---|---|---|
pending |
Replica record created, job submitted to Streamer | No |
pulling_docker_image |
Execlet is downloading Docker images | No |
initing_docker_container |
Execlet is starting containers (before healthy) | No |
running |
Containers healthy and accepting requests | No |
failed |
Replica failed (see Failure Reasons below) | Yes |
stopped |
User-terminated or aborted by Gateway | Yes |
Source: shared_models/src/lyceum/shared_models/inference.py ⧉
State machine
STARTING STARTING RUNNING
"Pulling image" "Model server..." "healthy"
│ │ │
pending ─────────────►│ │ │
▼ │ │
pulling_docker_image ──────────┼──────────────────────┤
▼ │
initing_docker_container ────────────┘
│
▼
running
│
(crash, OOM, SYSTEM_FAILURE)
│
▼
failed
Any non-terminal state can also transition to:
- stopped (user calls /inference/stop; Gateway writes status then aborts)
- failed (see Failure Transitions table for specific conditions)
Triggers:
pending→pulling_docker_image: Execlet progress callback with message "Pulling image..."pulling_docker_image→initing_docker_container: Execlet progress callback with message "Model server is starting..."initing_docker_container→running: Execlet progress callback with "healthy" in message- Any →
stopped: User calls Gateway/inference/stop→ Gateway writesstatus='stopped'→ Gateway aborts job - Any →
failed: See Failure Transitions below
Failure transitions and fatality
Replicas can fail for different reasons, and some failures are fatal (indicate permanent misconfiguration) while others are non-fatal (transient infrastructure issues).
| From State | Execlet Status | Fail Reason | Fatal? | Description |
|---|---|---|---|---|
pulling_docker_image |
FAILED |
image_pull_failed |
Yes | Bad HuggingFace token, private model without access, or wrong image tag |
initing_docker_container |
FAILED |
container_init_failed |
Yes | Application crashes during startup, OOM during init, or bad configuration |
running |
SYSTEM_FAILURE |
container_crashed |
No | Container exited unexpectedly after becoming healthy (crash, OOM, etc.) |
| Any | VM_CRASH |
container_crashed |
No | Execlet process died (VM crash, network partition). Transient infrastructure failure |
Fatal Failures:
- If a replica fails with a fatal reason and no other replicas remain in non-terminal state, the deployment is automatically marked as
status='failed' - This prevents DeploymentScaler from spawning new replicas that will fail the same way
- User must create a new deployment with corrected configuration (valid HF token, accessible model, etc.)
Non-Fatal Failures:
- DeploymentScaler will eventually spawn replacement replicas when
current_replicas < min_replicas - These are treated as transient issues that auto-scaling can recover from
Current Replicas Count:
- Incremented by +1 when replica transitions to
running - Decremented by -1 when replica transitions from
running→failed(crashed) - Not decremented when replica fails before reaching
running(never counted)
Source:
- State machine:
app/src/app/api/v2_streaming/internal/callbacks/inference_dedicated.py⧉ - Transition logic:
app/src/app/api/v2_streaming/internal/callbacks/inference_dedicated.py⧉
Billing
A dedicated replica is billed for the time it is running, whether or not it
serves requests. Tokens sent through a dedicated deployment are not billed. The
price comes from the replica's hardware profile, dedicated.<gpu_type>.<n>x
(for example dedicated.h100.1x).
Running windows
Each stretch a dedicated replica spends in running is one row in
replica_running_windows. A trigger on deployment_replicas opens the row when
the replica enters running and sets ended_at when it leaves. A container
restart (running → restarting → running) closes one window and opens a
new one, and both are billed. Serverless replicas get no windows.
Billing clock
Once a minute, the gateway's billing clock reads every window that is not yet
billed to its end and sends Croesus one dedicated_inference_running event per
minute of it. Events line up with clock minutes, so the first event of a window
and the last event of a closed window can be shorter than a minute. An open
window is billed only up to the start of the current minute.
billed_until on each window is the watermark: everything before it has been
billed. The clock moves it forward after Croesus accepts an event, and the next
tick starts from it. If the gateway or Croesus is down, billed_until stays at
the last accepted minute, and the clock bills the missed time once both are
back. Each event id is the replica id plus the interval start, so Croesus drops
a resend of a minute it already has. A window sends at most 60 events per tick,
so after a long outage it catches up an hour of running time per tick.
Heartbeats
While a replica is running, its health reports refresh
deployment_replicas.last_health_check. A window is billed only up to the last
heartbeat plus 90 s:
- If the heartbeats stop while the replica stays
running, billing pauses. When they come back, the paused time is billed. - If the replica leaves
running, its window keeps the heartbeat it had at that moment, and the time after that heartbeat plus 90 s is never billed. This covers a node that dies before the replica's status changes.
For example, a replica that runs from 12:00, sends its last heartbeat at 12:30
and is marked failed at 13:30 is billed from 12:00 to 12:31:30. The clock then
sets billed_until to 13:30, the window's ended_at, so the window counts as
done and the time in between is not billed.
Switches and dry-run
DEDICATED_BILLING_ENABLED starts the clock and makes deployment creation
refuse a hardware profile that Croesus has no price for. It needs
CROESUS_HOST, and the gateway refuses to start without it.
DEDICATED_BILLING_DRY_RUN requires DEDICATED_BILLING_ENABLED and does nothing
without it. With both on, the clock logs the next event of each window instead
of sending it, and logs a suspended org instead of stopping its deployments. It
leaves billed_until where it is, as Hydra's dry-run does, so when dry-run is
turned off the clock bills the dry-run period as well.
Missing prices and suspended orgs
A window whose hardware profile has no price is skipped and keeps its
billed_until, because one unpriced reading fails the Croesus tally for every
product. Once the price row exists, the clock bills the window from where it
stopped. When Croesus reports an org as suspended, the clock stops that org's
running deployments through the normal stop path.
Source:
- Windows and trigger:
db/migrations/iris/000011_add_replica_running_windows.up.sql⧉ - Billing clock:
app/src/app/api/v2_streaming/external/compute/inference/core/replica_billing.py⧉