Skip to content

Croesus

Croesus is our billing service. Every component that consumes a billable resource reports what happened. Croesus stores those reports, prices them and writes the result to a per-subject ledger. It also tops accounts up automatically via Stripe.

Code /croesus/ ⧉
Language Go
Deployment Cloud Run, internal ingress only, croesus.<dns-zone>
Database CloudSQL, LYC_DB_CROESUS_*

Why it is separate

Billable usage originates in several places: the Gateway (serverless inference, executions), Hydra (VM runtime), InferenceWorkers (tokens), and their variants. If each producer priced its own usage, then pricing logic would scatter across codebases. There would be many chances to double-charge, and no single place to audit a bill.

Producers therefore report what happened, never what it costs. Croesus owns pricing, aggregation and the ledger.

Data model

flowchart LR
    E[Events] -->|meter query| R[Meter readings]
    P[Prices] --> T
    R -->|tally| T[Transactions]
    T -->|sum| B[Account balance]
  • Event. An immutable record that something billable happened, e.g. "VM x ran for 300 seconds on hardware profile gpu.h100". Identified by the composite key (source, id), because IDs are generated by producers and are not unique across them. Recording an event Croesus already has keeps the stored one and answers success. A producer that derives its IDs can therefore retry an event it is unsure about without risking a double charge. The first write wins. An ID must therefore name one billable event, and must never be reused for another event from the same source.
  • Meter. A query over events. It selects one event type, extracts a value from the event data via JSONPath, groups events into bins and aggregates them. Meter readings are never persisted: they are a deterministic function of the events, which we keep anyway.
  • Price. A unit price per meter and group-by combination.
  • Transaction. A credit or debit against a subject's account, with an idempotency key.
  • Account balance. The sum of all credits minus the sum of all debits. There is no separate balance column to drift out of sync.

Amounts are always USD, held as decimals. Never cast a balance to a float.

Ingesting events

Producers POST a CloudEvent ⧉ to /events/v1/record. Croesus accepts and stores any well-formed event, whether or not a meter exists for its type. Unmetered events are never billed.

The CloudEvent fields map onto the internal event as follows:

CloudEvent Croesus Meaning
source + id Key Composite identity of the event
subject SubjectID Who pays (user or org)
type Kind Which meter, if any, processes this event
time OccurredAt When it happened, per the producer
data Data Arbitrary JSON the meter reads via JSONPath

Go producers should use the croesus/event ⧉ client rather than building requests by hand. Python producers post directly. See app/src/app/api/v2_streaming/external/vms/billing_service.py for the established shape.

Warning

Croesus performs no authentication. Its only access control is the Cloud Run internal-ingress boundary. Anything inside the VPC can record events for any subject and credit any account.

Meters and billing policy

Both live in croesus/config.yaml ⧉ and are loaded at startup. The config is validated on load. A meter referring to a non-existent group-by key aborts the process, and so does a billing policy naming an unknown meter.

A meter is modelled after a simplified subset of the OpenMeter ⧉ API, so that we can move to OpenMeter later without breaking the config schema. We do not use OpenMeter today because it drags in ClickHouse.

meters:
    -   slug: vm_running          # unique, immutable
        eventType: vm_running     # which events this meter consumes
        valueProperty: $.duration_seconds
        groupBy:
            vm_id: $.vm_id
            hardware_profile: $.hardware_profile
        aggregation: SUM

The billing policy splits groupBy into two independent roles. Conflating them is the usual source of confusion:

  • pricing_keys. The subset used to look a unit price up. For vm_running this is hardware_profile: the GPU type determines the price per second.
  • transaction_keys. The subset used to split transactions. For vm_running this is vm_id. Each VM gets its own line so a bill can be explained, even though all VMs of a profile share one price.

An empty transaction_keys means all readings for a subject collapse into one transaction.

Dedicated inference is metered the same way as VMs: dedicated_inference_running counts the seconds a replica is Ready, priced by hardware_profile and split into one transaction per deployment_id. Its profiles carry a dedicated. prefix (dedicated.h100.1x) so the customer price never depends on how the node was procured and never collides with a vm_running row.

Tallying

Tallying turns readings into debits. Cloud Scheduler POSTs to /api/v1/tally every five minutes. Nothing runs it in-process.

For each pass:

  • Read the tally cursor. That is the arrival timestamp everything before which has been billed.
  • Walk forward from the cursor in tally_interval steps (default five minutes) up to now. A period that would extend past now is left for the next pass.
  • For every period, meter and subject: fetch readings, look up unit prices, group readings by transaction_keys, and create one debit per group.
  • Advance the cursor.

Each transaction carries the idempotency key <meter-slug>:<subject-id>:<group-label>:<from>:<to>, so a retried or overlapping pass cannot double-charge. This is why the Cloud Scheduler job retries freely. A failed tally resumes where it stopped, and a later scheduled pass fixes what the retries did not.

Reading groups that price to zero or less are dropped with a warning rather than failing the pass. A zero-duration event should not stop billing for everyone else.

Note

A missing price is not tolerated. If a meter produces a reading whose pricing_keys combination has no row in prices, the whole tally pass fails and the cursor does not advance. Adding a new hardware profile or serverless model therefore means adding its price first.

Auto top-up

A subject can save a card and a threshold. When the balance falls below it, Croesus charges the card and credits the proceeds. Unlike tallying, this runs in-process on a ticker (auto_top_up_interval, default ten seconds) with a worker pool (auto_top_up_workers, default eight).

Guards against charging a card repeatedly:

Setting Default Purpose
auto_top_up_min_cooldown 60s Minimum spacing between two successful charges
auto_top_up_failure_cooldown 5m Backoff after a declined charge. Validated to be at least the min cooldown
auto_top_up_charge_timeout 30s After this a charge is reconciled, not retried blindly
auto_top_up_max_failed_attempts 3 Consecutive declines before auto top-up is disabled

A charge is claimed in the database before it is attempted. A unique index enforces one open charge per subject, so two workers cannot charge the same card concurrently.

When auto top-up is disabled, Croesus publishes to the croesus.events RabbitMQ exchange with routing key auto_top_up.disabled. The Gateway consumes it and emails the customer. If LYC_RABBITMQ_HOST is unset, the publisher is a no-op and Croesus runs normally without notifications.

Charging is invoice-first through Stripe. LegacyProcessor exists only to reconcile charges opened before that switch and is scheduled for deletion.

API

All routes are under /api/v1 except event recording and metrics.

Route Method Purpose
/events/v1/record POST Record a CloudEvent
/metrics GET Prometheus metrics
/health GET Health check
/prices GET List unit prices
/subjects/{id}/info GET Balance and suspension status
/subjects/{id}/events GET Raw events, optionally aggregated via /aggregate
/subjects/{id}/transactions GET Ledger entries, optionally aggregated via /aggregate
/subjects/{id}/credit POST Credit an account
/subjects/{id}/debit POST Debit an account
/subjects/{id}/reset-balance POST Set the balance to a given amount
/subjects/{id}/billing-profile GET, PUT Billing profile
/subjects/{id}/billing-details GET, PUT Billing details (PII, encrypted at rest)
/subjects/{id}/auto-top-up-settings GET, PUT Threshold and top-up amount
/subjects/{id}/setup-intent POST Stripe client secret for saving a card
/tally POST Run a tally pass

Accounts are created lazily: every /subjects/{id}/… route creates the account if it does not exist, as does the first event for a subject. Account creation also provisions a Stripe customer, so that top-ups always target the same one.

Configuration

Connection details come from LYC_* environment variables. Secrets are read from files under --secret-dir (default /secrets).

Variable / file Purpose
LYC_DB_CROESUS_{HOST,PORT,NAME,USER} CloudSQL connection
lyc-db-croesus-password/value Database password
LYC_ENCRYPTION_KEY_URI KMS key used to encrypt billing-detail PII
lyc-stripe-secret-key/value Stripe API key
LYC_RABBITMQ_*, lyc-rabbitmq-password/value Event publishing. Optional

Known rough edges

  • EnsureAccounts() is called on nearly every path to create accounts on demand. It is marked temporary by a TODO in croesus/cmd/croesus/main.go and should disappear once accounts are provisioned explicitly.
  • The config is read once at startup. Changing a meter or price dimension needs a redeploy.
  • /subjects/{id}/setup-intent leaks a Stripe client secret to its caller. This goes away when card handling moves fully into Croesus.