Croesus
Croesus is our billing service. Every component that consumes a billable resource reports what happened. Croesus stores those reports, prices them and writes the result to a per-subject ledger. It also tops accounts up automatically via Stripe.
| Code | /croesus/ ⧉ |
| Language | Go |
| Deployment | Cloud Run, internal ingress only, croesus.<dns-zone> |
| Database | CloudSQL, LYC_DB_CROESUS_* |
Why it is separate
Billable usage originates in several places: the Gateway (serverless inference, executions), Hydra (VM runtime), InferenceWorkers (tokens), and their variants. If each producer priced its own usage, then pricing logic would scatter across codebases. There would be many chances to double-charge, and no single place to audit a bill.
Producers therefore report what happened, never what it costs. Croesus owns pricing, aggregation and the ledger.
Data model
flowchart LR
E[Events] -->|meter query| R[Meter readings]
P[Prices] --> T
R -->|tally| T[Transactions]
T -->|sum| B[Account balance]
- Event. An immutable record that something billable happened, e.g. "VM
xran for 300 seconds on hardware profilegpu.h100". Identified by the composite key(source, id), because IDs are generated by producers and are not unique across them. Recording an event Croesus already has keeps the stored one and answers success. A producer that derives its IDs can therefore retry an event it is unsure about without risking a double charge. The first write wins. An ID must therefore name one billable event, and must never be reused for another event from the same source. - Meter. A query over events. It selects one event type, extracts a value from the event data via JSONPath, groups events into bins and aggregates them. Meter readings are never persisted: they are a deterministic function of the events, which we keep anyway.
- Price. A unit price per meter and group-by combination.
- Transaction. A credit or debit against a subject's account, with an idempotency key.
- Account balance. The sum of all credits minus the sum of all debits. There is no separate balance column to drift out of sync.
Amounts are always USD, held as decimals. Never cast a balance to a float.
Ingesting events
Producers POST a CloudEvent ⧉ to /events/v1/record. Croesus accepts and stores any
well-formed event, whether or not a meter exists for its type. Unmetered events are never billed.
The CloudEvent fields map onto the internal event as follows:
| CloudEvent | Croesus | Meaning |
|---|---|---|
source + id |
Key |
Composite identity of the event |
subject |
SubjectID |
Who pays (user or org) |
type |
Kind |
Which meter, if any, processes this event |
time |
OccurredAt |
When it happened, per the producer |
data |
Data |
Arbitrary JSON the meter reads via JSONPath |
Go producers should use the croesus/event ⧉ client
rather than building requests by hand. Python producers post directly. See
app/src/app/api/v2_streaming/external/vms/billing_service.py for the established shape.
Warning
Croesus performs no authentication. Its only access control is the Cloud Run internal-ingress boundary. Anything inside the VPC can record events for any subject and credit any account.
Meters and billing policy
Both live in croesus/config.yaml ⧉ and are loaded
at startup. The config is validated on load. A meter referring to a non-existent group-by key aborts the process, and so
does a billing policy naming an unknown meter.
A meter is modelled after a simplified subset of the OpenMeter ⧉ API, so that we can move to OpenMeter later without breaking the config schema. We do not use OpenMeter today because it drags in ClickHouse.
meters:
- slug: vm_running # unique, immutable
eventType: vm_running # which events this meter consumes
valueProperty: $.duration_seconds
groupBy:
vm_id: $.vm_id
hardware_profile: $.hardware_profile
aggregation: SUM
The billing policy splits groupBy into two independent roles. Conflating them is the usual source of confusion:
pricing_keys. The subset used to look a unit price up. Forvm_runningthis ishardware_profile: the GPU type determines the price per second.transaction_keys. The subset used to split transactions. Forvm_runningthis isvm_id. Each VM gets its own line so a bill can be explained, even though all VMs of a profile share one price.
An empty transaction_keys means all readings for a subject collapse into one transaction.
Dedicated inference is metered the same way as VMs: dedicated_inference_running counts the seconds a replica is
Ready, priced by hardware_profile and split into one transaction per deployment_id. Its profiles carry a
dedicated. prefix (dedicated.h100.1x) so the customer price never depends on how the node was procured and never
collides with a vm_running row.
Tallying
Tallying turns readings into debits. Cloud Scheduler POSTs to /api/v1/tally every five minutes. Nothing runs it
in-process.
For each pass:
- Read the tally cursor. That is the arrival timestamp everything before which has been billed.
- Walk forward from the cursor in
tally_intervalsteps (default five minutes) up to now. A period that would extend past now is left for the next pass. - For every period, meter and subject: fetch readings, look up unit prices, group readings by
transaction_keys, and create one debit per group. - Advance the cursor.
Each transaction carries the idempotency key <meter-slug>:<subject-id>:<group-label>:<from>:<to>, so a retried or
overlapping pass cannot double-charge. This is why the Cloud Scheduler job retries freely. A failed tally resumes where
it stopped, and a later scheduled pass fixes what the retries did not.
Reading groups that price to zero or less are dropped with a warning rather than failing the pass. A zero-duration event should not stop billing for everyone else.
Note
A missing price is not tolerated. If a meter produces a reading whose pricing_keys combination has no row in
prices, the whole tally pass fails and the cursor does not advance. Adding a new hardware profile or serverless
model therefore means adding its price first.
Auto top-up
A subject can save a card and a threshold. When the balance falls below it, Croesus charges the card and credits the
proceeds. Unlike tallying, this runs in-process on a ticker (auto_top_up_interval, default ten seconds) with a worker
pool (auto_top_up_workers, default eight).
Guards against charging a card repeatedly:
| Setting | Default | Purpose |
|---|---|---|
auto_top_up_min_cooldown |
60s | Minimum spacing between two successful charges |
auto_top_up_failure_cooldown |
5m | Backoff after a declined charge. Validated to be at least the min cooldown |
auto_top_up_charge_timeout |
30s | After this a charge is reconciled, not retried blindly |
auto_top_up_max_failed_attempts |
3 | Consecutive declines before auto top-up is disabled |
A charge is claimed in the database before it is attempted. A unique index enforces one open charge per subject, so two workers cannot charge the same card concurrently.
When auto top-up is disabled, Croesus publishes to the croesus.events RabbitMQ exchange with routing key
auto_top_up.disabled. The Gateway consumes it and emails the customer. If LYC_RABBITMQ_HOST is unset, the publisher
is a no-op and Croesus runs normally without notifications.
Charging is invoice-first through Stripe. LegacyProcessor exists only to reconcile charges opened before that switch
and is scheduled for deletion.
API
All routes are under /api/v1 except event recording and metrics.
| Route | Method | Purpose |
|---|---|---|
/events/v1/record |
POST | Record a CloudEvent |
/metrics |
GET | Prometheus metrics |
/health |
GET | Health check |
/prices |
GET | List unit prices |
/subjects/{id}/info |
GET | Balance and suspension status |
/subjects/{id}/events |
GET | Raw events, optionally aggregated via /aggregate |
/subjects/{id}/transactions |
GET | Ledger entries, optionally aggregated via /aggregate |
/subjects/{id}/credit |
POST | Credit an account |
/subjects/{id}/debit |
POST | Debit an account |
/subjects/{id}/reset-balance |
POST | Set the balance to a given amount |
/subjects/{id}/billing-profile |
GET, PUT | Billing profile |
/subjects/{id}/billing-details |
GET, PUT | Billing details (PII, encrypted at rest) |
/subjects/{id}/auto-top-up-settings |
GET, PUT | Threshold and top-up amount |
/subjects/{id}/setup-intent |
POST | Stripe client secret for saving a card |
/tally |
POST | Run a tally pass |
Accounts are created lazily: every /subjects/{id}/… route creates the account if it does not exist, as does the first
event for a subject. Account creation also provisions a Stripe customer, so that top-ups always target the same one.
Configuration
Connection details come from LYC_* environment variables. Secrets are read from files under --secret-dir (default
/secrets).
| Variable / file | Purpose |
|---|---|
LYC_DB_CROESUS_{HOST,PORT,NAME,USER} |
CloudSQL connection |
lyc-db-croesus-password/value |
Database password |
LYC_ENCRYPTION_KEY_URI |
KMS key used to encrypt billing-detail PII |
lyc-stripe-secret-key/value |
Stripe API key |
LYC_RABBITMQ_*, lyc-rabbitmq-password/value |
Event publishing. Optional |
Known rough edges
EnsureAccounts()is called on nearly every path to create accounts on demand. It is marked temporary by aTODOincroesus/cmd/croesus/main.goand should disappear once accounts are provisioned explicitly.- The config is read once at startup. Changing a meter or price dimension needs a redeploy.
/subjects/{id}/setup-intentleaks a Stripe client secret to its caller. This goes away when card handling moves fully into Croesus.