Skip to content

Neph

Neph (from ancient Greek néphos, meaning cloud) is a Lyceum-internal tool for interacting with our infrastructure. In essence its purpose is to simplify common tasks when working with Terraform, Ansible, Kubernetes... that are either Lyceum-specific or frequently required.

Warning

Neph can operate on all of our deploy environments (via the --env option) including production! Only make modifications to the latter if you're absolutely sure what you're doing.

Installation

Neph is a Python program that is part of the monorepo and thus managed by uv. I.e. from within the monorepo you can always execute it via uv run neph ....

Configuration

Neph automatically reads config key-value pairs from /infra/inventory/*/config.yml whenever it needs a config value. Internally, it uses SOPS for reading encrypted values.

Planned direction

Neph is the infrastructure CLI and stays that: terraform, ansible, github, db, k3s, instances, ssh, tailscale and secrets access. Build commands are Mise's, and Makefiles are being retired, so nothing is offloaded onto them. Three things still move out: the billing commands, the duplicated people list-emails, and whichever commands turn out to be dead. Folding /infra/ansible and /infra/k8s invocation into neph is contemplated, not decided. /infra/inventory is also poorly named and slated for renaming. See /infra/neph and /infra/inventory.

It also provides convenience commands for working with config key-value pairs:

List all known config keys:

neph config list --env <deploy-env>

Get a specific config value:

neph config get LYC_<...> --env <deploy-env>

Finally we can dump all config key-value pairs with:

neph config dump --env <deploy-env>

The dump command also accepts --output shell to emit shell-compatible exports, or --output github to emit lines suitable for $GITHUB_ENV in a GitHub Actions workflow.

For example, in your shell:

. <(neph config dump --env <deploy-env> --output shell)

Or in a GitHub Actions workflow:

run: neph config dump --env <deploy-env> --output github >> "$GITHUB_ENV"

Terraform

neph runs Terragrunt for one component of one environment, with that environment's config and secrets:

neph terraform --env <deploy-env> --component <component>

<component> is a directory under infra/terraform, e.g. gcp/iris for one unit or gcp for the whole stack. The command runs init, plan and apply, and prompts before the apply. These options change that:

  • --plan-only stops after the plan.
  • --init-only stops after init.
  • --yes skips the prompt. A set $CI does the same.
  • --feature name=value sets a Terragrunt feature flag.

Two more commands drive Terragrunt with that same environment:

neph terragrunt --env <deploy-env> --component <component> -- <terragrunt args>
neph terraform-import --env <deploy-env> --component <component> <address> <id>

neph terragrunt runs any Terragrunt command and streams its output unchanged, so show -json pipes into a file. neph terraform-import runs init, then imports one existing resource into the unit's state. Both prompt in production like neph terraform does.

After a unit's state bucket or prefix changes, move its state before any plan:

neph terraform-move-state --env <deploy-env> --component <component> --from gs://<bucket>/<old prefix>

neph inits with -reconfigure, which drops an old backend instead of migrating it. Without the move, the next plan starts from an empty state. The command copies the state into the backend the unit names now and checks that it arrived. It works for a Terragrunt unit and for a legacy Terraform root.

  • It refuses when the new backend already holds a different state. An empty state, such as the one init writes to a new GCS prefix, does not count.
  • It is a no-op when the new backend already holds this state, or when the old state is empty.
  • It leaves the old object in place. Delete it once a plan of the unit is clean.
  • It reads the old location through a throwaway init, which writes an empty state there when it holds none. Delete that object too.

mise run //infra/neph:move-state takes the same flags.

If the environment's config holds a key in LYC_GCP_SERVICE_ACCOUNT_CREDENTIALS, then Terraform authenticates as that service account. Otherwise it authenticates with your Application Default Credentials. In GitHub Actions these are the credentials the auth step names in GOOGLE_APPLICATION_CREDENTIALS.

Ansible

neph runs a playbook against one environment, with a dynamic inventory built from that environment's config and VMs:

neph ansible --env <deploy-env> --playbook <playbook> [--extra-arg <arg>]...

<playbook> names a file under infra/ansible/playbooks, e.g. ubuntu-gateway.yml. A glob such as 'ubuntu-*' runs several. The devops SSH key is written from the environment's config before the run.

Deploy

Publish lyceum-cli to PyPI from the main branch:

neph deploy lyceum-cli --env <deploy-env>

The version comes from app/src/lyceum-cli/lyceum/_version.py and must not exist on PyPI yet. The upload token is LYC_PYPI_TOKEN_SECRET in infra/inventory/common/config.yml. Pass --dry-run to build without uploading.

The preferred way to release is the Ops - Publish lyceum-cli GitHub Actions workflow (ops-publish-lyceum-cli.yml), which runs this command from main. Trigger it via gh workflow run ops-publish-lyceum-cli.yml or from the Actions tab.

Kubernetes

The main way to interact with our k3s clusters is via kubectl. To bootstrap that, Neph can automatically create a kubeconfig file for you:

neph k8s fetch-kubeconfig --env <deploy-env> (--file <file>) (--overwrite)

If --file is not specified, the kubeconfig is written to ~/.kube/config. If you already have that file it is automatically updated to include the k3s cluster. By default, existing entries for a cluster will not be overwritten. If the cluster is recreated or the cert is rotated, you'll have to specify --overwrite to do so.

You can then switch to the cluster with:

kubectl config use-context lyceum-k3s-<deploy-env>

Interacting with databases

Neph can be used to quickly drop into a psql session for one of our databases. To do this simply run:

neph db connect <db-name> --env <deploy-env> (--admin-access)

--admin-access is typically not required. You need it when the default database user lacks the privileges, for example to delete from or drop tables.

You can also run DB migrations with:

neph db migrate <db-name> --env <deploy-env> (--force-version <force-version>)

If --force-version is set, migrations will continue from the next version after the specified one or the very first if --force-version is set to zero.

Billing

In order to simplify testing, Neph can access some of Croesus' (TODO(Timo Nicolai): Add link) internal endpoints. Since billing happens on an org basis, personal org IDs have to be supplied to interact with the billing accounts of specific users.

Supported operations are:

neph billing org-info <org-id> --env <deploy-env>

which returns info about the org's current balance and suspension status, and:

neph billing (credit-org|debit-org|reset-org-balance) <org-id> <amount> --env <deploy-env>

where <amount> must be a valid decimal number.