Neph
Neph (from ancient Greek néphos, meaning cloud) is a Lyceum-internal tool for interacting with our infrastructure. In essence its purpose is to simplify common tasks when working with Terraform, Ansible, Kubernetes... that are either Lyceum-specific or frequently required.
Warning
Neph can operate on all of our deploy environments (via the --env option)
including production! Only make modifications to the latter if you're
absolutely sure what you're doing.
Installation
Neph is a Python program that is part of the monorepo and thus managed by uv.
I.e. from within the monorepo you can always execute it via uv run neph ....
Configuration
Neph automatically reads config key-value pairs from
/infra/inventory/*/config.yml whenever it needs a config value. Internally,
it uses SOPS for reading encrypted values.
Planned direction
Neph is the infrastructure CLI and stays that: terraform, ansible, github, db, k3s, instances,
ssh, tailscale and secrets access. Build commands are Mise's, and Makefiles are being retired, so nothing
is offloaded onto them. Three things still move out: the billing commands, the duplicated
people list-emails, and whichever commands turn out to be dead. Folding /infra/ansible and /infra/k8s
invocation into neph is contemplated, not decided. /infra/inventory is also poorly named and slated for
renaming. See
/infra/neph and
/infra/inventory.
It also provides convenience commands for working with config key-value pairs:
List all known config keys:
neph config list --env <deploy-env>
Get a specific config value:
neph config get LYC_<...> --env <deploy-env>
Finally we can dump all config key-value pairs with:
neph config dump --env <deploy-env>
The dump command also accepts --output shell to emit shell-compatible
exports, or --output github to emit lines suitable for $GITHUB_ENV in a
GitHub Actions workflow.
For example, in your shell:
. <(neph config dump --env <deploy-env> --output shell)
Or in a GitHub Actions workflow:
run: neph config dump --env <deploy-env> --output github >> "$GITHUB_ENV"
Terraform
neph runs Terragrunt for one component of one environment, with that environment's config and secrets:
neph terraform --env <deploy-env> --component <component>
<component> is a directory under infra/terraform, e.g. gcp/iris for one unit or gcp for
the whole stack. The command runs init, plan and apply, and prompts before the apply. These options
change that:
--plan-onlystops after the plan.--init-onlystops after init.--yesskips the prompt. A set$CIdoes the same.--feature name=valuesets a Terragrunt feature flag.
Two more commands drive Terragrunt with that same environment:
neph terragrunt --env <deploy-env> --component <component> -- <terragrunt args>
neph terraform-import --env <deploy-env> --component <component> <address> <id>
neph terragrunt runs any Terragrunt command and streams its output unchanged, so show -json
pipes into a file. neph terraform-import runs init, then imports one existing resource into the
unit's state. Both prompt in production like neph terraform does.
After a unit's state bucket or prefix changes, move its state before any plan:
neph terraform-move-state --env <deploy-env> --component <component> --from gs://<bucket>/<old prefix>
neph inits with -reconfigure, which drops an old backend instead of migrating it. Without the move,
the next plan starts from an empty state. The command copies the state into the backend the unit
names now and checks that it arrived. It works for a Terragrunt unit and for a legacy Terraform root.
- It refuses when the new backend already holds a different state. An empty state, such as the one init writes to a new GCS prefix, does not count.
- It is a no-op when the new backend already holds this state, or when the old state is empty.
- It leaves the old object in place. Delete it once a plan of the unit is clean.
- It reads the old location through a throwaway init, which writes an empty state there when it holds none. Delete that object too.
mise run //infra/neph:move-state takes the same flags.
If the environment's config holds a key in LYC_GCP_SERVICE_ACCOUNT_CREDENTIALS, then Terraform
authenticates as that service account. Otherwise it authenticates with your Application Default
Credentials. In GitHub Actions these are the credentials the auth step names in
GOOGLE_APPLICATION_CREDENTIALS.
Ansible
neph runs a playbook against one environment, with a dynamic inventory built from that environment's config and VMs:
neph ansible --env <deploy-env> --playbook <playbook> [--extra-arg <arg>]...
<playbook> names a file under infra/ansible/playbooks, e.g. ubuntu-gateway.yml. A glob such
as 'ubuntu-*' runs several. The devops SSH key is written from the environment's config before
the run.
Deploy
Publish lyceum-cli to PyPI from the main branch:
neph deploy lyceum-cli --env <deploy-env>
The version comes from app/src/lyceum-cli/lyceum/_version.py and must not exist on PyPI yet. The upload token is LYC_PYPI_TOKEN_SECRET in infra/inventory/common/config.yml. Pass --dry-run to build without uploading.
The preferred way to release is the Ops - Publish lyceum-cli GitHub Actions workflow (ops-publish-lyceum-cli.yml), which runs this command from main. Trigger it via gh workflow run ops-publish-lyceum-cli.yml or from the Actions tab.
Kubernetes
The main way to interact with our k3s clusters is via kubectl. To bootstrap
that, Neph can automatically create a kubeconfig file for you:
neph k8s fetch-kubeconfig --env <deploy-env> (--file <file>) (--overwrite)
If --file is not specified, the kubeconfig is written to ~/.kube/config. If
you already have that file it is automatically updated to include the k3s
cluster. By default, existing entries for a cluster will not be overwritten. If
the cluster is recreated or the cert is rotated, you'll have to specify
--overwrite to do so.
You can then switch to the cluster with:
kubectl config use-context lyceum-k3s-<deploy-env>
Interacting with databases
Neph can be used to quickly drop into a psql session for one of our
databases. To do this simply run:
neph db connect <db-name> --env <deploy-env> (--admin-access)
--admin-access is typically not required. You need it when the default
database user lacks the privileges, for example to delete from or drop tables.
You can also run DB migrations with:
neph db migrate <db-name> --env <deploy-env> (--force-version <force-version>)
If --force-version is set, migrations will continue from the next version
after the specified one or the very first if --force-version is set to zero.
Billing
In order to simplify testing, Neph can access some of Croesus' (TODO(Timo Nicolai): Add link) internal endpoints. Since billing happens on an org basis, personal org IDs have to be supplied to interact with the billing accounts of specific users.
Supported operations are:
neph billing org-info <org-id> --env <deploy-env>
which returns info about the org's current balance and suspension status, and:
neph billing (credit-org|debit-org|reset-org-balance) <org-id> <amount> --env <deploy-env>
where <amount> must be a valid decimal number.