ai-assisted-todolist

The database#

Migration is a one-off ECS task running Liquibase, never the application. migrate-at-start stays on for dev and test, where overlap cannot happen, and off everywhere else.

Rollback of a migration is a snapshot restore, taken immediately before the migration ran. Liquibase rollbacks exist and are not exercised here; a restore is the honest answer and it costs the downtime that this deployment model already accepts.

task_state and task_importance are native PostgreSQL enums, and ALTER TYPE … ADD VALUE cannot be rolled back at all. A release adding an enum value is one-way: it must ship before anything uses the value, and the only way back is the snapshot. That is the price of the enum decision in decisions/domain-and-backend.md, and it is the sharpest edge in this plan.

Idle cost is handled by destroying the environment, database included, with a final snapshot. Stopping is not enough: a stopped RDS instance still bills storage and AWS restarts it automatically after seven days.

A down/up cycle keeps the data, not the instance. env.sh down destroys the instance and leaves a final snapshot named taskfest-<env>-db-final-<timestamp>; env.sh up looks up the newest one and passes it as db_restore_snapshot, so the new instance starts from it. With no snapshot it starts empty, and it says which on screen. To start empty on purpose, or from an older snapshot, run the plan by hand with -var db_restore_snapshot=… (or =null). Snapshots are never deleted automatically; at this size each is cents a month, and old ones are removed in the console when wanted.

The application logs in with IAM, not a password. The ECS task role (taskfest-<env>-task) may rds-db:connect as one database user, taskfest_<env>, and nothing else. On every new connection RdsIamCredentialsProvider signs a token from the task role's credentials, valid for 15 minutes; RDS checks it against IAM. Nothing secret is configured, injected or rotated, and the counterpart on EKS would be a service account. The backend switches it on with TASKFEST_DATASOURCE_CREDENTIALS_PROVIDER=rds-iam; without it, as in the Compose stacks, the password is used as before. Migrations run as the same user.

The master password is RDS's, and is for bootstrapping only. manage_master_user_password has RDS generate it, keep it in Secrets Manager and rotate it every seven days, so it is never in state. Injecting it into the running service was rejected for exactly that rotation: ECS reads a secret once, at task start, so a long-running task would keep the old password and fail on its next connection after a rotation.

Creating the database user is one command, once per environment. taskfest_<env> does not exist until the master user creates it, and nothing outside the VPC can reach the database to do that, so it is a one-off ECS task:

./env.sh up qa             # the database, and the two task definitions that point at it
./env.sh db-bootstrap qa   # creates taskfest_qa: CREATE ROLE, GRANT rds_iam, schema grants
./env.sh migrate qa        # Liquibase, logged in as taskfest_qa with an IAM token
./env.sh up qa             # again, on a brand-new environment only -- see below

On a brand-new environment the first up fails, and that is expected. It creates the database, then waits for the service, whose backend cannot log in as a user that does not exist yet: FATAL: password authentication failed for user "taskfest_qa". Because it is the service's first deployment, ECS has nothing to roll back to and up stops with No rollback candidate was found. Everything up to the service is in place by then, so db-bootstrap and migrate work, and the second up replaces the failed service and succeeds. This happened when the environments were rebuilt under the TaskFest name (#128); an environment restored from a snapshot already has its user and never sees it.

db-bootstrap runs psql from the official PostgreSQL image (pulled from ECR Public's mirror, which has no Docker Hub rate limit) with the SQL in modules/environment/db-bootstrap.sql. It is the only task that receives the master credentials, and the only one the deploy role cannot run. It verifies the server with the region's RDS CA bundle, which is committed under modules/environment/rds-ca/ and passed in as a variable, because the image does not carry it. Every statement is idempotent, so running it again — or on a restored database, which already has the user — changes nothing. It is needed once per environment: every later up restores the user along with the data.

migrate is the backend image with quarkus.init-and-exit, so Quarkus runs Liquibase and stops instead of serving. It logs in exactly as the application will, with verify-full against the RDS bundle that backend/src/main/jib/opt/rds/ puts in the image, so a run that exits 0 proves the whole path: the user exists, RDS accepts the token, and the certificate checks out. Both commands print the task's log and exit with its exit code. The tasks run quay.io/ghilling/taskfest-backend:latest until the deploy change pins a release.

A restore needs its managed password re-established. For PostgreSQL, RDS cannot turn on managed credentials during a snapshot restore — AWS supports that for Oracle only — so a restored instance comes back with the master password from the time of the snapshot, whose secret was deleted with the old instance. The AWS provider follows the restore with a ModifyDBInstance that turns managed credentials back on and creates a new secret. That is expected rather than verified; if the first restore shows no db_master_secret_arn, the fallback is aws rds modify-db-instance --db-instance-identifier taskfest-<env>-db --manage-master-user-password --apply-immediately.