ai-assisted-todolist

Deployment#

Status: a plan. Nothing in this document exists yet, and issue #60 deliberately creates nothing — the point was to settle the irreversible choices on paper first.

This is also a proof of concept for a more commercial project, which is why it prefers managed services and a real account boundary over the cheapest possible demo.

The shape#

Two environments, qa and prod, in separate AWS accounts. Each holds the same thing:

  todolist.cloud.hilling.de
  todolist-qa.cloud.hilling.de
            │
            ▼
       CloudFront ──── /*      ──> S3                    (the built SPA)
            │
            └───────── /api/*  ──> ALB ──> ECS Fargate ──> RDS PostgreSQL
                                (internal)   (the Quarkus backend)

One distribution, two origins. That is not a detail — see the next section.

The load balancer is internal: CloudFront reaches it through a VPC origin, so it has no public address at all. That closes the hole where someone finds the load balancer's hostname and bypasses CloudFront entirely, and it costs nothing — it replaces the older trick of having CloudFront send a shared secret header that the load balancer checks.

Google is the identity provider in both environments, as it already is in the prod profile. Keycloak stays local-only.

What changes from the local stack, and the one thing that must not break#

Locally, httpd does two jobs: it serves the built SPA and it reverse-proxies /api to the backend. architecture.md explains why the second one matters — the backend builds its OIDC redirect_uri from the Host header, and the session cookie is same-origin.

Moving the SPA to S3 removes httpd. If CloudFront served only S3, the API would become cross-origin and sign-in would break — after deployment, in a way nothing local would catch.

So CloudFront takes both roles:

Path Origin Behaviour
/* S3 (private, OAC) cached; 403/404 rewritten to /index.html with status 200
/api/* ALB caching disabled; Host, Authorization and cookies forwarded

The error-response rule replaces httpd's FallbackResource /index.html, which is what makes a deep link work. The cache rule on /api/* is what stops a logged-in user's response being served to someone else — the failure mode worth being most careful about.

The frontend container image is no longer deployed anywhere. It stays in deployment/docker/ for local work, which means the Compose stack is no longer the same shape as production; the local-development.md claim that it is "the stack for checking the deployment shape" needs qualifying when this is built.

Why there is a load balancer at all#

It is the largest fixed item here, so it was challenged on cost twice. It stays, and the arithmetic is worth writing down because the intuition is wrong.

The alternative is not free. An API Gateway HTTP API needs a VPC link to reach the tasks, and a VPC link costs $0.01/hour — about $7.20 a month, charged at zero traffic. Requests are $1.00 per million on top, which is nothing here. So the comparison is roughly $16 against $7.20: a saving of about $9 a month per environment, not the $16 it first looks like.

What that $9 costs is blue/green. ECS shifts traffic between two target groups, which is a load balancer's job; an API Gateway private integration cannot do it. Deployments would become rolling or stop-and-start — acceptable, since downtime is acceptable here, but the demonstration goes.

The real lever is uptime, not the load balancer. An environment that exists costs about $35 a month before the load balancer; one that has been destroyed costs cents, including its load balancer. Nine dollars a month applies only while an environment is up, and the design already says QA should not be. Keeping the load balancer therefore costs little in practice and keeps a capability that is worth demonstrating in a proof of concept.

Also checked and not chosen:

Names and certificates#

The zone cloud.hilling.de already exists in Route 53.

Environment Host
QA todolist-qa.cloud.hilling.de
Prod todolist.cloud.hilling.de

Each is an A and AAAA alias record pointing at that environment's CloudFront distribution; alias records to CloudFront are not charged.

Two things that catch people, both worth knowing before building:

Alias records, not CNAMEs, and the cost difference is real though small here. Route 53 does not charge for queries to an alias record that points at an AWS resource; a CNAME is a standard record at $0.40 per million queries, and a CNAME pointing at another name in the same zone is billed as two queries because the resolver has to ask twice. At demo traffic that is fractions of a cent either way — but an alias is free, avoids the second lookup, and is the only one of the two that would work at a zone apex. There is no case for the CNAME.

The Google OIDC configuration#

Fixed hostnames make these knowable now, so they can be set up once the environments exist. The backend uses quarkus.oidc.authentication.redirect-path=/api/auth/callback, and /api/* routes through the same distribution, so the authorised redirect URIs are:

This is manual configuration in the Google console and cannot be automated here.

What is built (#120), per environment:

sh aws secretsmanager put-secret-value --secret-id todolist-qa-google-client-secret \ --secret-string '<the client secret from the Google console>'

The task execution role may read that one secret and injects it at task start as TODO_OIDC_GOOGLE_CLIENT_SECRET. It does not rotate by itself, unlike the RDS master secret, so injection at start is fine; after changing it, start new tasks. About $0.40 a month, and it stays while the environment is down. - Behind CloudFront and the ALB the backend sees plain HTTP, so it is told to believe the load balancer's X-Forwarded-Proto (quarkus.http.proxy.*, set in the task definition, trusted from inside the VPC only). Without that it asks Google to return to http://…, which Google rejects as an unregistered redirect URI. SignInBehindProxyTest pins it. - Who may sign in is decided in the Google console: with the app in Testing, only the Google accounts listed as test users can.

The documentation site's own name#

Google's consent screen asks for an application home page, a privacy policy and terms of service, all under the authorised domain hilling.de. They are the GitHub Pages site, under https://todolist-docs.cloud.hilling.de/: the home page is the rendered README, and privacy.md and terms.md are linked from the footer of every page. The name is a CNAME to guhilling.github.io, managed in the account/ root (docs_hostname), and the repository's Pages settings name the same host, which is what makes GitHub serve the site there and issue its certificate.

One account, and what that costs in guarantees#

QA and prod live in one AWS account, separated by IAM roles and resource tags. This reverses an earlier decision; the reason is that the hosted zone and the certificates live there too, and splitting the account turned every DNS record and every ACM validation into a cross-account operation for two hostnames.

Be clear about what is given up. With separate accounts, "prod is never created by an automated agent" was enforceable by there being nothing in the prod account to assume. In one account it becomes a policy guarantee, which is weaker:

Who may do what#

Three identities, in order of how much they may do. A human IAM admin creates the free foundation once: the VPC, subnets, security groups, IAM, the certificate and the DNS records. A role that same human assumes creates and destroys everything that bills: the database, the load balancer and the Fargate service. Below the line, the only unattended identity is the deploy role, assumed from GitHub Actions through OIDC, which may update the service and the site bucket and cannot create, change or delete any infrastructure. Separately, the task roles are what the container runs as rather than an identity anyone assumes, and a read-only user exists for looking at things.

The ordering principle is what a mistake costs, which is not the same as what a resource does. The foundation is free and effectively permanent, so it is created once and rarely touched. The billing resources are the ones torn down and rebuilt whenever an environment is not being demoed, which makes that a routine operation rather than a one-off. And redeploying the application happens many times a day, unattended — so it gets the identity that can do the least.

The line in the picture is the claim: the only identity that runs unattended is the one that cannot create or destroy anything. Everything with a bill attached needs a person.

Three technical identities, one job each:

Identity May do Used by
gunnar (IAM admin user) anything, including the free foundation a person, with MFA. Not the account root user, which has no access keys and is used for nothing
todolist-<env>-lifecycle (role) create and destroy what bills: RDS, the load balancer, the Fargate service — and the CloudFront VPC origin that follows the load balancer a person assuming it, with MFA
todolist-qa-deploy (role) redeploy the qa application; no infrastructure GitHub Actions, unattended
todolist-prod-deploy (role) redeploy the prod application; no infrastructure GitHub Actions, only from the prod environment, which requires approval
todolist-monitoring (user) read-only, everywhere dashboards and a local CLI profile; can change nothing

Why a separate lifecycle role, rather than just using the admin user. Standing an environment up and tearing it down again is the routine operation in this project's cost model, not a one-off — so it is worth having an identity whose blast radius is that operation, and an audit trail that says which of the two things a given session was doing. It is also the seam at which a "bring qa up" workflow could later exist.

It is assumed rather than attached, and the trust policy requires MFA — so using it is a deliberate act with a timestamp, which is most of the point. Sessions last an hour.

Be honest about what that role is, though. Creating an RDS instance and a load balancer needs create and delete on those services plus iam:PassRole for the task roles, which is close to administrator for the environment. Scoping it by env tag helps and is worth doing, but it is a scoping and audit boundary, not a security boundary: it will not survive a determined misuse the way the deploy role's policy will. The practical consequence is that it stays human-assumed. If a qa up/down workflow is ever built, it may use this role because qa is disposable; prod must not, because "delete the database" is one API call and the snapshot is the only way back. The split is by lifecycle stage, not by tag. Infrastructure — the VPC, the database, the load balancer, the cluster — is created by a human with their own privileges. The deploy identities cannot touch any of it. This is stronger than scoping an apply credential by tag, and for a structural reason: a role that could run tofu apply needs create and delete on everything the configuration manages, which is administrator for that environment under a different name. Not creating such a credential is a guarantee; constraining one is an argument.

A deploy is therefore four things and nothing else: register a task definition, point the service at it, run the Liquibase task, and sync the frontend artifact into the site bucket.

The deploy identities are roles assumed through GitHub OIDC, not users with access keys. A workflow job trades a signed token describing itself for credentials that expire in minutes, so no secret is stored in GitHub and there is nothing to rotate. The trust condition names the repository and the GitHub environment, which is what makes the prod approval gate an AWS refusal rather than only a GitHub courtesy: GitHub will not mint a token claiming environment:prod unless the job declares it, and a protected environment holds the job until approved.

Four consequences worth being plain about:

State is one S3 bucket with a key per environment. There is no lock table: OpenTofu's S3 backend takes the lock from S3 itself, so the bootstrap is one bucket and nothing else.

Where the code is#

deployment/aws-tofu/ implements this, and its README.md covers running it — bootstrapping the state bucket, the credential-free checks that CI runs, and the provider lock file.

The tool is OpenTofu, not Terraform. The language is identical, the binary is tofu, and two of its features are used deliberately: a backend block may interpolate a variable, so the environment roots share one backend.tf; and the S3 backend takes its lock from S3, so there is no DynamoDB table. decisions.md records the choice.

The layout is one module holding every resource, instantiated by a thin root per environment:

modules/environment/     every resource, parameterised
environments/qa/         a root that instantiates the module
environments/prod/       the same root, different values
account/                 the things there is one of per AWS account

account/ holds the GitHub OIDC provider and the monitoring user. Its permissions come from a read-only group: ReadOnlyAccess, everything denied without MFA except registering a device, and the state bucket's objects denied even with it. ReadOnlyAccess reads every bucket, so without the MFA rule a leaked access key would have read prod's state. From the CLI the user therefore needs a session from aws sts get-session-token --serial-number … --token-code …; the bare access key is refused. It exists because those are singular: instantiating them from a module applied twice would have the two environments fighting over one provider. It is applied before either environment, which look the provider up by URL.

The two roots are identical apart from terraform.tfvars — the state key is interpolated from a variable, which Terraform could not do — and check-environments-match.py fails CI when they are not. Making prod a parameter change is the intent of this plan; that check is what keeps it from decaying into two codebases that drift. Anything which must differ between environments becomes a module variable instead.

Cost attribution moves to tags. With one account there is no per-account bill, so every resource carries env=qa or env=prod and the budgets filter on that. This only works if the tagging is complete: an untagged resource is invisible to both environment budgets and shows up only in the total.

Resource inventory#

Per environment, identical unless noted:

Resource Why
VPC, two public subnets in two AZs Two AZs because the load balancer requires it, not for availability
No NAT gateway See Cost — this is the single biggest saving, and it is why the tasks sit in the public subnets
Security groups ALB open on 443; ECS open only to the ALB; RDS open only to ECS
VPC flow logs, to a private S3 bucket All traffic, kept 30 days; see Observability. Not behind the teardown switch, because delivery is charged per GB and a torn-down VPC sends almost nothing
Internal ALB, two target groups Blue/green shifts traffic between them; reachable only from CloudFront
Two private subnets For the load balancer only. No NAT gateway: nothing in them makes outbound calls
ACM certificate in us-east-1 CloudFront accepts certificates from that region only
Route 53 alias records In the same account as everything else; free to query
ECS cluster, one Fargate service The Quarkus backend, 0.5 vCPU / 1 GB
ECS cluster, and two one-off tasks migrate (the backend image with quarkus.init-and-exit) and db-bootstrap (psql, once per environment); see The database
RDS PostgreSQL, single-AZ, db.t4g.micro HA is explicitly not required
Secrets Manager The RDS-managed master password, for bootstrapping and administration only, and the Google client secret. The application has no database password at all
Three IAM users and their policies The technical identities above, scoped by tag
S3 bucket, private The built SPA; reachable only through CloudFront's origin access control
CloudFront distribution The two origins above
CloudWatch log groups, alarms, dashboard See Observability
AWS Budgets See Cost

ECS tasks run in public subnets with a public IP, which is what avoids the NAT gateway. They are not reachable from outside: the security group admits only the ALB. This is a deliberate demo-scale trade — a commercial deployment would use private subnets and pay for NAT or VPC endpoints, and the reasoning is in decisions.md.

Deploying: two paths, chosen by the release#

Blue/green and "take the downtime" are opposites — blue/green keeps both versions live, which is exactly the overlap that makes a schema change unsafe. Both are wanted, so the release picks:

A release with no migration — blue/green. ECS shifts traffic to a new task set, bakes, and rolls back by shifting back. Zero downtime, and the rollback is instant because the old task set is still there. This is the common case.

A release with a migration — downtime. Scale the service to zero, snapshot the database, run Liquibase as a one-off ECS task, deploy the new version, scale back up. Nothing old is running, so nothing has to be backward-compatible and expand-and-contract is not needed.

The image decides which. It carries the changelog it expects as a label and a file; the deployment compares that with the database and takes the corresponding path. Hibernate already runs schema-management.strategy=validate, so a mismatch already fails fast — this makes it fail before any traffic moves, and say why.

Deployment runs from GitHub Actions as the environment's technical user. No CodePipeline; the prod gate is a protected environment, which is also what keeps the prod key unreadable from an unapproved run.

The SPA deploys as aws s3 sync of a release artifact, the same way openapi.yaml is already attached to each release. Rolling the frontend back is re-syncing the previous release, which is why the build has to be an artifact rather than something rebuilt at deploy time.

This half exists (#119). release.yml attaches todo-frontend-<version>.tar.gz, the built dist/, to every release, and deploy-frontend.yml — started by hand from the Actions tab with an environment and a tag — syncs it into that environment's site bucket as the deploy role. It uploads in three passes so no viewer ever sees an index.html naming an asset that is not there yet: the hashed assets/ first, cached for a year and never deleted (a browser holding the old page keeps working, and so does a rollback); then the remaining files, cached five minutes; index.html last with no-cache, which under CloudFront's CachingOptimized policy means it is held for its one-second minimum — so a release is visible at once and no invalidation is needed.

Deep links are a CloudFront Function on the default behaviour (spa-routing.js): a path whose last segment has no dot gets index.html, and a path that names a file is passed through, so a missing asset is still the bucket's 403. It replaces httpd's FallbackResource. CloudFront's custom error responses would have been the usual way, and are not used because they apply to every origin: /api/* errors would have become the app with a 200. While an environment is down there is no /api/* behaviour, so API requests reach this function too; it answers them with a plain 503, "The backend is not running in this environment", rather than the app — the first deployment showed /api/auth/providers coming back as index.html with a 200 until it did.

How an environment is torn down#

Teardown is a parameter, not a tofu destroy. Each environment has a running variable; the resources that cost money — the database, the load balancer, the service — exist only when it is true. The foundation ignores it and is always there, because it is free and destroying it buys nothing.

The CloudFront distribution stays; its /api/* origin does not. A CloudFront VPC origin names one load balancer ARN and cannot be changed or deleted while a distribution uses it, and the load balancer is destroyed on every down. So down drops the distribution's /api/* behaviour and deletes the VPC origin, and up creates a new one and adds the behaviour back. Deploying a VPC origin takes up to 15 minutes and a distribution change several more, so up and down each take roughly 25–35 minutes. That is the price of a down that costs nothing; keeping the load balancer up instead would have cost about $20 a month while nobody was using the environment. While down, the distribution still answers, from the site bucket alone.

deployment/aws-tofu/env.sh up qa       # create what bills
deployment/aws-tofu/env.sh down qa     # destroy it; VPC, subnets and IAM stay
deployment/aws-tofu/env.sh status qa   # what the last apply recorded

Three properties of this are deliberate:

The alternative was a separate root and state for the billable layer, where tofu destroy could not reach the foundation. It was rejected: the environment roots are already applied, so their state keys would have had to move, and a parameter is the same mechanism this directory already uses for the difference between qa and prod.

env.sh shows the plan and waits for an answer on every run, and applies the saved plan file rather than re-evaluating — so what is applied is exactly what was displayed. It defaults to the lifecycle profile for up and down, and deliberately not for status, since reading what the last apply recorded needs nothing but the state bucket.

The database#

Migration is a one-off ECS task running Liquibase, never the application. migrate-at-start stays on for dev and test, where overlap cannot happen, and off everywhere else.

Rollback of a migration is a snapshot restore, taken immediately before the migration ran. Liquibase rollbacks exist and are not exercised here; a restore is the honest answer and it costs the downtime that this deployment model already accepts.

task_state and task_importance are native PostgreSQL enums, and ALTER TYPE … ADD VALUE cannot be rolled back at all. A release adding an enum value is one-way: it must ship before anything uses the value, and the only way back is the snapshot. That is the price of the enum decision in decisions.md, and it is the sharpest edge in this plan.

Idle cost is handled by destroying the environment, database included, with a final snapshot. Stopping is not enough: a stopped RDS instance still bills storage and AWS restarts it automatically after seven days.

A down/up cycle keeps the data, not the instance. env.sh down destroys the instance and leaves a final snapshot named todolist-<env>-db-final-<timestamp>; env.sh up looks up the newest one and passes it as db_restore_snapshot, so the new instance starts from it. With no snapshot it starts empty, and it says which on screen. To start empty on purpose, or from an older snapshot, run the plan by hand with -var db_restore_snapshot=… (or =null). Snapshots are never deleted automatically; at this size each is cents a month, and old ones are removed in the console when wanted.

The application logs in with IAM, not a password. The ECS task role (todolist-<env>-task) may rds-db:connect as one database user, todolist_<env>, and nothing else. On every new connection RdsIamCredentialsProvider signs a token from the task role's credentials, valid for 15 minutes; RDS checks it against IAM. Nothing secret is configured, injected or rotated, and the counterpart on EKS would be a service account. The backend switches it on with TODO_DATASOURCE_CREDENTIALS_PROVIDER=rds-iam; without it, as in the Compose stacks, the password is used as before. Migrations run as the same user.

The master password is RDS's, and is for bootstrapping only. manage_master_user_password has RDS generate it, keep it in Secrets Manager and rotate it every seven days, so it is never in state. Injecting it into the running service was rejected for exactly that rotation: ECS reads a secret once, at task start, so a long-running task would keep the old password and fail on its next connection after a rotation.

Creating the database user is one command, once per environment. todolist_<env> does not exist until the master user creates it, and nothing outside the VPC can reach the database to do that, so it is a one-off ECS task:

./env.sh up qa             # the database, and the two task definitions that point at it
./env.sh db-bootstrap qa   # creates todolist_qa: CREATE ROLE, GRANT rds_iam, schema grants
./env.sh migrate qa        # Liquibase, logged in as todolist_qa with an IAM token

db-bootstrap runs psql from the official PostgreSQL image (pulled from ECR Public's mirror, which has no Docker Hub rate limit) with the SQL in modules/environment/db-bootstrap.sql. It is the only task that receives the master credentials, and the only one the deploy role cannot run. It verifies the server with the region's RDS CA bundle, which is committed under modules/environment/rds-ca/ and passed in as a variable, because the image does not carry it. Every statement is idempotent, so running it again — or on a restored database, which already has the user — changes nothing. It is needed once per environment: every later up restores the user along with the data.

migrate is the backend image with quarkus.init-and-exit, so Quarkus runs Liquibase and stops instead of serving. It logs in exactly as the application will, with verify-full against the RDS bundle that backend/src/main/jib/opt/rds/ puts in the image, so a run that exits 0 proves the whole path: the user exists, RDS accepts the token, and the certificate checks out. Both commands print the task's log and exit with its exit code. The tasks run quay.io/ghilling/todo-backend:latest until the deploy change pins a release.

A restore needs its managed password re-established. For PostgreSQL, RDS cannot turn on managed credentials during a snapshot restore — AWS supports that for Oracle only — so a restored instance comes back with the master password from the time of the snapshot, whose secret was deleted with the old instance. The AWS provider follows the restore with a ModifyDBInstance that turns managed credentials back on and creates a new secret. That is expected rather than verified; if the first restore shows no db_master_secret_arn, the fallback is aws rds modify-db-instance --db-instance-identifier todolist-<env>-db --manage-master-user-password --apply-immediately.

Observability#

Named, because "best practices are applied" plans nothing:

Cost#

These are estimates from published US East rates in September 2026, not a quote, and eu-central-1 differs. They are here for shape and for setting alert thresholds; the plan does not pretend to a precision it cannot have without a live account.

One environment, running continuously:

Estimate/month Note
ALB ~$20 Fixed, whether or not anyone uses it, while the environment is up. Unavoidable while blue/green is demonstrated. eu-central-1 estimate
Fargate 0.5 vCPU / 1 GB ~$20 One task, x86. eu-central-1 estimate
RDS db.t4g.micro single-AZ ~$13 Plus ~$3 for 20 GB gp3. Seven days of backups and Performance Insights are within the free allowances. An estimate until qa has run for a month
S3 + CloudFront ~$1 At demo traffic
CloudWatch, Secrets Manager ~$2
Route 53 $0 The zone already exists; alias queries are not charged
NAT gateway ~$33 — avoided Would have been the largest line item
Total ~$50 ~$83 with a NAT gateway

Both environments up at once is therefore about $100 a month, which is the number the total budget alarm is really guarding against.

Idle, after tofu destroy: a few cents of snapshot and S3 storage. This is the point of making QA disposable, and it is worth more than any per-resource tuning.

Proposed budget alarms, for sign-off: QA $25, prod $40, total $75, alerting at 80% of forecast and again at 100% of actual. They are set below the running-continuously estimate on purpose — an environment left up is exactly what the alarm is for.

Deliberately not in this plan#

To verify before building#

Honest gaps, because a plan that hides them is worse than one that names them: