Repository navigation
Make Terraform rebuild the production stack, and document the process - #6
Merged
Merged
Conversation
The configuration had never been applied, and a first run surfaced several defects. Applying it now rebuilds the stack exactly as the AWS CLI did. Placement was non-deterministic. Both instances and the Instance Connect Endpoint selected data.aws_subnets.default.ids[0], but that list comes back in AWS API order with no stability guarantee, so the zone could change between runs and force replacement. A dedicated data source now resolves the subnet by availability zone, which restores the ap-south-1a placement the original build used. Engine versions were unpinned, so a rebuild would take whatever AWS defaults to at the time. Postgres is pinned to 18 and Redis to 7.1 to match what ran in production, keeping the database aligned with the Drizzle migrations. The admin target group never passed its health check. The panel redirects the root path to /admin, so the default matcher of 200 saw a 307 and the target stayed unhealthy. It now accepts 200-399. A load balancer routes to every target when all of them are unhealthy, so this failed silently rather than causing a visible outage. The images bucket name is now a variable, since S3 names are globally unique and a rebuild in another account needs its own. Also corrects the documented key derivation command, which named a file that has never existed, records provider checksums for both platforms so the lock file works on Windows and Linux, and excludes saved plans from git because they embed the variable values they were created with, including the database password. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01G7gAqBoQbfVpoMfBp3iV2u
Rebuilding the platform required knowledge that lived only in people's heads and in scattered per-service notes. docs/PROVISIONING.md now covers the whole path from an empty AWS account to serving traffic: prerequisites, Terraform, reaching the instances through the Instance Connect tunnel, the environment files, copying the compose artifacts, the two derived files the backend needs, starting both stacks, migrations, and verification. It records the failure modes an actual rebuild hit, so they do not have to be rediscovered. The API crash-loops on a fresh database until migrations run, and because the container is restarting, docker exec is unreliable and a one-off container is needed instead. Provider checksum verification fails intermittently on Windows when antivirus rewrites the provider binary. Health checks look fine while a target group is entirely unhealthy, because the load balancer then routes to every target anyway. Saved plan files are as sensitive as state. Linked from the documentation index. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01G7gAqBoQbfVpoMfBp3iV2u
secret/ was ignored wholesale, so the one genuinely useful thing in it, the tunnel helper, could not be shared. ssh.sh and a rewritten README are now tracked; everything else, above all the SSH private keys, stays out. Because a directory-level ignore stops git descending into it, a negation alone would not have worked. The rule is now secret/* with explicit exceptions, followed by a second exclusion of *.pem and *.key so key material cannot be added even if a broader exception is introduced later. ssh.sh no longer pins instance and endpoint identifiers, which changed with every rebuild and were already stale. It resolves them by Name tag at run time, so it keeps working after the infrastructure is replaced. The previous README could not be published as it stood: it named instances that no longer exist and described where administrator credentials live. It has been rewritten around what the folder is, how to connect, and why a tunnel is required at all. The credential note moved to the gitignored docs/ACCESS.md rather than being dropped. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01G7gAqBoQbfVpoMfBp3iV2u
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The Terraform in
terraform/had never been applied. This branch fixes thedefects a first real apply surfaced, documents the whole rebuild, and makes
the instance access helper shareable.
Verified by provisioning the full stack from an empty account and deploying
onto it: 44 resources created, both compose stacks running, all five target
groups healthy, and every hostname serving over HTTPS.
Terraform (
c35a44a)Placement was non-deterministic. Both instances and the Instance Connect
Endpoint used
data.aws_subnets.default.ids[0]. That list comes back in AWSAPI order with no stability guarantee, so the zone could move between runs and
force replacement. Confirmed live: in one account the API returned the subnets
as
1b, 1c, 1a, makingids[0]the wrong zone entirely. A dedicated datasource now resolves by availability zone, and the apply landed both instances
in
subnet-0dc831fafa6ea08bf, the same subnet the original CLI build used.Engine versions were unpinned, so a rebuild would take whatever AWS
defaults to that day. Pinned to Postgres 18 and Redis 7.1 to match production
and keep the database aligned with the Drizzle migrations.
The admin target group never passed its health check. The panel redirects
/to/admin, so the default matcher of200saw a307. Now200-399.Worth noting how this stayed invisible: a load balancer routes to every target
when all targets in a group are unhealthy, so the admin panel kept working
in production on that fallback rather than failing loudly.
The images bucket name is now a variable, since S3 names are globally
unique and a rebuild in another account needs its own.
Also: corrected a documented command that referenced
secret/wcl.pem, a filethat has never existed; recorded provider checksums for Windows and Linux so
the lock file works on both; and excluded saved plans from git, because a
tfplanembeds the variable values it was built with, including the databasepassword.
Documentation (
d994bb3)docs/PROVISIONING.mdcovers the path from an empty AWS account to servingtraffic: prerequisites, Terraform, the Instance Connect tunnel, environment
files, copying artifacts, the two derived files the backend needs, starting
both stacks, migrations, and verification.
It records the failure modes the real rebuild hit rather than leaving them to
be rediscovered: the API crash-loops on a fresh database until migrations run
(and
docker execis unreliable while the container restarts, so a one-offcontainer is needed); provider checksum verification fails intermittently on
Windows when antivirus rewrites the provider binary; and saved plans are as
sensitive as state.
Access helper (
93c0f41)secret/was ignored wholesale, so the tunnel helper could not be shared.ssh.shand a rewritten README are now tracked; the SSH private keys are not.Because a directory-level ignore stops git descending into it, a negation
alone would not work. The rule is
secret/*with explicit exceptions, then asecond exclusion of
*.pemand*.keyso key material cannot be added evenif a broader exception appears later.
ssh.shno longer pins instance and endpoint identifiers. They changed withevery rebuild and were already stale, pointing at resources destroyed in the
teardown. It resolves them by
Nametag at run time instead.The old README could not be published as it stood: it named instances that no
longer exist and described where administrator credentials live. Rewritten
around what the folder is and how to connect; the credential note moved to the
gitignored
docs/ACCESS.mdrather than being dropped.Not included
The temporary test-domain values live only in the gitignored
terraform.tfvars. Documentation continues to referencerbuexam.inthroughout, and
variables.tfstill defaults to it.🤖 Generated with Claude Code
https://claude.ai/code/session_01G7gAqBoQbfVpoMfBp3iV2u