Skip to content

Make Terraform rebuild the production stack, and document the process - #6

Merged
MasterBhuvnesh merged 3 commits into
mainfrom
iac/terraform-provisioning
Aug 18, 2026
Merged

MasterBhuvnesh merged 3 commits into
mainfrom
iac/terraform-provisioning

Conversation

@MasterBhuvnesh

Copy link
Copy Markdown
Owner

The Terraform in terraform/ had never been applied. This branch fixes the
defects a first real apply surfaced, documents the whole rebuild, and makes
the instance access helper shareable.

Verified by provisioning the full stack from an empty account and deploying
onto it: 44 resources created, both compose stacks running, all five target
groups healthy, and every hostname serving over HTTPS.

Terraform (c35a44a)

Placement was non-deterministic. Both instances and the Instance Connect
Endpoint used data.aws_subnets.default.ids[0]. That list comes back in AWS
API order with no stability guarantee, so the zone could move between runs and
force replacement. Confirmed live: in one account the API returned the subnets
as 1b, 1c, 1a, making ids[0] the wrong zone entirely. A dedicated data
source now resolves by availability zone, and the apply landed both instances
in subnet-0dc831fafa6ea08bf, the same subnet the original CLI build used.

Engine versions were unpinned, so a rebuild would take whatever AWS
defaults to that day. Pinned to Postgres 18 and Redis 7.1 to match production
and keep the database aligned with the Drizzle migrations.

The admin target group never passed its health check. The panel redirects
/ to /admin, so the default matcher of 200 saw a 307. Now 200-399.
Worth noting how this stayed invisible: a load balancer routes to every target
when all targets in a group are unhealthy, so the admin panel kept working
in production on that fallback rather than failing loudly.

The images bucket name is now a variable, since S3 names are globally
unique and a rebuild in another account needs its own.

Also: corrected a documented command that referenced secret/wcl.pem, a file
that has never existed; recorded provider checksums for Windows and Linux so
the lock file works on both; and excluded saved plans from git, because a
tfplan embeds the variable values it was built with, including the database
password.

Documentation (d994bb3)

docs/PROVISIONING.md covers the path from an empty AWS account to serving
traffic: prerequisites, Terraform, the Instance Connect tunnel, environment
files, copying artifacts, the two derived files the backend needs, starting
both stacks, migrations, and verification.

It records the failure modes the real rebuild hit rather than leaving them to
be rediscovered: the API crash-loops on a fresh database until migrations run
(and docker exec is unreliable while the container restarts, so a one-off
container is needed); provider checksum verification fails intermittently on
Windows when antivirus rewrites the provider binary; and saved plans are as
sensitive as state.

Access helper (93c0f41)

secret/ was ignored wholesale, so the tunnel helper could not be shared.
ssh.sh and a rewritten README are now tracked; the SSH private keys are not.

Because a directory-level ignore stops git descending into it, a negation
alone would not work. The rule is secret/* with explicit exceptions, then a
second exclusion of *.pem and *.key so key material cannot be added even
if a broader exception appears later.

ssh.sh no longer pins instance and endpoint identifiers. They changed with
every rebuild and were already stale, pointing at resources destroyed in the
teardown. It resolves them by Name tag at run time instead.

The old README could not be published as it stood: it named instances that no
longer exist and described where administrator credentials live. Rewritten
around what the folder is and how to connect; the credential note moved to the
gitignored docs/ACCESS.md rather than being dropped.

Not included

The temporary test-domain values live only in the gitignored
terraform.tfvars. Documentation continues to reference rbuexam.in
throughout, and variables.tf still defaults to it.

🤖 Generated with Claude Code

https://claude.ai/code/session_01G7gAqBoQbfVpoMfBp3iV2u

MasterBhuvnesh and others added 3 commits August 19, 2026 02:19
The configuration had never been applied, and a first run surfaced several
defects. Applying it now rebuilds the stack exactly as the AWS CLI did.

Placement was non-deterministic. Both instances and the Instance Connect
Endpoint selected data.aws_subnets.default.ids[0], but that list comes back
in AWS API order with no stability guarantee, so the zone could change
between runs and force replacement. A dedicated data source now resolves the
subnet by availability zone, which restores the ap-south-1a placement the
original build used.

Engine versions were unpinned, so a rebuild would take whatever AWS defaults
to at the time. Postgres is pinned to 18 and Redis to 7.1 to match what ran
in production, keeping the database aligned with the Drizzle migrations.

The admin target group never passed its health check. The panel redirects the
root path to /admin, so the default matcher of 200 saw a 307 and the target
stayed unhealthy. It now accepts 200-399. A load balancer routes to every
target when all of them are unhealthy, so this failed silently rather than
causing a visible outage.

The images bucket name is now a variable, since S3 names are globally unique
and a rebuild in another account needs its own.

Also corrects the documented key derivation command, which named a file that
has never existed, records provider checksums for both platforms so the lock
file works on Windows and Linux, and excludes saved plans from git because
they embed the variable values they were created with, including the database
password.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01G7gAqBoQbfVpoMfBp3iV2u
Rebuilding the platform required knowledge that lived only in people's heads
and in scattered per-service notes. docs/PROVISIONING.md now covers the whole
path from an empty AWS account to serving traffic: prerequisites, Terraform,
reaching the instances through the Instance Connect tunnel, the environment
files, copying the compose artifacts, the two derived files the backend needs,
starting both stacks, migrations, and verification.

It records the failure modes an actual rebuild hit, so they do not have to be
rediscovered. The API crash-loops on a fresh database until migrations run,
and because the container is restarting, docker exec is unreliable and a
one-off container is needed instead. Provider checksum verification fails
intermittently on Windows when antivirus rewrites the provider binary. Health
checks look fine while a target group is entirely unhealthy, because the load
balancer then routes to every target anyway. Saved plan files are as sensitive
as state.

Linked from the documentation index.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01G7gAqBoQbfVpoMfBp3iV2u
secret/ was ignored wholesale, so the one genuinely useful thing in it, the
tunnel helper, could not be shared. ssh.sh and a rewritten README are now
tracked; everything else, above all the SSH private keys, stays out.

Because a directory-level ignore stops git descending into it, a negation
alone would not have worked. The rule is now secret/* with explicit
exceptions, followed by a second exclusion of *.pem and *.key so key material
cannot be added even if a broader exception is introduced later.

ssh.sh no longer pins instance and endpoint identifiers, which changed with
every rebuild and were already stale. It resolves them by Name tag at run
time, so it keeps working after the infrastructure is replaced.

The previous README could not be published as it stood: it named instances
that no longer exist and described where administrator credentials live. It
has been rewritten around what the folder is, how to connect, and why a
tunnel is required at all. The credential note moved to the gitignored
docs/ACCESS.md rather than being dropped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01G7gAqBoQbfVpoMfBp3iV2u
@MasterBhuvnesh
MasterBhuvnesh merged commit 08c778e into main Aug 18, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant