Skip to content

QA AWS networking lab for parity with Azure and GCP - #29

Merged
rishabkumar7 merged 2 commits into
mainfrom
aws-qa
Sep 18, 2026
Merged

rishabkumar7 merged 2 commits into
mainfrom
aws-qa

Conversation

@rishabkumar7

@rishabkumar7 rishabkumar7 commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Bring the AWS lab into parity with the four incidents QA'd for Azure (#26) and GCP (#27), using AWS networking semantics. Closes #28. All changes are under aws/; Azure and GCP are unchanged, and the shared scripts/dns-validation.sh is used as-is.

  • Bootstrap with a working private-subnet NAT route, wait for cloud-init and healthy local services on all four instances, then delete the private route table's default route to prepare INC-4521.
  • Verify the API's effective route table (explicit or main), an available public NAT gateway in the lab VPC, and that the observed outbound IP is that gateway's Elastic IP; reject a public IP on the API or an internet-gateway route.
  • Check Route 53 private DNS through the Amazon resolver (169.254.169.253) and the system resolver on all four instances, including the bastion. Shorten the zone's SOA negative TTL to 60s so repaired records become visible quickly.
  • Replace the generic http.server with a Flask health API plus postgresql-client, run real PostgreSQL on the database instance, and validate API health JSON and pg_isready over private IPs.
  • Add sg-policy.py, an effective security-group evaluator: unions all attached groups, port ranges and all-protocol rules, expands security-group references to current member interfaces, resolves managed prefix lists, flags IPv6 sources, and errors on dual-stack or cross-account/peered topologies. Corroborate with live allowed/denied probes, including internet SSH to the web server's public IP.
  • Distinguish unresolved incidents (1) from execution/service errors (2), gate token export on all four passing, and rewrite aws/README.md with the platform-correct topology, diagnostics, expected results, caveats, and cleanup.

Confirmed bugs fixed

  • The private route table had no default route from the outset, so the web/API cloud-init package installs could not run; the web server sat in the private subnet with a public IP that AWS routing could not serve. The web server now lives in the public subnet (see deviations).
  • The API ran python3 -m http.server, so /health returned 404 and nc was the only probe; the database ran a bare TCP listener.
  • The validator matched only FromPort == 22 rules on the Terraform-named groups, so port ranges, all-protocol rules, extra attached groups, prefix lists, and referenced-group membership were not evaluated.
  • A data source inside the compute module was deferred behind the network module's depends_on, which would have replaced all four instances on every setup re-run.
  • INC-4523 ticket text claimed connection-refused errors; dropped packets time out.

Live QA results (us-east-1, deployment 50b686de)

Each case was driven with the AWS CLI as a student would, and corroborated with live probes over SSH (curl, dig/getent, pg_isready, nc, ping) in addition to the validator's own result. Exit codes: 0 resolved, 1 unresolved, 2 validation error.

Area Case Expected Observed
Setup First deployment through setup.sh Healthy bootstrap via NAT, both application paths blocked with healthy local services, egress fault prepared, public DNS on all four instances ✅ exit 0; outbound IP was the NAT gateway's EIP before the route was removed
Setup Baseline validate.sh 0/4, exit 1, per-incident detail ✅
Setup Partial failure (invalid Route 53 change mid-apply) Reports "NOT ready", exit 1, recovery instructions ✅ observed live
Setup Re-run on an existing lab Reapplies Terraform without replacing instances, re-prepares faults ✅ instance IDs unchanged; SGs/NACL/route reset
INC-4521 NAT default route restored resolved; outbound IP = NAT EIP ✅
INC-4521 EIP on the API + IGW route (works live, HTTP 200) unresolved: API must stay private ✅
INC-4521 IGW route without public IP unresolved: route target is not a NAT gateway ✅
INC-4521 Repair removed unresolved: no 0.0.0.0/0 route; SSH and local API health intact ✅
INC-4522 Correct A records with NAT still broken resolved on all four instances (Amazon resolver + system resolver) ✅
INC-4522 Wrong IP / extra A address / missing record unresolved ✅
INC-4522 Wrong /etc/hosts entry on the bastion with correct cloud DNS unresolved ✅
INC-4522 Hosts-file-only workaround with records deleted unresolved ✅
INC-4522 Zone associated only with a temporary QA VPC unresolved; resolved again after re-association (took ~5 min) ✅ temp VPC deleted
INC-4523 Web→API only, then API→DB only unresolved each time ✅
INC-4523 API SG egress 5432 alone with NACL deny still present unresolved ✅
INC-4523 Both repaired with NAT and DNS still broken resolved ✅
INC-4523 Generic http.server on 8080 (404) / API stopped / PostgreSQL stopped / pg_isready removed error (2) ✅
INC-4523 dig removed from the web instance validator preflight exit 2 ✅
INC-4523 Faults restored, reproduced, repaired again unresolved then resolved ✅
INC-4524 SG-reference policy applied resolved; internet SSH to web public IP blocked, bastion ping works ✅
INC-4524 TCP 20-25 from 0.0.0.0/0; all-protocol from 10.0.0.0/8; extra SG on API ENI; prefix list with 0.0.0.0/0; bastion /24 instead of /32; IPv6 ::/0 unresolved, naming the leaked source ✅
INC-4524 Bastion SG also attached to the web ENI (unauthorized member of a referenced group) unresolved ✅
INC-4524 /32 of the bastion's private IP instead of the SG reference resolved (documented as equivalent) ✅
INC-4524 Database 5432 rule removed; web ICMP removed unresolved: required source blocked ✅
INC-4524 Invalid AWS credentials; private trusted IP error (2) ✅
Combined All four resolved together exit 0 ✅
Token Export blocked with one unresolved / with a validation error exit 1 / exit 2, no token ✅
Token Export, verify, altered payload valid; two altered tokens rejected ✅ challenge networking-lab-aws, 4 challenges, correct deployment ID

Platform-specific deviations

  • Web server in the public subnet. Azure and GCP serve an instance-level public IP from the private subnet; AWS only routes a public IP when the subnet's route table points at the internet gateway, and a NAT-routed subnet produces asymmetric replies. The public subnet now holds the bastion, web server, and NAT gateway; the API is the only private-subnet instance. INC-4521 semantics are unchanged.
  • INC-4522 fault is the missing records, not the VPC association. The Route 53 private zone is associated with the lab VPC from creation (AWS cannot remove a private zone's last association); students add the A records. The wrong-association case was exercised in QA with a temporary VPC.
  • INC-4523 has three faults on AWS: the web SG has no 8080 egress and the API SG has no 8080 ingress; the API SG has no 5432 egress and the database subnet's network ACL denies inbound 5432 at rule 100. The ticket text now describes timeouts.
  • INC-4524 policy is evaluated by address. Security-group references are the intended fix, and a /32 of the bastion's or API's current private IP is accepted as equivalent (documented in the README). Network ACLs and outbound rules do not substitute for tight inbound rules on the destination group.
  • Bastion ICMP is required only to the web server (the AWS policy table in QA AWS networking lab for parity with Azure and GCP #28); GCP also requires API/database ICMP.

Tested scope and limitations

  • terraform fmt -check/validate, shellcheck -x on scripts and cloud-init templates, bash -n, python3 -m py_compile. No regression test suite added.
  • Not exercised live: DNS/TLS failure of example.com itself, IAM principals with partial permissions (a bad-credential principal was tested), dual-stack VPCs, cross-account/peered SG references, and regions other than us-east-1. Those topologies produce explicit validation errors.
  • Route 53 negative caching: a repaired record was visible within about a minute after setup lowered the SOA minimum TTL; re-associating the zone with the VPC took about five minutes.
  • Destroying a VPC that still contains an unmanaged security group makes Terraform retry for its full 20-minute timeout. destroy.sh now removes the instances and the lab's own security groups first (revoking rules that reference unmanaged groups), deletes unmanaged groups (identified by state resources, not by text search), then destroys the rest, and resolves IDs from the state file when outputs are gone after a partial destroy. Verified on a third deployment (8bc6cde5) with a student-created group attached to the API and cross-referenced by the API group: one uninterrupted run finished in under three minutes with nothing left behind.
  • Existing instances ignore AMI changes, so a newer Ubuntu image does not replace them on a setup re-run (verified by planning with an older AMI ID).

Cleanup evidence

Three QA deployments (50b686de, 651d8279, 8bc6cde5) were destroyed with destroy.sh, each with an unmanaged, cross-referenced security group planted (the last two also attached to the API interface). After each: no lab instances, volumes, NAT gateways, Elastic IPs, ENIs, security groups, network ACLs, route tables, VPCs, key pairs, or internal.test hosted zones remain; the temporary QA VPC and prefix list were deleted; Terraform state is empty; ~/.ssh/netlab-key was removed (none existed before QA). No completion tokens are included.

🤖 Generated with Claude Code

Bring the AWS lab in line with the Azure (#26) and GCP (#27) QA passes
while keeping AWS networking semantics. All changes are under aws/.

- Bootstrap with a working private-subnet NAT route, wait for cloud-init
  and healthy local services, then remove the route to prepare INC-4521.
- Move the web server to the public subnet: AWS only serves an instance
  public IP when the subnet routes to the internet gateway.
- Replace the generic HTTP listener with a Flask health API plus
  postgresql-client, and run real PostgreSQL on the database instance.
- Add common.sh and sg-policy.py: verify the API's effective route
  table, NAT gateway, and observed outbound IP; check Route 53 private
  DNS through the Amazon and system resolvers on all four instances;
  probe application health over private IPs; evaluate effective
  security-group ingress sources across all attached groups, port
  ranges, references, prefix lists, and IPv6, corroborated by live
  allowed/denied traffic.
- Return 0/1/2 for resolved/unresolved/error, gate token export on all
  four incidents, and lower the zone's SOA negative TTL during setup.
- Look up the AMI in the root module so setup re-runs do not replace
  instances; make destroy remove instances first, delete unmanaged
  security groups, and recover IDs from state after a partial destroy.
- Rewrite aws/README.md with the platform-correct topology, diagnostics,
  expected results, propagation caveats, and cleanup guidance.

Closes #28

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

AMI drift can replace instances, service errors can be misreported, and cross-referenced groups can still delay cleanup.

Get a fresh assessment by requesting another Copilot review.

Pull request overview

Brings the AWS networking lab into parity with the Azure and GCP incident workflows using AWS-native networking behavior.

Changes:

  • Adds reliable bootstrap, real Flask/PostgreSQL services, and AWS-specific topology.
  • Expands validation for NAT, DNS, application paths, and effective security-group policy.
  • Improves setup, cleanup, diagnostics, and student documentation.
File summaries
File Description
.gitignore Ignores Python caches and Terraform plans.
aws/README.md Documents topology, incidents, validation, and cleanup.
aws/scripts/common.sh Adds shared connectivity and policy checks.
aws/scripts/destroy.sh Handles unmanaged resources during cleanup.
aws/scripts/setup.sh Verifies bootstrap before preparing faults.
aws/scripts/sg-policy.py Evaluates effective security-group ingress.
aws/scripts/validate.sh Adds detailed incident validation and exit codes.
aws/terraform/main.tf Moves AMI lookup and sequences compute creation.
aws/terraform/outputs.tf Exposes diagnostic resource and instance IDs.
aws/terraform/modules/compute/main.tf Moves web to public subnet and consumes AMI input.
aws/terraform/modules/compute/outputs.tf Exposes instance IDs.
aws/terraform/modules/compute/variables.tf Adds the AMI input.
aws/terraform/modules/compute/templates/api-init.sh Installs and runs the Flask API.
aws/terraform/modules/compute/templates/bastion-init.sh Adds diagnostics and readiness marker.
aws/terraform/modules/compute/templates/database-init.sh Installs and configures PostgreSQL.
aws/terraform/modules/compute/templates/web-init.sh Updates web bootstrap and subnet guidance.
aws/terraform/modules/dns/main.tf Clarifies private-zone behavior.
aws/terraform/modules/dns/records.tf Documents intentionally absent records.
aws/terraform/modules/network/main.tf Adds bootstrap NAT routing and clarifies NACL behavior.
aws/terraform/modules/network/outputs.tf Exposes routing, gateway, and NACL IDs.
aws/terraform/modules/network/security_groups.tf Documents intentional SG faults.
Review details
  • Files reviewed: 20/21 changed files
  • Comments generated: 3
  • Review effort level: Balanced

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread aws/scripts/common.sh Outdated
Comment thread aws/scripts/destroy.sh
Comment thread aws/terraform/main.tf
# depends on the whole network module, and a data source inside it would be
# deferred (and force instance replacement) on every re-run.
data "aws_ami" "ubuntu" {
most_recent = true
- destroy.sh: destroy the lab's own security groups together with the
  instances in the first targeted pass so their rules referencing
  unmanaged groups are revoked, then delete unmanaged groups before the
  full destroy. Verified live: with a student-created group attached to
  the API and cross-referenced by the API group, an uninterrupted
  destroy completed in under three minutes with nothing left behind.
- compute: ignore AMI changes on existing instances so a newer Ubuntu
  image does not replace them on a setup re-run (verified by planning
  with an older AMI ID: no replacements).
- common.sh: check local API and PostgreSQL health separately so each
  failure names the right service.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

The broad AWS infrastructure, security-policy, and destructive-cleanup changes warrant final human review despite the documented live QA.

Review details
  • Files reviewed: 20/21 changed files
  • Comments generated: 0 new
  • Review effort level: Balanced

@madebygps

Copy link
Copy Markdown
Contributor

just fyi in recent merge I removed skills taht do lab qa and replaced to maintence report.

@rishabkumar7

Copy link
Copy Markdown
Contributor Author

yeah, I will be doing a full end-to-end manual test today

@rishabkumar7

Copy link
Copy Markdown
Contributor Author

Manual verification

I ran the lab by hand from the aws-qa separately from the automated QA recorded in #29:

Environment: us-east-1, deployment afcd7d22, deployed through aws/scripts/setup.sh.

Step Expected Observed
Baseline validate.sh after setup 0/4 resolved, exit 1, per-incident detail ✅ 0/4, exit 1. No default route for the API subnet, names unresolved, both application paths timing out, all six broad sources flagged
Bastion SSH with the saved key Session opens ✅
INC-4521: create-route to the NAT gateway Resolved through NAT ✅ External HTTPS works through nat-0914d35db2494b925, and the observed outbound IP is that gateway's Elastic IP
INC-4522: three A records in the private zone Resolved on all four instances ✅ Amazon resolver and system resolver agree on every instance
INC-4523: web SG egress + API SG ingress on 8080, API SG egress on 5432, NACL deny rule 100 removed Both paths healthy over private IPs ✅ API health JSON and pg_isready both pass
INC-4524: bastion SSH to my /32; web/API/database SSH and web ICMP to the bastion SG; 5432 to the API SG Resolved with live allowed and denied probes ✅
All four together 4/4, exit 0 ✅
validate.sh export Token issued only after 4/4 ✅ challenge networking-lab-aws, 4 challenges, instance ID afcd7d22 (token omitted per the issue)
validate.sh verify on the real token VALID ✅
validate.sh verify on two corrupted tokens Rejected ✅ rejected, see note below
Re-run setup.sh on the live lab Faults re-prepared without replacing instances ✅ Plan was 0 to add, 5 to change, 0 to destroy; the same four instance IDs were kept; security groups and the NACL reset; NAT route removed again
validate.sh after the re-run INC-4521/4523/4524 unresolved, INC-4522 still resolved (records persist, as documented), exit 1 ✅ 1/4, exit 1
destroy.sh Instances and lab security groups first, then the rest ✅ 14 then 17 resources destroyed, "CLEANUP COMPLETE", local SSH key removed
Post-destroy check of the account Nothing left ✅ No lab VPC, instances, NAT gateway, Elastic IPs, security groups, internal.test zone, key pair, or available volumes; Terraform state empty

Note on token verification: a corrupted token (broken base64 or JSON) is rejected with a non-zero exit and is never reported as valid, but the script exits silently after "Verifying token..." instead of printing "Invalid token format". set -eo pipefail aborts on the failed decode before the message. A token with intact encoding and a wrong signature does print "Token is INVALID". The same code is in the GCP validator. Cosmetic, not a security issue; worth a small follow-up.

Not covered in this manual pass (covered by the automated QA in #29): validating after each incident individually, partial INC-4523 repairs, and the rejection cases (internet-gateway route or public IP for INC-4521, wrong or extra DNS answers, hosts-file workarounds, broad or alternate security-group rules).

@rishabkumar7
rishabkumar7 merged commit f20a69d into main Sep 18, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

QA AWS networking lab for parity with Azure and GCP

3 participants