diff --git a/.gitignore b/.gitignore index 57594b2..a948fe9 100644 --- a/.gitignore +++ b/.gitignore @@ -36,4 +36,9 @@ Thumbs.db # GCP DNS Transactions -*./transaction.yaml \ No newline at end of file +*./transaction.yaml +# Python bytecode +__pycache__/ + +# Terraform plan files written by setup.sh +tfplan diff --git a/aws/README.md b/aws/README.md index 2b8b4a3..1655972 100644 --- a/aws/README.md +++ b/aws/README.md @@ -2,17 +2,45 @@ A realistic network troubleshooting exercise. You're the on-call engineer—diagnose and fix. +```mermaid +flowchart TB + client["Your machine"] + internet["Internet / external APIs"] + igw["Internet gateway"] + dns["Route 53 private hosted zone
internal.test"] + + subgraph vpc["VPC 10.0.0.0/16"] + subgraph public["Public subnet: 10.0.1.0/24
route table: 0.0.0.0/0 → internet gateway"] + bastion["Bastion
Elastic IP + private IP"] + web["Web
Elastic IP + private IP"] + nat["NAT gateway
Elastic IP"] + end + subgraph private["Private subnet: 10.0.2.0/24
route table: 0.0.0.0/0 → NAT gateway"] + api["API
Private IP only"] + end + subgraph database["Database subnet: 10.0.3.0/24
route table: 0.0.0.0/0 → NAT gateway
custom network ACL"] + db["Database
Private IP only"] + end + end + + client -->|"SSH 22"| bastion + client -->|"HTTP 80 / HTTPS 443"| web + bastion -->|"SSH 22 / ICMP"| web + bastion -->|"SSH 22"| api + bastion -->|"SSH 22"| db + web -->|"TCP 8080"| api + api -->|"TCP 5432"| db + api -.->|"default route"| nat + db -.->|"default route"| nat + nat --> igw + igw --> internet + dns -.->|"VPC association"| vpc ``` -┌───────────────────────────────────────────────────────────────┐ -│ VPC (10.0.0.0/16) │ -│ ┌────────────────┐ ┌────────────────┐ ┌────────────────┐ │ -│ │ Public Subnet │ │ Private Subnet │ │ Database │ │ -│ │ 10.0.1.0/24 │ │ 10.0.2.0/24 │ │ Subnet │ │ -│ │ - Bastion │ │ - Web App │ │ 10.0.3.0/24 │ │ -│ │ - NAT GW │ │ - API Server │ │ - Database │ │ -│ └────────────────┘ └────────────────┘ └────────────────┘ │ -└───────────────────────────────────────────────────────────────┘ -``` + +Intended traffic after repairs. Dashed lines show route targets and resource +associations. In AWS an instance's public IP only works when its subnet routes +`0.0.0.0/0` to the internet gateway, so the web server lives in the public +subnet; the API is the only instance in the private subnet. ## Contents @@ -27,10 +55,15 @@ A realistic network troubleshooting exercise. You're the on-call engineer—diag ## Prerequisites -- **AWS CLI** installed and authenticated (e.g., `aws configure`) +- **AWS CLI** installed and authenticated (e.g., `aws configure`); confirm the + account and region with `aws sts get-caller-identity` and `aws configure get region` +- **IAM permissions** to manage VPC networking, EC2, Route 53 private zones, and NAT gateways +- **Bash**, **OpenSSH client**, **Python 3.9+**, **curl**, and **OpenSSL** installed locally - **Terraform** installed (1.4+) - **jq** installed for JSON parsing ([Download jq](https://jqlang.org/download/)) -- **AWS credentials** available to Terraform/CLI (env vars or shared config) + +The default region is `us-east-1`. Set `TF_VAR_aws_region` before running setup +to deploy elsewhere; the scripts read the region from Terraform outputs. --- @@ -51,9 +84,14 @@ A realistic network troubleshooting exercise. You're the on-call engineer—diag ./setup.sh ``` -The setup script will display SSH connection instructions when complete. +Wait for **READY TO START** and the SSH connection instructions. Setup deploys +with a working NAT route so the instances can install packages, waits for +cloud-init and healthy local services on all four instances, and only then +prepares the incidents. If it fails, inspect the error and retry, or run +`./destroy.sh` to avoid charges. -**Cost**: ~$0.50-1.00/session. Destroy when done. +**Cost**: ~$0.50-1.00/session (four `t3.micro` instances, a NAT gateway, and +three public IPv4 addresses). Destroy when done. --- @@ -61,20 +99,26 @@ The setup script will display SSH connection instructions when complete. This lab has **two separate activities**: -- **Diagnose via SSH** — The setup script gives you an SSH command to connect through the bastion host. Use it to hop into VMs and check what's broken (test connectivity, resolve DNS, curl endpoints, etc.). -- **Fix via AWS CLI** — Once you know the root cause, open a separate terminal on your **local machine** and fix the misconfigured cloud resources using `aws` commands (e.g., fix security groups, route tables, DNS records). +- **Diagnose via SSH** — The setup script gives you an SSH command to connect through the bastion host. Use it to hop into instances and check what's broken (test connectivity, resolve DNS, curl endpoints, etc.). +- **Fix via AWS CLI** — Once you know the root cause, open a separate terminal on your **local machine** and fix the misconfigured cloud resources using `aws` commands (e.g., fix security groups, route tables, network ACLs, DNS records). -Do **not** edit Terraform files to fix issues. Do **not** try to fix things from inside the VMs. The cloud infrastructure is what's broken — fix it with the cloud CLI. +Do **not** edit Terraform files to fix issues. Do **not** try to fix things from inside the instances. The cloud infrastructure is what's broken — fix it with the cloud CLI. After fixing, run `./validate.sh` to confirm. +Use `./setup.sh`, not Terraform alone, to prepare the incidents. Rerunning setup +reapplies Terraform (which resets the lab's security groups, network ACL, and +routes, but keeps the instances) and prepares the faults again; Route 53 +records you added remain. It is not a progress checker. Recreate older labs +rather than updating them in place. + --- ## Getting Help If you run into issues (broken instructions, validation failures you can’t explain, or suspected bugs), please open a **GitHub Issue** in this repo: -- [Open an issue](../issues/new/choose) +- [Open an issue](https://github.com/learntocloud/networking-lab/issues/new/choose) - Include: incident ID (e.g., INC-4521), what you tried, and `./validate.sh` output (redact secrets/tokens). --- @@ -83,16 +127,31 @@ If you run into issues (broken instructions, validation failures you can’t exp You're on call. Four tickets just came in. Your job: diagnose and fix. +Replace `<..._IP>` with the addresses from setup. Run each diagnostic on the +machine listed, using `ubuntu` unless you configured a different username. +Setup also prints the IDs of the private route table, NAT gateway, database +network ACL, Route 53 zone, and security groups for use with the AWS CLI. + ### 🎫 INC-4521: API service can't pull external data **Priority:** High **Reported by:** Backend Team **Time:** 09:47 AM -> "Our API service that runs on the private subnet stopped being able to fetch data from external APIs this morning. We didn't change anything on our end. Requests to third-party services just hang and timeout. Internal calls between our services still work fine." +> "Our API service that runs on the private subnet stopped being able to fetch data from external APIs this morning. We didn't change anything on our end. Requests to third-party services just hang and timeout. SSH, public DNS, and local API health checks still work." **Affected system:** API server (private subnet) +**Done when:** The API can reach external HTTPS through the lab's NAT gateway +while remaining private (no public IP). Check the route table that actually +applies to the API's subnet (an explicit association, or the VPC's main route +table), its `0.0.0.0/0` target, and the NAT gateway's state and Elastic IP. + +| Run from | Command | Expected result | +|----------|---------|-----------------| +| API | `curl -4 --noproxy '*' -sS -o /dev/null -w '%{http_code}\n' --max-time 10 https://example.com` | `200` | +| API | `curl -4 --noproxy '*' -fsS --max-time 10 https://api.ipify.org` | The NAT gateway's Elastic IP. | + --- ### 🎫 INC-4522: Service discovery broken @@ -103,7 +162,20 @@ You're on call. Four tickets just came in. Your job: diagnose and fix. > "Our applications can't resolve internal hostnames anymore. We've been using `web.internal.test`, `api.internal.test`, and `db.internal.test` for service discovery but they stopped resolving. Public DNS works fine - we can resolve google.com. This is blocking deployments." -**Affected system:** All VMs +**Affected systems:** Bastion, web, API, and database + +**Done when:** All three service names resolve exclusively to their correct +private IPv4 addresses through the Amazon-provided DNS resolver and the system +resolver on all four instances. Check the Route 53 private hosted zone, its +association with the lab VPC, the VPC's DNS settings, and the zone's records. +Hosts-file-only workarounds do not count. + +| Run from | Command | Expected result | +|----------|---------|-----------------| +| Each instance | `dig +short @169.254.169.253 web.internal.test A` | Only the web instance's private IP. | +| Each instance | `getent ahostsv4 web.internal.test` | The same private IP; repeated rows are normal. | + +Repeat for `api.internal.test` and `db.internal.test`, expecting their respective IPs. --- @@ -113,10 +185,24 @@ You're on call. Four tickets just came in. Your job: diagnose and fix. **Reported by:** Web Team **Time:** 10:32 AM -> "The web frontend suddenly can't connect to the API backend. We're getting connection refused errors on port 8080. The API health endpoint works when we curl localhost on the API server itself, so the service is running. Also, the API team says they can't reach the database on port 5432." +> "The web frontend suddenly can't connect to the API backend. Connections to port 8080 time out. The API health endpoint works when we curl localhost on the API server itself, so the service is running. The API team also reports timeouts reaching the database on port 5432, although PostgreSQL accepts connections locally on the database server." **Affected systems:** Web server → API server, API server → Database +**Done when:** Both application paths respond, not just their TCP ports. + +| Run from | Command | Expected result | +|----------|---------|-----------------| +| Web | `curl --noproxy '*' -fsS --max-time 5 http://:8080/health` | JSON with `"status": "healthy"`. | +| API | `pg_isready -h -p 5432 -U labuser -d labdb -t 3` | `accepting connections` | + +These checks use private IPs, independently of NAT and DNS repairs. The API's +`/db-check` endpoint also needs INC-4522. A packet has to pass the source +security group's **outbound** rules, the destination subnet's network ACL (in +rule-number order, and stateless, so return traffic needs its own allowance), +and the destination security group's **inbound** rules. Security groups are +stateful, so a permitted request's reply is always allowed back. + --- ### 🎫 INC-4524: Security audit findings @@ -127,19 +213,76 @@ You're on call. Four tickets just came in. Your job: diagnose and fix. > "Our quarterly security scan flagged several issues with the network segmentation: > -> 1. SSH is accessible from the internet on all hosts — bastion should only allow SSH from your trusted source IP/CIDR, and internal hosts (web, API, database) should only allow SSH from the bastion security group -> 2. Database accepts connections on port 5432 from too broad a range — it should only accept connections from the API security group -> 3. ICMP is open from anywhere on the web server — it should only be allowed from the bastion security group +> 1. SSH rules are too broad. Restrict bastion SSH to your current public IPv4 address (`/32`), and web, API, and database SSH to the bastion security group. +> 2. Database accepts connections on port 5432 from too broad a range — it should only accept connections from the API security group. +> 3. ICMP is open from anywhere on the web server — it should only be allowed from the bastion security group. > > These need to be tightened up before our compliance review next week." -**Affected systems:** Security groups / NACLs +**Affected systems:** Security groups + +**Done when:** Only approved sources have access, and required traffic still +works. Complete INC-4523 first and preserve its fixes. + +| Traffic | Approved source | +|---------|-----------------| +| Bastion SSH (TCP 22) | Your current public IPv4 `/32`, as seen by the bastion | +| Web, API, database SSH (TCP 22) | Bastion security group | +| Database TCP 5432 | API security group | +| Web ICMP | Bastion security group | + +| Run from | Command | Expected result | +|----------|---------|-----------------| +| Your machine | `ssh -i ~/.ssh/netlab-key ubuntu@` | SSH session opens. | +| Bastion | `ssh ` | SSH works to web, API, and database. | +| Your machine | `curl --noproxy '*' -kI --max-time 5 http:///health https:///health` | HTTP `200` from both endpoints. | +| Bastion | `ping -c 3 -W 2 ` | Echo replies from the web server. | +| Your machine | `nc -zvw3 22` | Connection fails or times out. | +| API | `ping -c 3 -W 2 ` | No echo replies. | +| Bastion | `nc -zvw3 5432` | Connection fails or times out. | + +The trusted `/32` is the client address seen by the bastion over SSH. If it +changes, update the bastion rule with the AWS CLI before validating. Validation +checks this address without echoing it in status output. + +Security groups are stateful, allow-only, and unordered; every group attached +to an interface is additive, and there are no deny rules or priorities. The +validator evaluates the effective policy by address: it unions all matching +rules on **all** groups attached to each instance (including any port range or +all-protocol rule that covers the port), expands a security-group reference to +the interfaces that currently hold that group, resolves prefix lists, and +compares the result with the approved sources above. A `/32` for the bastion's +or API's current private IP is accepted as equivalent, and narrower rules are +fine as long as the required client keeps access. Extra groups, rules, or +referenced-group members that add any other source are rejected. Network ACLs +and outbound rules can block traffic, but they do not replace tight inbound +rules on the destination security group. The HTTPS check uses `-k` for the +lab's self-signed certificate. --- ## Verify Your Fixes -The validation script tests actual connectivity—not just configuration. It SSHs into the VMs and runs the same checks a user would to confirm services are reachable. Sometimes, you may need to wait a minute or two for changes to propogate before validating. +The commands above are spot checks. The validator checks cloud configuration +and live traffic: the API's effective route table and NAT gateway, DNS through +the Amazon resolver and the system resolver on all four instances, application +health over private IPs, and effective security-group sources corroborated by +allowed and denied probes. It supports the lab's single-interface, IPv4-only +topology; dual-stack VPCs, additional interfaces, and cross-account or peered +security-group references are reported as validation errors rather than +silently ignored. + +Allow route, security-group, and DNS changes to propagate before retrying. +Security-group changes do not interrupt connections that are already tracked, +so use fresh connections when testing. Cached DNS answers, including cached +"no such name" answers from before a repair, can persist until their TTL +expires; setup lowers the zone's negative-caching TTL to 60 seconds, but a +change to the zone's VPC association can take several minutes to take effect. +Validation requires SSH and working diagnostic tools on all four instances. +NAT checks depend on `example.com` and `api.ipify.org` being available. + +**Exit codes:** `0` = all resolved, `1` = unresolved incidents, `2` = validation +error. Completion tokens are available only when all four pass. **When to use it:** - After fixing an incident to confirm it's resolved @@ -175,7 +318,13 @@ The validation script tests actual connectivity—not just configuration. It SSH If `./destroy.sh` exits with errors, it is most likely because you created cloud resources while resolving the incidents that are not tracked by Terraform. Terraform cannot delete resources it does not manage, and some AWS resources cannot be deleted while dependent resources still exist. -Read the error message carefully — it will name the resource that is blocking deletion. Delete that resource manually with the AWS CLI, then run `./destroy.sh` again. +The destroy script removes the instances first, deletes security groups in the +lab VPC that Terraform does not manage, destroys the rest, and retries once if +something still blocks deletion. It also works after a partial destroy. Read any +remaining error message carefully — it will name the resource that is blocking +deletion (for example an Elastic IP you allocated, or a route table you +created). Delete that resource manually with the AWS CLI, then run +`./destroy.sh` again. ## Clean Up @@ -191,4 +340,7 @@ When finished, destroy resources to avoid charges: ./destroy.sh ``` -> **Note:** If `terraform destroy` fails, it's likely because you created resources via the AWS CLI (e.g., security group rules, route table entries) that Terraform doesn't know about. Delete those resources manually with `aws` first, then re-run `./destroy.sh`. +3. Confirm nothing is left: check EC2 instances, NAT gateways, Elastic IPs, + security groups, and Route 53 private hosted zones in the region you used. + +> **Note:** If `terraform destroy` fails, it's likely because you created resources via the AWS CLI (e.g., security groups, Elastic IPs, route tables) that Terraform doesn't know about. Delete those resources manually with `aws` first, then re-run `./destroy.sh`. diff --git a/aws/scripts/common.sh b/aws/scripts/common.sh new file mode 100644 index 0000000..3ba72e3 --- /dev/null +++ b/aws/scripts/common.sh @@ -0,0 +1,354 @@ +#!/bin/bash +# Shared helpers for the AWS networking lab setup and validation scripts. +# shellcheck disable=SC2034 # detail/state variables are consumed by the sourcing scripts +# Checks return 0 (pass), 1 (unresolved), or 2 (execution/service error). + +SSH_OPTS=(-o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null + -o ConnectTimeout=10 -o ConnectionAttempts=1 -o BatchMode=yes + -o ControlMaster=no -o ControlPath=none + -o ServerAliveInterval=10 -o ServerAliveCountMax=3 -q) + +EGRESS_DETAIL="" +PORTS_DETAIL="" +HARDENING_DETAIL="" +SERVICE_DETAIL="" + +require_commands() { + local COMMAND + for COMMAND in "$@"; do + if ! command -v "$COMMAND" >/dev/null 2>&1; then + echo "Error: Required command '$COMMAND' is not installed." >&2 + return 2 + fi + done +} + +get_terraform_output() { + terraform -chdir="$TERRAFORM_DIR" output -raw "$1" +} + +load_lab_outputs() { + DEPLOYMENT_ID=$(get_terraform_output deployment_id) || return 2 + REGION=$(get_terraform_output region) || return 2 + VPC_ID=$(get_terraform_output vpc_id) || return 2 + ADMIN_USERNAME=$(get_terraform_output admin_username) || return 2 + BASTION_IP=$(get_terraform_output bastion_public_ip) || return 2 + BASTION_PRIVATE_IP=$(get_terraform_output bastion_private_ip) || return 2 + WEB_IP=$(get_terraform_output web_server_private_ip) || return 2 + WEB_PUBLIC_IP=$(get_terraform_output web_server_public_ip) || return 2 + API_IP=$(get_terraform_output api_server_private_ip) || return 2 + DB_IP=$(get_terraform_output database_server_private_ip) || return 2 + BASTION_INSTANCE_ID=$(get_terraform_output bastion_instance_id) || return 2 + WEB_INSTANCE_ID=$(get_terraform_output web_instance_id) || return 2 + API_INSTANCE_ID=$(get_terraform_output api_instance_id) || return 2 + DB_INSTANCE_ID=$(get_terraform_output database_instance_id) || return 2 + PRIVATE_ROUTE_TABLE_ID=$(get_terraform_output private_route_table_id) || return 2 + NAT_GATEWAY_ID=$(get_terraform_output nat_gateway_id) || return 2 + DNS_ZONE_ID=$(get_terraform_output dns_zone_id) || return 2 + SSH_KEY="$HOME/.ssh/netlab-key" + if [ -z "$DEPLOYMENT_ID" ] || [ -z "$REGION" ] || [ -z "$VPC_ID" ] || + [ -z "$ADMIN_USERNAME" ] || [ -z "$BASTION_IP" ] || [ -z "$BASTION_PRIVATE_IP" ] || + [ -z "$WEB_IP" ] || [ -z "$WEB_PUBLIC_IP" ] || [ -z "$API_IP" ] || [ -z "$DB_IP" ] || + [ -z "$BASTION_INSTANCE_ID" ] || [ -z "$WEB_INSTANCE_ID" ] || + [ -z "$API_INSTANCE_ID" ] || [ -z "$DB_INSTANCE_ID" ] || + [ -z "$PRIVATE_ROUTE_TABLE_ID" ] || [ -z "$NAT_GATEWAY_ID" ] || [ -z "$DNS_ZONE_ID" ]; then + echo "Error: Terraform outputs are incomplete. Finish setup first." >&2 + return 2 + fi +} + +aws_json() { + aws --region "$REGION" --output json "$@" +} + +run_on_vm() { + local TARGET_IP="$1" COMMAND PROXY + local -a HOP=("${SSH_OPTS[@]}" -i "$SSH_KEY") + printf -v COMMAND '%q' "$2" + if [ "$TARGET_IP" != "$BASTION_IP" ]; then + printf -v PROXY '%q ' ssh "${SSH_OPTS[@]}" -i "$SSH_KEY" \ + -W '%h:%p' "$ADMIN_USERNAME@$BASTION_IP" + HOP+=(-o "ProxyCommand=$PROXY") + fi + ssh -n "${HOP[@]}" "$ADMIN_USERNAME@$TARGET_IP" \ + "timeout ${3:-60} bash -o pipefail -c $COMMAND" +} + +check_vm_tools() { + run_on_vm "$1" ' + for TOOL in curl python3 nc dig getent timeout jq sort grep cut ping; do + command -v "$TOOL" >/dev/null || { + echo "Error: Required diagnostic tool $TOOL is missing." >&2 + exit 127 + } + done + ' +} + +# Local TCP connect test from the validation client: 0 open, 1 blocked, 2 error. +tcp_probe_local() { + python3 - "$1" "$2" <<'PYPROBE' +import socket +import sys + +try: + with socket.create_connection((sys.argv[1], int(sys.argv[2])), timeout=5): + sys.exit(0) +except (socket.timeout, ConnectionRefusedError, ConnectionResetError, TimeoutError): + sys.exit(1) +except OSError as error: + # EHOSTUNREACH/ENETUNREACH also mean the path is blocked. + sys.exit(1 if error.errno in (110, 111, 113, 101) else 2) +PYPROBE +} + +probe_api_https() { + local RESPONSE STATUS + if RESPONSE=$(run_on_vm "$API_IP" \ + 'curl -4 --noproxy "*" -fsSL --max-redirs 3 --connect-timeout 5 --max-time 10 --retry 2 --retry-all-errors --retry-delay 2 --retry-max-time 35 -o /dev/null -w "%{http_code}" https://example.com' 50); then + if [[ "$RESPONSE" == 2[0-9][0-9] ]]; then return 0; fi + EGRESS_DETAIL="External HTTPS returned unexpected status $RESPONSE." + return 2 + else + STATUS=$? + fi + EGRESS_DETAIL="API external HTTPS failed (exit $STATUS)." + case "$STATUS" in + 7|28) return 1 ;; + *) EGRESS_DETAIL+=" Check SSH, tools, public DNS, TLS, and the external endpoint."; return 2 ;; + esac +} + +# Resolve the route table that actually applies to a subnet (explicit or main). +effective_route_table() { + local SUBNET_ID="$1" TABLE + TABLE=$(aws_json ec2 describe-route-tables \ + --filters "Name=vpc-id,Values=$VPC_ID" "Name=association.subnet-id,Values=$SUBNET_ID" \ + --query 'RouteTables[0]') || return 2 + if [ "$TABLE" = null ]; then + TABLE=$(aws_json ec2 describe-route-tables \ + --filters "Name=vpc-id,Values=$VPC_ID" "Name=association.main,Values=true" \ + --query 'RouteTables[0]') || return 2 + fi + if [ "$TABLE" = null ]; then return 2; fi + printf '%s' "$TABLE" +} + +check_api_egress() { + local ENI SUBNET_ID TABLE TABLE_ID ROUTE NAT_ID NAT NAT_IPS OUTBOUND_IP + EGRESS_DETAIL="" + if ! check_vm_tools "$API_IP"; then + EGRESS_DETAIL="Cannot reach the API over SSH or diagnostic tools are missing." + return 2 + fi + if ! ENI=$(aws_json ec2 describe-network-interfaces \ + --filters "Name=attachment.instance-id,Values=$API_INSTANCE_ID" \ + --query 'NetworkInterfaces'); then + EGRESS_DETAIL="Cannot read the API network interfaces from AWS." + return 2 + fi + if ! jq -e 'length == 1 and (.[0].Ipv6Addresses | length == 0) and + (.[0].PrivateIpAddresses | length == 1)' <<< "$ENI" >/dev/null; then + EGRESS_DETAIL="Only the lab single-interface, primary-IPv4 API topology is supported." + return 2 + fi + if jq -e '.[0].Association.PublicIp != null' <<< "$ENI" >/dev/null; then + EGRESS_DETAIL="The API must stay private; a public IP is not a NAT gateway repair." + return 1 + fi + SUBNET_ID=$(jq -r '.[0].SubnetId' <<< "$ENI") + if ! TABLE=$(effective_route_table "$SUBNET_ID"); then + EGRESS_DETAIL="Cannot determine the route table that applies to the API subnet." + return 2 + fi + TABLE_ID=$(jq -r '.RouteTableId' <<< "$TABLE") + ROUTE=$(jq -c '[.Routes[] | select(.DestinationCidrBlock == "0.0.0.0/0")] | .[0]' <<< "$TABLE") + if [ "$ROUTE" = null ]; then + EGRESS_DETAIL="Route table $TABLE_ID for the API subnet has no 0.0.0.0/0 route." + return 1 + fi + if [ "$(jq -r '.State' <<< "$ROUTE")" != active ]; then + EGRESS_DETAIL="The API subnet's 0.0.0.0/0 route in $TABLE_ID is not active (blackhole)." + return 1 + fi + NAT_ID=$(jq -r '.NatGatewayId // empty' <<< "$ROUTE") + if [ -z "$NAT_ID" ]; then + EGRESS_DETAIL="The API subnet's 0.0.0.0/0 route in $TABLE_ID does not target a NAT gateway; the API must stay private." + return 1 + fi + if ! NAT=$(aws_json ec2 describe-nat-gateways \ + --filter "Name=nat-gateway-id,Values=$NAT_ID" --query 'NatGateways[0]'); then + EGRESS_DETAIL="Cannot read NAT gateway $NAT_ID from AWS." + return 2 + fi + if [ "$NAT" = null ] || [ "$(jq -r '.VpcId' <<< "$NAT")" != "$VPC_ID" ]; then + EGRESS_DETAIL="NAT gateway $NAT_ID is not in the lab VPC." + return 1 + fi + if [ "$(jq -r '.State' <<< "$NAT")" != available ] || + [ "$(jq -r '.ConnectivityType // "public"' <<< "$NAT")" != public ]; then + EGRESS_DETAIL="NAT gateway $NAT_ID is not an available public NAT gateway." + return 1 + fi + NAT_IPS=$(jq -r '[.NatGatewayAddresses[]? | .PublicIp | select(. != null)] | unique | join(" ")' <<< "$NAT") + if [ -z "$NAT_IPS" ]; then + EGRESS_DETAIL="NAT gateway $NAT_ID has no public IPv4 address." + return 1 + fi + probe_api_https || return $? + if ! OUTBOUND_IP=$(run_on_vm "$API_IP" \ + 'curl -4 --noproxy "*" -fsS --connect-timeout 5 --max-time 10 --retry 2 --retry-all-errors --retry-delay 2 --retry-max-time 35 https://api.ipify.org' 50); then + EGRESS_DETAIL="Could not obtain the API outbound IP from api.ipify.org." + return 2 + fi + if [ -z "$OUTBOUND_IP" ] || [[ " $NAT_IPS " != *" $OUTBOUND_IP "* ]]; then + EGRESS_DETAIL="Observed outbound IP $OUTBOUND_IP is not an address of NAT gateway $NAT_ID." + return 1 + fi + EGRESS_DETAIL="External HTTPS works through NAT gateway $NAT_ID (outbound IP $OUTBOUND_IP)." +} + +probe_api_health() { + local SOURCE="$1" TARGET="$2" RESPONSE STATUS + if RESPONSE=$(run_on_vm "$SOURCE" \ + "curl --noproxy '*' -fsS --connect-timeout 3 --max-time 5 http://$TARGET:8080/health" 15); then + if jq -se 'length == 1 and (.[0] | type == "object" and .status == "healthy")' \ + <<< "$RESPONSE" >/dev/null 2>&1; then + SERVICE_DETAIL="API health on $TARGET:8080 is healthy." + return 0 + fi + SERVICE_DETAIL="The endpoint on $TARGET:8080 did not return healthy API JSON." + return 2 + else + STATUS=$? + fi + SERVICE_DETAIL="$SOURCE cannot verify API health on $TARGET:8080 (exit $STATUS)." + case "$STATUS" in 7|28) return 1 ;; *) return 2 ;; esac +} + +probe_postgres() { + local SOURCE="$1" TARGET="$2" STATUS + if run_on_vm "$SOURCE" "pg_isready -q -h $TARGET -p 5432 -U labuser -d labdb -t 3" 15; then + SERVICE_DETAIL="PostgreSQL on $TARGET:5432 is accepting connections." + return 0 + else + STATUS=$? + fi + SERVICE_DETAIL="$SOURCE could not confirm PostgreSQL readiness on $TARGET:5432 (exit $STATUS)." + if [ "$STATUS" -eq 2 ]; then return 1; fi + SERVICE_DETAIL+=" Check SSH, pg_isready, and database readiness." + return 2 +} + +check_application_paths() { + local ATTEMPT STATUS WEB_DETAIL + WEB_API_STATE=error + API_DB_STATE=error + if ! probe_api_health "$API_IP" 127.0.0.1; then + PORTS_DETAIL="Local API health failed; network rules cannot be assessed. $SERVICE_DETAIL" + return 2 + fi + if ! probe_postgres "$DB_IP" 127.0.0.1; then + PORTS_DETAIL="Local database health failed; network rules cannot be assessed. $SERVICE_DETAIL" + return 2 + fi + if ! run_on_vm "$API_IP" 'command -v pg_isready >/dev/null'; then + PORTS_DETAIL="Cannot run pg_isready on the API; check SSH and postgresql-client." + return 2 + fi + for ATTEMPT in {1..3}; do + if probe_api_health "$WEB_IP" "$API_IP"; then + WEB_API_STATE=resolved + else + STATUS=$? + if [ "$STATUS" -eq 1 ]; then WEB_API_STATE=unresolved; else WEB_API_STATE=error; fi + fi + WEB_DETAIL="$SERVICE_DETAIL" + if probe_postgres "$API_IP" "$DB_IP"; then + API_DB_STATE=resolved + else + STATUS=$? + if [ "$STATUS" -eq 1 ]; then API_DB_STATE=unresolved; else API_DB_STATE=error; fi + fi + PORTS_DETAIL="Web -> API: $WEB_DETAIL API -> database: $SERVICE_DETAIL" + if [ "$WEB_API_STATE" = error ] || [ "$API_DB_STATE" = error ]; then return 2; fi + if [ "$WEB_API_STATE" = resolved ] && [ "$API_DB_STATE" = resolved ]; then return 0; fi + if [ "$ATTEMPT" -lt 3 ]; then sleep 2; fi + done + return 1 +} + +check_hardening() { + local CONNECTION TRUSTED_IP CLIENT_PORT BASTION_PRIVATE SERVER_PORT POLICY STATUS + local SOURCE TARGET PORT SCHEME + if ! CONNECTION=$(run_on_vm "$BASTION_IP" 'printf "%s\n" "$SSH_CONNECTION"'); then + HARDENING_DETAIL="Cannot determine the current SSH client address from the bastion." + return 2 + fi + read -r TRUSTED_IP CLIENT_PORT BASTION_PRIVATE SERVER_PORT <<< "$CONNECTION" + if POLICY=$(python3 "$SCRIPT_DIR/sg-policy.py" --region "$REGION" --vpc-id "$VPC_ID" \ + --bastion "$BASTION_INSTANCE_ID" --web "$WEB_INSTANCE_ID" \ + --api "$API_INSTANCE_ID" --database "$DB_INSTANCE_ID" --trusted-ip "$TRUSTED_IP"); then + : + else + STATUS=$? + HARDENING_DETAIL="${POLICY:-Effective security-group policy could not be assessed; see stderr.}" + if [ "$STATUS" -eq 1 ] && [ -n "$POLICY" ]; then return 1; else return 2; fi + fi + if check_application_paths; then :; else + STATUS=$? + HARDENING_DETAIL="Source policy passed, but application traffic is not healthy. Complete INC-4523 and preserve its fixes. $PORTS_DETAIL" + return "$STATUS" + fi + for SCHEME in http https; do + if curl -4 --noproxy '*' -kfsS --connect-timeout 5 --max-time 10 \ + "$SCHEME://$WEB_PUBLIC_IP/health" >/dev/null; then :; else + STATUS=$? + HARDENING_DETAIL="Required public web $SCHEME access failed (exit $STATUS)." + case "$STATUS" in 7|28) return 1 ;; *) return 2 ;; esac + fi + done + if run_on_vm "$BASTION_IP" "ping -4 -n -c 3 -W 2 $WEB_IP >/dev/null" 15; then :; else + STATUS=$? + HARDENING_DETAIL="Required bastion ICMP to the web server failed (exit $STATUS)." + if [ "$STATUS" -eq 1 ]; then return 1; else return 2; fi + fi + for SOURCE in "$API_IP" "$DB_IP"; do + if run_on_vm "$SOURCE" "ping -4 -n -c 1 -W 2 $WEB_IP >/dev/null" 10; then + HARDENING_DETAIL="Unauthorized ICMP from $SOURCE to the web server still succeeds." + return 1 + else + STATUS=$? + fi + if [ "$STATUS" -ne 1 ]; then + HARDENING_DETAIL="Negative ICMP check could not run (exit $STATUS)." + return 2 + fi + done + # Unauthorized SSH to the web server's public IP from the validation client. + if tcp_probe_local "$WEB_PUBLIC_IP" 22; then + HARDENING_DETAIL="Unauthorized SSH to the web server's public IP still succeeds from the internet." + return 1 + else + STATUS=$? + fi + if [ "$STATUS" -ne 1 ]; then + HARDENING_DETAIL="Negative public SSH check could not run (exit $STATUS)." + return 2 + fi + for POLICY in "$API_IP $WEB_IP 22" "$WEB_IP $API_IP 22" "$WEB_IP $DB_IP 22" \ + "$WEB_IP $BASTION_PRIVATE 22" "$BASTION_IP $DB_IP 5432"; do + read -r SOURCE TARGET PORT <<< "$POLICY" + if run_on_vm "$SOURCE" "nc -zw3 $TARGET $PORT" 10; then + HARDENING_DETAIL="Unauthorized TCP from $SOURCE to $TARGET:$PORT still succeeds." + return 1 + else + STATUS=$? + fi + if [ "$STATUS" -ne 1 ]; then + HARDENING_DETAIL="Negative TCP check could not run (exit $STATUS)." + return 2 + fi + done + HARDENING_DETAIL="Effective security-group sources and live allowed/denied traffic passed (bastion SSH limited to the current client's public IPv4 /32)." +} diff --git a/aws/scripts/destroy.sh b/aws/scripts/destroy.sh index 55151cf..0e3713b 100755 --- a/aws/scripts/destroy.sh +++ b/aws/scripts/destroy.sh @@ -31,10 +31,30 @@ fi cd "$TERRAFORM_DIR" -# Get VPC for confirmation -VPC_ID=$(terraform output -raw vpc_id 2>/dev/null || echo "unknown") +# Identify the deployment. After a partial destroy the outputs are gone, so fall +# back to the state file, and only accept well-formed values. +state_attribute() { # resource type, attribute + jq -r --arg type "$1" --arg attr "$2" \ + '.resources[]? | select(.type == $type) | .instances[]?.attributes[$attr] // empty' \ + terraform.tfstate 2>/dev/null | head -n 1 +} +VPC_ID=$( (terraform output -raw vpc_id 2>/dev/null || true) | grep -Eo '^vpc-[0-9a-f]+$' || true) +if [ -z "$VPC_ID" ]; then + VPC_ID=$(state_attribute aws_vpc id | grep -Eo '^vpc-[0-9a-f]+$' || true) +fi +REGION=$( (terraform output -raw region 2>/dev/null || true) | grep -Eo '^[a-z]{2}(-[a-z]+)+-[0-9]$' || true) +if [ -z "$REGION" ]; then + REGION=$(state_attribute aws_vpc arn | sed -n 's/^arn:aws:ec2:\([a-z0-9-]*\):.*/\1/p') +fi +if [ -z "$REGION" ]; then + REGION=$(aws configure get region 2>/dev/null || true) +fi +AWS_ARGS=() +if [ -n "$REGION" ]; then + AWS_ARGS=(--region "$REGION") +fi -echo "VPC to destroy: $VPC_ID" +echo "VPC to destroy: ${VPC_ID:-unknown} (region: ${REGION:-default})" echo "" read -p "Are you sure you want to destroy all resources? (yes/N) " -r echo "" @@ -48,9 +68,12 @@ fi echo "Cleaning up dependencies (Route53 records, SG references)..." # Use this deployment's zone, not a name search that could select another lab. -ZONE_ID=$(terraform output -raw dns_zone_id) +ZONE_ID=$( (terraform output -raw dns_zone_id 2>/dev/null || true) | grep -Eo '^(/hostedzone/)?Z[0-9A-Z]+$' || true) +if [ -z "$ZONE_ID" ]; then + ZONE_ID=$(state_attribute aws_route53_zone zone_id | grep -Eo '^Z[0-9A-Z]+$' || true) +fi ZONE_ID="${ZONE_ID#/hostedzone/}" -if [ -n "$ZONE_ID" ] && [ "$ZONE_ID" != "None" ]; then +if [ -n "$ZONE_ID" ]; then for _ in {1..5}; do RECORDS_JSON=$(aws route53 list-resource-record-sets \ --hosted-zone-id "$ZONE_ID" \ @@ -68,9 +91,85 @@ if [ -n "$ZONE_ID" ] && [ "$ZONE_ID" != "None" ]; then done fi +# Security groups Terraform manages (by resource, not by reference: managed +# groups can reference unmanaged ones, so a plain text search is not enough). +managed_security_groups() { + jq -r '.resources[]? | select(.type == "aws_security_group") | .instances[]?.attributes.id // empty' \ + terraform.tfstate 2>/dev/null +} + +unmanaged_security_groups() { + local SG MANAGED + [ -n "$VPC_ID" ] || return 0 + MANAGED=" $(managed_security_groups | tr '\n' ' ') " + for SG in $(aws "${AWS_ARGS[@]}" ec2 describe-security-groups \ + --filters "Name=vpc-id,Values=$VPC_ID" \ + --query "SecurityGroups[?GroupName!='default'].GroupId" --output text 2>/dev/null | tr -d '\r'); do + case "$MANAGED" in + *" $SG "*) ;; + *) echo "$SG" ;; + esac + done +} + +# Remove rules from security groups in the lab VPC that Terraform does not +# manage (created during INC-4523/INC-4524 repairs), so cross-references do not +# block deletion. Terraform revokes rules on its own groups. +sweep_extra_security_groups() { + local SG PERMISSIONS + for SG in $(unmanaged_security_groups); do + echo "Removing rules from unmanaged security group $SG" + PERMISSIONS=$(aws "${AWS_ARGS[@]}" ec2 describe-security-groups --group-ids "$SG" \ + --query 'SecurityGroups[0].IpPermissions' --output json 2>/dev/null || echo "[]") + if [ "$PERMISSIONS" != "[]" ] && [ -n "$PERMISSIONS" ]; then + aws "${AWS_ARGS[@]}" ec2 revoke-security-group-ingress --group-id "$SG" \ + --ip-permissions "$PERMISSIONS" >/dev/null 2>&1 || true + fi + PERMISSIONS=$(aws "${AWS_ARGS[@]}" ec2 describe-security-groups --group-ids "$SG" \ + --query 'SecurityGroups[0].IpPermissionsEgress' --output json 2>/dev/null || echo "[]") + if [ "$PERMISSIONS" != "[]" ] && [ -n "$PERMISSIONS" ]; then + aws "${AWS_ARGS[@]}" ec2 revoke-security-group-egress --group-id "$SG" \ + --ip-permissions "$PERMISSIONS" >/dev/null 2>&1 || true + fi + done +} -echo "Destroying infrastructure..." +delete_extra_security_groups() { + local SG + for SG in $(unmanaged_security_groups); do + echo "Deleting leftover security group $SG" + aws "${AWS_ARGS[@]}" ec2 delete-security-group --group-id "$SG" >/dev/null 2>&1 || true + done +} + +sweep_extra_security_groups + +# Remove the instances and the lab's own security groups first. Terraform +# revokes the managed groups' rules on delete, which also drops any rule that +# references a group created outside Terraform, so those groups can then be +# deleted before the VPC; otherwise an unmanaged group blocks VPC deletion and +# Terraform retries for its full 20-minute timeout before failing. +echo "Destroying instances and lab security groups..." +TARGETS=(-target=module.compute) +while IFS= read -r ADDRESS; do + [ -n "$ADDRESS" ] && TARGETS+=("-target=$ADDRESS") +done < <(terraform state list 2>/dev/null | grep '\.aws_security_group\.' || true) +terraform destroy "${TARGETS[@]}" -auto-approve +delete_extra_security_groups + +echo "Destroying remaining infrastructure..." +set +e terraform destroy -auto-approve +DESTROY_EXIT=$? +set -e + +if [ $DESTROY_EXIT -ne 0 ]; then + echo "" + echo "Terraform destroy failed; removing leftover security groups and retrying once..." + sweep_extra_security_groups + delete_extra_security_groups + terraform destroy -auto-approve +fi # Clean up SSH key if [ -f ~/.ssh/netlab-key ]; then @@ -83,6 +182,8 @@ echo -e "${GREEN}============================================${NC}" echo -e "${GREEN} CLEANUP COMPLETE${NC}" echo -e "${GREEN}============================================${NC}" echo "" -echo "All resources have been destroyed." +echo "All Terraform-managed resources have been destroyed." +echo "If you created extra resources with the AWS CLI (Elastic IPs, routes," +echo "security groups, VPCs), confirm they are gone in the AWS console." echo "Thanks for using the L2C Networking Lab!" echo "" diff --git a/aws/scripts/setup.sh b/aws/scripts/setup.sh index 8f60853..ef5f88a 100755 --- a/aws/scripts/setup.sh +++ b/aws/scripts/setup.sh @@ -1,13 +1,15 @@ #!/bin/bash # ============================================================================= # NETWORKING LAB - AWS SETUP SCRIPT -# Deploys the intentionally broken infrastructure for learning +# Deploys the infrastructure, waits for a healthy baseline, then prepares the +# intentional incidents. # ============================================================================= -set -e +set -euo pipefail SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" TERRAFORM_DIR="${SCRIPT_DIR}/../terraform" +source "${SCRIPT_DIR}/common.sh" # Colors RED='\033[0;31m' @@ -27,6 +29,7 @@ echo "" # ----------------------------------------------------------------------------- echo "Checking prerequisites..." +require_commands jq ssh python3 curl # Check AWS CLI if ! command -v aws &> /dev/null; then @@ -37,13 +40,14 @@ fi echo -e " ${GREEN}✓${NC} AWS CLI found" # Check AWS credentials -if ! aws sts get-caller-identity &> /dev/null; then +if ! IDENTITY=$(aws sts get-caller-identity --output json 2>/dev/null); then echo -e "${RED}Error: AWS credentials not configured.${NC}" echo "Run: aws configure" exit 1 fi -ACCOUNT=$(aws sts get-caller-identity --query Account --output text) -echo -e " ${GREEN}✓${NC} AWS credentials OK (Account: $ACCOUNT)" +ACCOUNT=$(jq -r '.Account' <<< "$IDENTITY") +IDENTITY_ARN=$(jq -r '.Arn' <<< "$IDENTITY") +echo -e " ${GREEN}✓${NC} AWS credentials OK (Account: $ACCOUNT, $IDENTITY_ARN)" # Check Terraform if ! command -v terraform &> /dev/null; then @@ -60,7 +64,8 @@ echo -e " ${GREEN}✓${NC} Terraform found: v$TF_VERSION" echo "" echo "Deploying infrastructure..." -echo -e "${YELLOW}This will create AWS resources that incur costs (~\$0.50-1.00/session).${NC}" +echo -e "${YELLOW}This will create AWS resources that incur costs (~\$0.50-1.00/session):${NC}" +echo -e "${YELLOW}four t3.micro instances, a NAT gateway, and three public IPv4 addresses.${NC}" echo "" read -p "Continue? (y/N) " -n 1 -r echo "" @@ -85,13 +90,21 @@ terraform plan -out=tfplan # Apply echo "" echo "Applying infrastructure..." +report_setup_failure() { + local STATUS=$? + if [ "$STATUS" -ne 0 ]; then + echo "Error: Lab preparation failed; the lab is NOT ready." >&2 + echo "Resources may remain. Retry setup or run $SCRIPT_DIR/destroy.sh to avoid charges." >&2 + fi +} +trap report_setup_failure EXIT terraform apply tfplan # Clean up plan file rm -f tfplan # ----------------------------------------------------------------------------- -# Post-deployment info +# Post-deployment checks and incident preparation # ----------------------------------------------------------------------------- echo "" @@ -102,12 +115,116 @@ echo -e "${GREEN}============================================${NC}" # Save SSH key echo "" echo "Saving SSH key..." -terraform output -raw ssh_private_key > ~/.ssh/netlab-key 2>/dev/null +mkdir -p "$HOME/.ssh" +(umask 077; terraform output -raw ssh_private_key > "$HOME/.ssh/netlab-key") chmod 600 ~/.ssh/netlab-key echo -e " ${GREEN}✓${NC} SSH key saved to ~/.ssh/netlab-key" -# Show deployment region -REGION=$(terraform output -raw region 2>/dev/null) +load_lab_outputs + +# Route 53 negative caching defaults to the SOA minimum TTL (15 minutes), which +# would hide repaired INC-4522 records for a long time. Keep the SOA content and +# lower only the TTLs. +echo "Shortening the private zone's negative-caching TTL..." +SOA=$(aws route53 list-resource-record-sets --hosted-zone-id "$DNS_ZONE_ID" --output json \ + --query "ResourceRecordSets[?Type=='SOA'] | [0]") +SOA_NAME=$(jq -r '.Name' <<< "$SOA") +SOA_VALUE=$(jq -r '.ResourceRecords[0].Value' <<< "$SOA" | awk '{ $7 = 60; print }') +aws route53 change-resource-record-sets --hosted-zone-id "$DNS_ZONE_ID" --output text \ + --query 'ChangeInfo.Status' --change-batch "$(jq -cn --arg name "$SOA_NAME" --arg value "$SOA_VALUE" \ + '{Changes: [{Action: "UPSERT", ResourceRecordSet: {Name: $name, Type: "SOA", TTL: 60, ResourceRecords: [{Value: $value}]}}]}')" + +wait_for_vm() { + local NAME="$1" IP="$2" ATTEMPT SSH_ERROR + echo "Waiting for $NAME SSH and cloud-init..." + for ATTEMPT in {1..36}; do + if SSH_ERROR=$(run_on_vm "$IP" true 10 2>&1); then + break + fi + if [ "$ATTEMPT" -eq 36 ]; then + printf 'Error: %s SSH did not become ready: %s\n' "$NAME" "$SSH_ERROR" >&2 + return 1 + fi + sleep 5 + done + run_on_vm "$IP" ' + cloud-init status --wait >/dev/null + STATUS=$? + if [ "$STATUS" -eq 2 ]; then + echo "Warning: cloud-init reported recoverable errors; checking tools and services." + elif [ "$STATUS" -ne 0 ]; then + echo "Error: cloud-init failed (status $STATUS). Inspect sudo cat /var/log/cloud-init-output.log" >&2 + exit "$STATUS" + fi + test -f /var/lib/netlab-startup-complete || { + echo "Error: The lab initialization script did not finish. Inspect sudo cat /var/log/cloud-init-output.log" >&2 + exit 1 + } + ' 900 + check_vm_tools "$IP" +} + +echo "" +echo "Checking instance initialization before preparing the incidents..." +wait_for_vm bastion "$BASTION_IP" +wait_for_vm web "$WEB_IP" +wait_for_vm API "$API_IP" +wait_for_vm database "$DB_IP" + +run_on_vm "$WEB_IP" 'curl -fsS --max-time 10 http://localhost/health && curl -kfsS --max-time 10 https://localhost/health' +echo "Checking local services and both blocked application paths..." +STATUS=0 +check_application_paths || STATUS=$? +if [ "$STATUS" -ne 1 ] || [ "$WEB_API_STATE" != unresolved ] || [ "$API_DB_STATE" != unresolved ]; then + echo "Error: Expected both application paths blocked with healthy local services. $PORTS_DETAIL" >&2 + exit 1 +fi +echo "$PORTS_DETAIL" + +if ! check_api_egress; then + echo "Error: Healthy bootstrap egress was not established. $EGRESS_DETAIL" >&2 + exit 1 +fi +echo "$EGRESS_DETAIL" + +echo "Preparing the outbound connectivity incident..." +if aws_json ec2 describe-route-tables --route-table-ids "$PRIVATE_ROUTE_TABLE_ID" \ + --query 'RouteTables[0].Routes[?DestinationCidrBlock==`0.0.0.0/0`]' | jq -e 'length > 0' >/dev/null; then + aws --region "$REGION" ec2 delete-route --route-table-id "$PRIVATE_ROUTE_TABLE_ID" \ + --destination-cidr-block 0.0.0.0/0 +fi +FAULT_READY=false +for ATTEMPT in {1..6}; do + if probe_api_https; then + sleep 5 + else + STATUS=$? + if [ "$STATUS" -eq 1 ]; then FAULT_READY=true; break; fi + echo "Error: Could not confirm the initial egress fault. $EGRESS_DETAIL" >&2 + exit 1 + fi +done +if [ "$FAULT_READY" != true ]; then + echo "Error: API external HTTPS still works after preparing the NAT incident." >&2 + exit 1 +fi +if ! probe_api_health "$API_IP" 127.0.0.1; then + echo "Error: API health failed after preparing the NAT incident. $SERVICE_DETAIL" >&2 + exit 1 +fi + +echo "Checking public DNS from every instance..." +for IP in "$BASTION_IP" "$WEB_IP" "$API_IP" "$DB_IP"; do + if ! run_on_vm "$IP" ' + ANSWER=$(dig +time=3 +tries=2 +short @169.254.169.253 google.com A) && + printf "%s\n" "$ANSWER" | grep -Eq "^[0-9]+(\.[0-9]+){3}$" && + timeout 10 getent ahostsv4 google.com >/dev/null + '; then + echo "Error: Public DNS baseline failed on $IP." >&2 + exit 1 + fi +done + echo -e " ${GREEN}✓${NC} Region: $REGION" echo "" @@ -117,10 +234,11 @@ echo -e "${BLUE}============================================${NC}" echo "" echo "Your broken infrastructure is deployed." echo "Work through the tasks in README.md to fix it." +terraform output -raw connection_instructions echo "" echo "Validate your progress anytime with:" -echo " ./scripts/validate.sh" +echo " $SCRIPT_DIR/validate.sh" echo "" echo "When done, clean up with:" -echo " ./scripts/destroy.sh" +echo " $SCRIPT_DIR/destroy.sh" echo "" diff --git a/aws/scripts/sg-policy.py b/aws/scripts/sg-policy.py new file mode 100755 index 0000000..bb3d146 --- /dev/null +++ b/aws/scripts/sg-policy.py @@ -0,0 +1,266 @@ +#!/usr/bin/env python3 +"""Assess effective AWS security-group ingress sources for the lab hardening policy. + +Security groups are stateful, allow-only, unordered, and additive across every +group attached to an interface. This evaluator therefore unions the sources of +all matching rules on all attached groups, expands security-group references to +the private IPs of the interfaces that currently hold those groups, resolves +managed prefix lists, and compares the result with the approved source set. +""" + +import argparse +import ipaddress +import json +import subprocess +import sys + + +class PolicyError(ValueError): + pass + + +ROLES = ("bastion", "web", "api", "database") +PROTOCOL_NUMBERS = {"tcp": {"tcp", "6"}, "icmp": {"icmp", "1"}} + + +def merge(ranges): + result = [] + for low, high in sorted(ranges): + if result and low <= result[-1][1] + 1: + result[-1] = (result[-1][0], max(high, result[-1][1])) + else: + result.append((low, high)) + return result + + +def subtract(ranges, removed): + result = [] + for low, high in ranges: + cursor = low + for start, end in removed: + if end < cursor: + continue + if start > high: + break + if start > cursor: + result.append((cursor, start - 1)) + cursor = max(cursor, end + 1) + if cursor > high: + break + if cursor <= high: + result.append((cursor, high)) + return result + + +def contains(ranges, address): + return any(low <= address <= high for low, high in ranges) + + +def cidr_range(value): + network = ipaddress.ip_network(value, strict=False) + if network.version != 4: + return None + return (int(network.network_address), int(network.broadcast_address)) + + +def aws_json(args, *arguments): + try: + result = subprocess.run( + ["aws", "--region", args.region, "--output", "json", *arguments], + check=True, capture_output=True, text=True, timeout=90, + ) + except subprocess.CalledProcessError as error: + raise PolicyError(f"AWS query failed: {error.stderr.strip()}") from error + except subprocess.TimeoutExpired as error: + raise PolicyError("AWS query timed out") from error + except FileNotFoundError as error: + raise PolicyError("The aws CLI is not installed") from error + try: + return json.loads(result.stdout or "null") + except json.JSONDecodeError as error: + raise PolicyError("AWS returned invalid JSON") from error + + +def describe_hosts(args): + ids = [getattr(args, role) for role in ROLES] + reservations = aws_json(args, "ec2", "describe-instances", "--instance-ids", *ids)["Reservations"] + instances = {vm["InstanceId"]: vm for item in reservations for vm in item["Instances"]} + hosts = {} + for role, instance_id in zip(ROLES, ids): + vm = instances.get(instance_id) + if vm is None: + raise PolicyError(f"Cannot identify the {role} instance {instance_id}") + if vm["State"]["Name"] != "running": + raise PolicyError(f"The {role} instance is not running") + if vm.get("VpcId") != args.vpc_id: + raise PolicyError(f"The {role} instance is not in the lab VPC") + nics = vm.get("NetworkInterfaces", []) + if (len(nics) != 1 or nics[0].get("Ipv6Addresses") or + len(nics[0].get("PrivateIpAddresses", [])) != 1): + raise PolicyError("Only the lab single-interface, primary-IPv4 topology is supported") + hosts[role] = { + "ip": int(ipaddress.IPv4Address(nics[0]["PrivateIpAddress"])), + "groups": [group["GroupId"] for group in nics[0].get("Groups", [])], + } + if not hosts[role]["groups"]: + raise PolicyError(f"The {role} interface has no security group") + return hosts + + +def group_members(args): + """Private IPv4 addresses of every interface in the VPC, keyed by security group.""" + vpc = aws_json(args, "ec2", "describe-vpcs", "--vpc-ids", args.vpc_id)["Vpcs"][0] + if any(item.get("Ipv6CidrBlockState", {}).get("State") not in (None, "disassociated") + for item in vpc.get("Ipv6CidrBlockAssociationSet", [])): + raise PolicyError("Dual-stack VPCs are outside this lab validator's scope") + interfaces = aws_json(args, "ec2", "describe-network-interfaces", + "--filters", f"Name=vpc-id,Values={args.vpc_id}")["NetworkInterfaces"] + members = {} + for interface in interfaces: + if interface.get("Ipv6Addresses"): + raise PolicyError("Dual-stack interfaces are outside this lab validator's scope") + addresses = [cidr_range(item["PrivateIpAddress"]) for item in interface.get("PrivateIpAddresses", [])] + for group in interface.get("Groups", []): + members.setdefault(group["GroupId"], []).extend(addresses) + return {group: merge(addresses) for group, addresses in members.items()} + + +def prefix_list_ranges(args, prefix_list_id, cache): + if prefix_list_id not in cache: + entries = aws_json(args, "ec2", "get-managed-prefix-list-entries", + "--prefix-list-id", prefix_list_id).get("Entries", []) + ranges, ipv6 = [], False + for entry in entries: + value = cidr_range(entry["Cidr"]) + if value is None: + ipv6 = True + else: + ranges.append(value) + cache[prefix_list_id] = (merge(ranges), ipv6) + return cache[prefix_list_id] + + +def rule_sources(args, rule, members, prefix_cache, account): + ranges, ipv6 = [], bool(rule.get("Ipv6Ranges")) + for item in rule.get("IpRanges", []): + value = cidr_range(item["CidrIp"]) + if value is None: + raise PolicyError(f"Invalid IPv4 source {item['CidrIp']}") + ranges.append(value) + for item in rule.get("PrefixListIds", []): + listed, listed_ipv6 = prefix_list_ranges(args, item["PrefixListId"], prefix_cache) + ranges.extend(listed) + ipv6 = ipv6 or listed_ipv6 + for pair in rule.get("UserIdGroupPairs", []): + if (pair.get("VpcPeeringConnectionId") or pair.get("PeeringStatus") or + (pair.get("UserId") and pair["UserId"] != account) or + pair.get("VpcId") not in (None, args.vpc_id)): + raise PolicyError("Cross-account or peered security-group references are outside this lab validator's scope") + group_id = pair.get("GroupId") + if not group_id: + raise PolicyError("Security-group reference without a group ID") + ranges.extend(members.get(group_id, [])) + return merge(ranges), ipv6 + + +def rule_matches(rule, protocol, port): + value = str(rule["IpProtocol"]).lower() + if value == "-1": + return True + if value not in PROTOCOL_NUMBERS[protocol]: + return False + if protocol == "icmp": + return True + low, high = rule.get("FromPort"), rule.get("ToPort") + if type(low) is not int or type(high) is not int or not 0 <= low <= high <= 65535: + raise PolicyError("Invalid security-group port range") + return low <= port <= high + + +def rule_allows_echo(rule): + value = str(rule["IpProtocol"]).lower() + if value == "-1": + return True + if value not in PROTOCOL_NUMBERS["icmp"]: + return False + return rule.get("FromPort", -1) in (-1, 8) and rule.get("ToPort", -1) in (-1, 0) + + +def effective_sources(args, groups, protocol, port, required_filter, members, prefix_cache, account): + """Union of sources allowed on any attached group; also the union for the required traffic.""" + allowed, required_allowed, ipv6 = [], [], False + for group in groups: + if group["VpcId"] != args.vpc_id: + raise PolicyError(f"Security group {group['GroupId']} is not in the lab VPC") + for rule in group.get("IpPermissions", []): + if not rule_matches(rule, protocol, port): + continue + sources, rule_ipv6 = rule_sources(args, rule, members, prefix_cache, account) + allowed.extend(sources) + ipv6 = ipv6 or rule_ipv6 + if required_filter(rule): + required_allowed.extend(sources) + return merge(allowed), merge(required_allowed), ipv6 + + +def assess(args): + trusted = ipaddress.IPv4Address(args.trusted_ip) + if not trusted.is_global: + raise PolicyError("Bastion SSH must originate from the validation client's public IPv4 address") + account = aws_json(args, "sts", "get-caller-identity")["Account"] + hosts = describe_hosts(args) + members = group_members(args) + group_ids = sorted({group for host in hosts.values() for group in host["groups"]}) + described = {group["GroupId"]: group for group in + aws_json(args, "ec2", "describe-security-groups", "--group-ids", *group_ids)["SecurityGroups"]} + if set(described) != set(group_ids): + raise PolicyError("Could not describe every attached security group") + for role, host in hosts.items(): + host["groups"] = [described[group] for group in host["groups"]] + prefix_cache = {} + bastion, api = hosts["bastion"], hosts["api"] + checks = [ + ("bastion", "tcp", 22, [(int(trusted), int(trusted))], int(trusted), "Bastion SSH"), + *[(role, "tcp", 22, [(bastion["ip"], bastion["ip"])], bastion["ip"], f"{role} SSH") + for role in ("web", "api", "database")], + ("database", "tcp", 5432, [(api["ip"], api["ip"])], api["ip"], "Database TCP 5432"), + ("web", "icmp", None, [(bastion["ip"], bastion["ip"])], bastion["ip"], "Web ICMP"), + ] + failures = [] + for role, protocol, port, approved, required, label in checks: + required_filter = rule_allows_echo if protocol == "icmp" else (lambda rule: True) + allowed, required_allowed, ipv6 = effective_sources( + args, hosts[role]["groups"], protocol, port, required_filter, members, prefix_cache, account) + leaked = subtract(allowed, approved) + if leaked: + failures.append(f"{label} permits unauthorized source {ipaddress.IPv4Address(leaked[0][0])} " + "across the attached security groups.") + if ipv6: + failures.append(f"{label} permits IPv6 sources; the lab policy is IPv4-only.") + if not contains(required_allowed, required): + failures.append(f"{label} blocks its required source.") + if failures: + print("Bastion SSH must be restricted to the current client's public IPv4 /32; " + "internal SSH, database, and ICMP sources must resolve to the bastion or API only. " + + " ".join(failures)) + return 1 + print("Effective security-group ingress sources match the policy; bastion SSH is limited to the current client's public IPv4 /32.") + return 0 + + +def main(): + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("--region", required=True) + parser.add_argument("--vpc-id", required=True) + parser.add_argument("--trusted-ip", required=True) + for role in ROLES: + parser.add_argument(f"--{role}", required=True, help=f"{role} EC2 instance ID") + try: + return assess(parser.parse_args()) + except (PolicyError, OSError, KeyError, TypeError, AttributeError, ValueError, IndexError) as error: + print(f"Cannot assess effective security-group policy: {error}", file=sys.stderr) + return 2 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/aws/scripts/validate.sh b/aws/scripts/validate.sh index 201e131..0237b4d 100755 --- a/aws/scripts/validate.sh +++ b/aws/scripts/validate.sh @@ -1,24 +1,24 @@ #!/bin/bash # ============================================================================= -# NETWORKING LAB - AWS VALIDATION SCRIPT +# NETWORKING LAB - VALIDATION SCRIPT (AWS) # Validates incident resolution by testing actual connectivity # ============================================================================= -set -e +set -eo pipefail SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" TERRAFORM_DIR="${SCRIPT_DIR}/../terraform" +source "${SCRIPT_DIR}/common.sh" source "${SCRIPT_DIR}/../../scripts/dns-validation.sh" # Colors for output RED='\033[0;31m' GREEN='\033[0;32m' -YELLOW='\033[1;33m' CYAN='\033[0;36m' NC='\033[0m' -# Incident tracking +# Incident tracking (POSIX-friendly) INC_4521="pending" INC_4522="pending" INC_4523="pending" @@ -27,19 +27,10 @@ INC_4524="pending" # Master secret for token generation (matches verification service) MASTER_SECRET="L2C_CTF_MASTER_2024" -# SSH options for non-interactive use (-n prevents stdin consumption) -SSH_OPTS="-n -o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -o ConnectTimeout=10 -o BatchMode=yes -q" - # ============================================================================= # Helper Functions # ============================================================================= -get_terraform_output() { - cd "$TERRAFORM_DIR" - terraform output -raw "$1" 2>/dev/null || echo "" -} - -# Cross-platform base64 encode (Linux uses -w 0, macOS does not support -w) base64_encode_no_wrap() { if printf "test" | base64 -w 0 >/dev/null 2>&1; then printf '%s' "$1" | base64 -w 0 @@ -48,7 +39,6 @@ base64_encode_no_wrap() { fi } -# Cross-platform base64 decode (Linux uses -d, macOS uses -D) base64_decode_stdin() { if printf "dGVzdA==" | base64 -d >/dev/null 2>&1; then base64 -d @@ -57,7 +47,6 @@ base64_decode_stdin() { fi } -# Cross-platform SHA-256 (Linux has sha256sum, macOS has shasum) sha256_hex() { if command -v sha256sum >/dev/null 2>&1; then printf '%s' "$1" | sha256sum | awk '{print $1}' @@ -68,60 +57,37 @@ sha256_hex() { fi } -# Run a command on a VM via SSH through bastion -run_on_vm() { - local TARGET_IP="$1" - local CMD="$2" - - ssh $SSH_OPTS -i "$SSH_KEY" "$ADMIN_USERNAME@$BASTION_IP" \ - "ssh -o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null $ADMIN_USERNAME@$TARGET_IP '$CMD' 2>/dev/null" 2>/dev/null | tr -d '\n\r' -} - # ============================================================================= -# Pre-flight checks (silent) +# Pre-flight checks # ============================================================================= preflight_check() { + require_commands aws terraform ssh jq python3 curl openssl base64 + if ! aws sts get-caller-identity --output text >/dev/null; then + echo "Error: AWS credentials are not configured. Run 'aws configure'." >&2 + exit 2 + fi # Check if terraform state exists if [ ! -f "${TERRAFORM_DIR}/terraform.tfstate" ]; then echo -e "${RED}Error: No terraform state found. Run './setup.sh' first.${NC}" - exit 1 + exit 2 fi # Check for SSH key if [ ! -f "$HOME/.ssh/netlab-key" ]; then echo -e "${RED}Error: SSH key not found at ~/.ssh/netlab-key${NC}" echo "Run: cd ../terraform && terraform output -raw ssh_private_key > ~/.ssh/netlab-key && chmod 600 ~/.ssh/netlab-key" - exit 1 - fi - - # Get outputs from terraform - DEPLOYMENT_ID=$(get_terraform_output "deployment_id") - BASTION_IP=$(get_terraform_output "bastion_public_ip") - WEB_IP=$(get_terraform_output "web_server_private_ip") - API_IP=$(get_terraform_output "api_server_private_ip") - DB_IP=$(get_terraform_output "database_server_private_ip") - SSH_KEY="$HOME/.ssh/netlab-key" - ADMIN_USERNAME=$(get_terraform_output "admin_username") - - BASTION_SG_ID=$(get_terraform_output "bastion_sg_id") - WEB_SG_ID=$(get_terraform_output "web_sg_id") - API_SG_ID=$(get_terraform_output "api_sg_id") - DB_SG_ID=$(get_terraform_output "db_sg_id") - - if [ -z "$BASTION_IP" ] || [ -z "$ADMIN_USERNAME" ]; then - echo -e "${RED}Error: Could not get terraform outputs. Is the infrastructure deployed?${NC}" - exit 1 - fi - - # Test bastion connectivity - if ! ssh $SSH_OPTS -i "$SSH_KEY" "$ADMIN_USERNAME@$BASTION_IP" "echo ok" >/dev/null 2>&1; then - echo -e "${RED}Error: Cannot reach bastion host${NC}" - exit 1 + exit 2 fi - export DEPLOYMENT_ID BASTION_IP WEB_IP API_IP DB_IP SSH_KEY ADMIN_USERNAME - export BASTION_SG_ID WEB_SG_ID API_SG_ID DB_SG_ID + load_lab_outputs + local IP + for IP in "$BASTION_IP" "$WEB_IP" "$API_IP" "$DB_IP"; do + if ! check_vm_tools "$IP"; then + echo "Error: Cannot run validation on $IP; check SSH and diagnostic tools." >&2 + exit 2 + fi + done } # ============================================================================= @@ -129,124 +95,40 @@ preflight_check() { # ============================================================================= validate_inc_4521() { - local RESULT=$(run_on_vm "$API_IP" "curl -s --max-time 10 -o /dev/null -w '%{http_code}' https://example.com 2>/dev/null || echo 'failed'") - if [ "$RESULT" == "200" ]; then - INC_4521="resolved" - else - INC_4521="unresolved" - fi + local STATUS=0 + check_api_egress || STATUS=$? + record_incident INC_4521 "$STATUS" "$EGRESS_DETAIL" } validate_inc_4522() { - if validate_private_dns "169.254.169.253"; then - INC_4522="resolved" - else - INC_4522="unresolved" - fi + local ATTEMPT STATUS + for ATTEMPT in {1..3}; do + STATUS=0 + validate_private_dns "169.254.169.253" "$BASTION_IP" "$WEB_IP" "$API_IP" "$DB_IP" || STATUS=$? + if [ "$STATUS" -ne 1 ]; then break; fi + if [ "$ATTEMPT" -lt 3 ]; then sleep 2; fi + done + record_incident INC_4522 "$STATUS" "$DNS_DETAIL" } validate_inc_4523() { - local WEB_TO_API=$(run_on_vm "$WEB_IP" "nc -zw3 $API_IP 8080 && echo 1 || echo 0") - local API_TO_DB=$(run_on_vm "$API_IP" "nc -zw3 $DB_IP 5432 && echo 1 || echo 0") - WEB_TO_API=${WEB_TO_API:-0} - API_TO_DB=${API_TO_DB:-0} - - if [ "$WEB_TO_API" -eq 1 ] 2>/dev/null && [ "$API_TO_DB" -eq 1 ] 2>/dev/null; then - INC_4523="resolved" - else - INC_4523="unresolved" - fi + local STATUS=0 + check_application_paths || STATUS=$? + record_incident INC_4523 "$STATUS" "$PORTS_DETAIL" } validate_inc_4524() { - local ALL_PASS=true - - # Check 1: Bastion SSH source restriction (must not be internet-open) - local BASTION_SSH_CIDRS=$(aws ec2 describe-security-groups --group-ids "$BASTION_SG_ID" \ - --query "SecurityGroups[0].IpPermissions[?FromPort==\`22\` && ToPort==\`22\`].IpRanges[].CidrIp" --output text 2>/dev/null) - local BASTION_SSH_SG_SOURCES=$(aws ec2 describe-security-groups --group-ids "$BASTION_SG_ID" \ - --query "SecurityGroups[0].IpPermissions[?FromPort==\`22\` && ToPort==\`22\`].UserIdGroupPairs[].GroupId" --output text 2>/dev/null) - BASTION_SSH_CIDRS=$(printf '%s' "$BASTION_SSH_CIDRS" | tr -d '\r') - BASTION_SSH_SG_SOURCES=$(printf '%s' "$BASTION_SSH_SG_SOURCES" | tr -d '\r') - - if [ -z "$BASTION_SSH_CIDRS" ] || [ -n "$BASTION_SSH_SG_SOURCES" ] || echo "$BASTION_SSH_CIDRS" | grep -q "0.0.0.0/0"; then - ALL_PASS=false - fi - - # Check 2: SSH source restriction on internal tiers (SG-scoped only; no CIDR) - for SG_ID in "$WEB_SG_ID" "$API_SG_ID" "$DB_SG_ID"; do - local SSH_CIDRS=$(aws ec2 describe-security-groups --group-ids "$SG_ID" \ - --query "SecurityGroups[0].IpPermissions[?FromPort==\`22\` && ToPort==\`22\`].IpRanges[].CidrIp" --output text 2>/dev/null) - local SSH_SG_SOURCES=$(aws ec2 describe-security-groups --group-ids "$SG_ID" \ - --query "SecurityGroups[0].IpPermissions[?FromPort==\`22\` && ToPort==\`22\`].UserIdGroupPairs[].GroupId" --output text 2>/dev/null) - SSH_CIDRS=$(printf '%s' "$SSH_CIDRS" | tr -d '\r') - SSH_SG_SOURCES=$(printf '%s' "$SSH_SG_SOURCES" | tr -d '\r') - - if [ -n "$SSH_CIDRS" ] || [ -z "$SSH_SG_SOURCES" ]; then - ALL_PASS=false - fi - - if [ -n "$SSH_SG_SOURCES" ]; then - for SRC_SG in $SSH_SG_SOURCES; do - if [ "$SRC_SG" != "$BASTION_SG_ID" ]; then - ALL_PASS=false - fi - done - fi - done - - # Check 3: Database source restriction (API SG only, TCP/5432 only, SG-scoped) - local DB_CIDRS=$(aws ec2 describe-security-groups --group-ids "$DB_SG_ID" \ - --query "SecurityGroups[0].IpPermissions[?IpProtocol=='tcp' && FromPort==\`5432\` && ToPort==\`5432\`].IpRanges[].CidrIp" --output text 2>/dev/null) - local DB_SG_SOURCES=$(aws ec2 describe-security-groups --group-ids "$DB_SG_ID" \ - --query "SecurityGroups[0].IpPermissions[?IpProtocol=='tcp' && FromPort==\`5432\` && ToPort==\`5432\`].UserIdGroupPairs[].GroupId" --output text 2>/dev/null) - local DB_WIDE_RULES=$(aws ec2 describe-security-groups --group-ids "$DB_SG_ID" \ - --query "SecurityGroups[0].IpPermissions[?IpProtocol=='tcp' && FromPort<=\`5432\` && ToPort>=\`5432\` && (FromPort!=\`5432\` || ToPort!=\`5432\`)]" --output text 2>/dev/null) - DB_CIDRS=$(printf '%s' "$DB_CIDRS" | tr -d '\r') - DB_SG_SOURCES=$(printf '%s' "$DB_SG_SOURCES" | tr -d '\r') - DB_WIDE_RULES=$(printf '%s' "$DB_WIDE_RULES" | tr -d '\r') - - if [ -n "$DB_WIDE_RULES" ]; then - ALL_PASS=false - fi - - if [ -n "$DB_CIDRS" ] || [ -z "$DB_SG_SOURCES" ]; then - ALL_PASS=false - fi - - if [ -n "$DB_SG_SOURCES" ]; then - for SRC_SG in $DB_SG_SOURCES; do - if [ "$SRC_SG" != "$API_SG_ID" ]; then - ALL_PASS=false - fi - done - fi - - # Check 4: ICMP restriction (SG-scoped only; no CIDR) - local ICMP_CIDRS=$(aws ec2 describe-security-groups --group-ids "$WEB_SG_ID" \ - --query "SecurityGroups[0].IpPermissions[?IpProtocol==\`icmp\`].IpRanges[].CidrIp" --output text 2>/dev/null) - local ICMP_SG_SOURCES=$(aws ec2 describe-security-groups --group-ids "$WEB_SG_ID" \ - --query "SecurityGroups[0].IpPermissions[?IpProtocol==\`icmp\`].UserIdGroupPairs[].GroupId" --output text 2>/dev/null) - ICMP_CIDRS=$(printf '%s' "$ICMP_CIDRS" | tr -d '\r') - ICMP_SG_SOURCES=$(printf '%s' "$ICMP_SG_SOURCES" | tr -d '\r') - - if [ -n "$ICMP_CIDRS" ] || [ -z "$ICMP_SG_SOURCES" ]; then - ALL_PASS=false - fi - - if [ -n "$ICMP_SG_SOURCES" ]; then - for SRC_SG in $ICMP_SG_SOURCES; do - if [ "$SRC_SG" != "$BASTION_SG_ID" ]; then - ALL_PASS=false - fi - done - fi + local STATUS=0 + check_hardening || STATUS=$? + record_incident INC_4524 "$STATUS" "$HARDENING_DETAIL" +} - if [ "$ALL_PASS" = true ]; then - INC_4524="resolved" - else - INC_4524="unresolved" - fi +record_incident() { + case "$2" in + 0) printf -v "$1" '%s' resolved ;; + 1) printf -v "$1" '%s' unresolved ;; + *) printf -v "$1" '%s' error; echo "Error: $1: $3" >&2 ;; + esac } # ============================================================================= @@ -256,18 +138,32 @@ validate_inc_4524() { generate_verification_token() { local GITHUB_USER="$1" - local TIMESTAMP=$(date +%s) - local COMPLETION_DATE=$(date -u +"%Y-%m-%d") - local COMPLETION_TIME=$(date -u +"%H:%M:%S") + # Get current timestamp + local TIMESTAMP + local COMPLETION_DATE + local COMPLETION_TIME + + TIMESTAMP=$(date +%s) + COMPLETION_DATE=$(date -u +"%Y-%m-%d") + COMPLETION_TIME=$(date -u +"%H:%M:%S") - local VERIFICATION_SECRET=$(sha256_hex "${MASTER_SECRET}:${DEPLOYMENT_ID}") + # Derive verification secret from master secret + instance ID + local VERIFICATION_SECRET + VERIFICATION_SECRET=$(sha256_hex "${MASTER_SECRET}:${DEPLOYMENT_ID}") - local PAYLOAD='{"github_username":"'"$GITHUB_USER"'","date":"'"$COMPLETION_DATE"'","time":"'"$COMPLETION_TIME"'","timestamp":'"$TIMESTAMP"',"challenge":"networking-lab-aws","challenges":4,"instance_id":"'"$DEPLOYMENT_ID"'"}' + # Create payload as single-line JSON + local PAYLOAD + PAYLOAD='{"github_username":"'"$GITHUB_USER"'","date":"'"$COMPLETION_DATE"'","time":"'"$COMPLETION_TIME"'","timestamp":'"$TIMESTAMP"',"challenge":"networking-lab-aws","challenges":4,"instance_id":"'"$DEPLOYMENT_ID"'"}' - local SIGNATURE=$(echo -n "$PAYLOAD" | openssl dgst -sha256 -hmac "$VERIFICATION_SECRET" | cut -d' ' -f2) + # Generate HMAC-SHA256 signature + local SIGNATURE + SIGNATURE=$(echo -n "$PAYLOAD" | openssl dgst -sha256 -hmac "$VERIFICATION_SECRET" | cut -d' ' -f2) - local TOKEN_DATA='{"payload":'"$PAYLOAD"',"signature":"'"$SIGNATURE"'"}' + # Create final token structure + local TOKEN_DATA + TOKEN_DATA='{"payload":'"$PAYLOAD"',"signature":"'"$SIGNATURE"'"}' + # Base64 encode the token base64_encode_no_wrap "$TOKEN_DATA" } @@ -314,6 +210,10 @@ show_status() { echo "" echo " Resolved: $RESOLVED / $TOTAL" + echo " INC-4521 ($INC_4521): $EGRESS_DETAIL" + echo " INC-4522 ($INC_4522): $DNS_DETAIL" + echo " INC-4523 ($INC_4523): $PORTS_DETAIL" + echo " INC-4524 ($INC_4524): $HARDENING_DETAIL" echo "" if [ $RESOLVED -eq $TOTAL ]; then @@ -325,16 +225,24 @@ show_status() { echo " your completion token." echo "" fi + if [[ " $INC_4521 $INC_4522 $INC_4523 $INC_4524 " == *" error "* ]]; then return 2; fi + [ "$RESOLVED" -eq "$TOTAL" ] } export_token() { - preflight_check > /dev/null 2>&1 - - validate_inc_4521 > /dev/null 2>&1 - validate_inc_4522 > /dev/null 2>&1 - validate_inc_4523 > /dev/null 2>&1 - validate_inc_4524 > /dev/null 2>&1 + preflight_check + + # Run all validations + validate_inc_4521 + validate_inc_4522 + validate_inc_4523 + validate_inc_4524 + if [[ " $INC_4521 $INC_4522 $INC_4523 $INC_4524 " == *" error "* ]]; then + echo "Error: Validation could not complete; no token was generated." >&2 + exit 2 + fi + # Check if all resolved local RESOLVED=0 [ "$INC_4521" == "resolved" ] && RESOLVED=$((RESOLVED + 1)) [ "$INC_4522" == "resolved" ] && RESOLVED=$((RESOLVED + 1)) @@ -352,6 +260,7 @@ export_token() { echo -e "${GREEN}============================================${NC}" echo "" + # Get GitHub username echo "Enter your GitHub username (must match your learntocloud.guide account):" echo -n "> " read GITHUB_USER @@ -365,7 +274,9 @@ export_token() { echo "Generating completion token..." echo "" - local TOKEN=$(generate_verification_token "$GITHUB_USER") + # Generate the token + local TOKEN + TOKEN=$(generate_verification_token "$GITHUB_USER") echo -e "${GREEN}Your completion token:${NC}" echo "" @@ -395,24 +306,36 @@ verify_token() { echo "Verifying token..." echo "" - local DECODED=$(echo "$TOKEN" | base64_decode_stdin 2>/dev/null) + # Decode the token + local DECODED + DECODED=$(echo "$TOKEN" | base64_decode_stdin 2>/dev/null) if [ -z "$DECODED" ]; then echo -e "${RED}Error: Invalid token format.${NC}" exit 1 fi - local PAYLOAD=$(echo "$DECODED" | jq -c '.payload' 2>/dev/null) - local PROVIDED_SIG=$(echo "$DECODED" | jq -r '.signature' 2>/dev/null) - local INSTANCE_ID=$(echo "$DECODED" | jq -r '.payload.instance_id' 2>/dev/null) + # Extract payload as compact JSON + local PAYLOAD + local PROVIDED_SIG + local INSTANCE_ID + + PAYLOAD=$(echo "$DECODED" | jq -c '.payload' 2>/dev/null) + PROVIDED_SIG=$(echo "$DECODED" | jq -r '.signature' 2>/dev/null) + INSTANCE_ID=$(echo "$DECODED" | jq -r '.payload.instance_id' 2>/dev/null) if [ -z "$PAYLOAD" ] || [ "$PAYLOAD" == "null" ] || [ -z "$PROVIDED_SIG" ] || [ -z "$INSTANCE_ID" ]; then echo -e "${RED}Error: Could not parse token.${NC}" exit 1 fi - local VERIFICATION_SECRET=$(sha256_hex "${MASTER_SECRET}:${INSTANCE_ID}") - local EXPECTED_SIG=$(echo -n "$PAYLOAD" | openssl dgst -sha256 -hmac "$VERIFICATION_SECRET" | cut -d' ' -f2) + # Derive verification secret + local VERIFICATION_SECRET + VERIFICATION_SECRET=$(sha256_hex "${MASTER_SECRET}:${INSTANCE_ID}") + + # Regenerate signature over the exact payload string + local EXPECTED_SIG + EXPECTED_SIG=$(echo -n "$PAYLOAD" | openssl dgst -sha256 -hmac "$VERIFICATION_SECRET" | cut -d' ' -f2) if [ "$PROVIDED_SIG" == "$EXPECTED_SIG" ]; then echo -e "${GREEN}✓ Token is VALID${NC}" @@ -476,4 +399,6 @@ main() { echo "" } -main "$@" +if [[ "${BASH_SOURCE[0]}" == "$0" ]]; then + main "$@" +fi diff --git a/aws/terraform/main.tf b/aws/terraform/main.tf index 591cd54..6818ba7 100644 --- a/aws/terraform/main.tf +++ b/aws/terraform/main.tf @@ -24,15 +24,34 @@ resource "random_id" "deployment" { byte_length = 4 } +# Looked up here rather than inside the compute module: the compute module +# depends on the whole network module, and a data source inside it would be +# deferred (and force instance replacement) on every re-run. Existing instances +# ignore later AMI changes (see the compute module's lifecycle blocks). +data "aws_ami" "ubuntu" { + most_recent = true + owners = ["099720109477"] + + filter { + name = "name" + values = ["ubuntu/images/hvm-ssd/ubuntu-jammy-22.04-amd64-server-*"] + } + + filter { + name = "virtualization-type" + values = ["hvm"] + } +} + module "network" { source = "./modules/network" - deployment_id = random_id.deployment.hex - vpc_cidr = var.vpc_cidr - public_subnet_cidr = var.public_subnet_cidr - private_subnet_cidr = var.private_subnet_cidr + deployment_id = random_id.deployment.hex + vpc_cidr = var.vpc_cidr + public_subnet_cidr = var.public_subnet_cidr + private_subnet_cidr = var.private_subnet_cidr database_subnet_cidr = var.database_subnet_cidr - aws_region = var.aws_region + aws_region = var.aws_region } module "dns" { @@ -46,20 +65,25 @@ module "dns" { } module "compute" { + # cloud-init installs packages through the NAT gateway, so wait for the + # complete network module (including the NAT route) before launching. + depends_on = [module.network] + source = "./modules/compute" - deployment_id = random_id.deployment.hex - aws_region = var.aws_region - admin_username = var.admin_username + ami_id = data.aws_ami.ubuntu.id + deployment_id = random_id.deployment.hex + aws_region = var.aws_region + admin_username = var.admin_username - public_subnet_id = module.network.public_subnet_id - private_subnet_id = module.network.private_subnet_id - database_subnet_id = module.network.database_subnet_id + public_subnet_id = module.network.public_subnet_id + private_subnet_id = module.network.private_subnet_id + database_subnet_id = module.network.database_subnet_id - bastion_sg_id = module.network.bastion_sg_id - web_sg_id = module.network.web_sg_id - api_sg_id = module.network.api_sg_id - db_sg_id = module.network.db_sg_id + bastion_sg_id = module.network.bastion_sg_id + web_sg_id = module.network.web_sg_id + api_sg_id = module.network.api_sg_id + db_sg_id = module.network.db_sg_id - vpc_id = module.network.vpc_id + vpc_id = module.network.vpc_id } diff --git a/aws/terraform/modules/compute/main.tf b/aws/terraform/modules/compute/main.tf index 068f750..79243c9 100644 --- a/aws/terraform/modules/compute/main.tf +++ b/aws/terraform/modules/compute/main.tf @@ -10,27 +10,12 @@ resource "aws_key_pair" "lab" { public_key = tls_private_key.ssh.public_key_openssh } -data "aws_ami" "ubuntu" { - most_recent = true - owners = ["099720109477"] - - filter { - name = "name" - values = ["ubuntu/images/hvm-ssd/ubuntu-jammy-22.04-amd64-server-*"] - } - - filter { - name = "virtualization-type" - values = ["hvm"] - } -} - # ============================================================================= # BASTION HOST (Public Subnet) # ============================================================================= resource "aws_instance" "bastion" { - ami = data.aws_ami.ubuntu.id + ami = var.ami_id instance_type = "t3.micro" subnet_id = var.public_subnet_id vpc_security_group_ids = [var.bastion_sg_id] @@ -48,6 +33,11 @@ resource "aws_instance" "bastion" { project = "networking-lab" role = "bastion" } + + lifecycle { + # A newer Ubuntu image must not replace instances when setup is re-run. + ignore_changes = [ami] + } } resource "aws_eip" "bastion" { @@ -65,13 +55,15 @@ resource "aws_eip_association" "bastion" { } # ============================================================================= -# WEB SERVER (Private Subnet + Public IP for HTTP/HTTPS testing) +# WEB SERVER (Public Subnet + Elastic IP for HTTP/HTTPS testing) +# In AWS an instance's public IP only works when its subnet routes 0.0.0.0/0 to +# the internet gateway, so the web server lives in the public subnet. # ============================================================================= resource "aws_instance" "web" { - ami = data.aws_ami.ubuntu.id + ami = var.ami_id instance_type = "t3.micro" - subnet_id = var.private_subnet_id + subnet_id = var.public_subnet_id vpc_security_group_ids = [var.web_sg_id] key_name = aws_key_pair.lab.key_name associate_public_ip_address = true @@ -86,6 +78,11 @@ resource "aws_instance" "web" { project = "networking-lab" role = "web" } + + lifecycle { + # A newer Ubuntu image must not replace instances when setup is re-run. + ignore_changes = [ami] + } } resource "aws_eip" "web" { @@ -107,7 +104,7 @@ resource "aws_eip_association" "web" { # ============================================================================= resource "aws_instance" "api" { - ami = data.aws_ami.ubuntu.id + ami = var.ami_id instance_type = "t3.micro" subnet_id = var.private_subnet_id vpc_security_group_ids = [var.api_sg_id] @@ -124,6 +121,11 @@ resource "aws_instance" "api" { project = "networking-lab" role = "api" } + + lifecycle { + # A newer Ubuntu image must not replace instances when setup is re-run. + ignore_changes = [ami] + } } # ============================================================================= @@ -131,7 +133,7 @@ resource "aws_instance" "api" { # ============================================================================= resource "aws_instance" "database" { - ami = data.aws_ami.ubuntu.id + ami = var.ami_id instance_type = "t3.micro" subnet_id = var.database_subnet_id vpc_security_group_ids = [var.db_sg_id] @@ -148,4 +150,9 @@ resource "aws_instance" "database" { project = "networking-lab" role = "database" } + + lifecycle { + # A newer Ubuntu image must not replace instances when setup is re-run. + ignore_changes = [ami] + } } diff --git a/aws/terraform/modules/compute/outputs.tf b/aws/terraform/modules/compute/outputs.tf index d7ca24a..ce342bd 100644 --- a/aws/terraform/modules/compute/outputs.tf +++ b/aws/terraform/modules/compute/outputs.tf @@ -26,3 +26,19 @@ output "ssh_private_key" { value = tls_private_key.ssh.private_key_pem sensitive = true } + +output "bastion_instance_id" { + value = aws_instance.bastion.id +} + +output "web_instance_id" { + value = aws_instance.web.id +} + +output "api_instance_id" { + value = aws_instance.api.id +} + +output "db_instance_id" { + value = aws_instance.database.id +} diff --git a/aws/terraform/modules/compute/templates/api-init.sh b/aws/terraform/modules/compute/templates/api-init.sh index 9671ad0..44d5da7 100644 --- a/aws/terraform/modules/compute/templates/api-init.sh +++ b/aws/terraform/modules/compute/templates/api-init.sh @@ -1,6 +1,7 @@ #!/bin/bash -# API server initialization script +# API server initialization script (runs via cloud-init user data) set -e +export DEBIAN_FRONTEND=noninteractive admin_username="${admin_username}" ssh_public_key="${ssh_public_key}" @@ -19,17 +20,50 @@ SSHKEY chmod 600 /home/${admin_username}/.ssh/authorized_keys chown -R ${admin_username}:${admin_username} /home/${admin_username}/.ssh -# Create simple API application (offline) +# Package installation needs the private subnet's NAT route, which setup.sh +# removes only after this script has finished (INC-4521). +apt-get -o DPkg::Lock::Timeout=600 update +apt-get -o DPkg::Lock::Timeout=600 install -y \ + python3 \ + python3-flask \ + postgresql-client \ + iputils-ping \ + net-tools \ + dnsutils \ + traceroute \ + netcat-openbsd \ + curl \ + jq \ + vim + mkdir -p /opt/api -cat > /opt/api/index.html << 'HTML' - - Networking Lab API - -

API Server

-

Status: running

- - -HTML +cat > /opt/api/app.py << 'PYAPP' +from flask import Flask, jsonify +import os +import socket + +app = Flask(__name__) + +@app.route('/') +def home(): + return jsonify(service='API Server', status='running', hostname=socket.gethostname()) + +@app.route('/health') +def health(): + return jsonify(status='healthy') + +@app.route('/db-check') +def db_check(): + host = os.environ.get('DB_HOST', 'db.internal.test') + try: + with socket.create_connection((host, 5432), timeout=5): + return jsonify(database='reachable', host=host, port=5432) + except OSError as error: + return jsonify(database='unreachable', host=host, port=5432, error=str(error)), 503 + +if __name__ == '__main__': + app.run(host='0.0.0.0', port=8080) +PYAPP # Create systemd service cat > /etc/systemd/system/api.service << 'SVCFILE' @@ -41,7 +75,8 @@ After=network.target Type=simple User=root WorkingDirectory=/opt/api -ExecStart=/usr/bin/python3 -m http.server 8080 --bind 0.0.0.0 +Environment=DB_HOST=db.internal.test +ExecStart=/usr/bin/python3 /opt/api/app.py Restart=always [Install] @@ -59,17 +94,20 @@ cat > /etc/motd << 'EOF' ============================================================ You are on the API server in the PRIVATE subnet. -This server runs a simple HTTP server on port 8080. +This server runs a Flask API on port 8080. Check API service: sudo systemctl status api - curl http://localhost:8080/ + curl http://localhost:8080/health + curl http://localhost:8080/db-check Test connectivity: - nc -zv db.internal.test 5432 - curl http://localhost:8080/ + curl -sS -o /dev/null -w '%%{http_code}\n' --max-time 10 https://example.com + dig +short @169.254.169.253 db.internal.test A + pg_isready -h -p 5432 -U labuser -d labdb -t 3 ============================================================ EOF echo "API server setup complete" +touch /var/lib/netlab-startup-complete diff --git a/aws/terraform/modules/compute/templates/bastion-init.sh b/aws/terraform/modules/compute/templates/bastion-init.sh index 1291e38..8b0c260 100644 --- a/aws/terraform/modules/compute/templates/bastion-init.sh +++ b/aws/terraform/modules/compute/templates/bastion-init.sh @@ -1,6 +1,7 @@ #!/bin/bash # Bastion host initialization script set -e +export DEBIAN_FRONTEND=noninteractive admin_username="${admin_username}" ssh_public_key="${ssh_public_key}" @@ -12,8 +13,10 @@ if ! id -u "$admin_username" >/dev/null 2>&1; then fi # Install useful networking tools -apt-get update -apt-get install -y \ +apt-get -o DPkg::Lock::Timeout=600 update +apt-get -o DPkg::Lock::Timeout=600 install -y \ + python3 \ + iputils-ping \ net-tools \ dnsutils \ traceroute \ @@ -47,7 +50,7 @@ cat > /etc/motd << 'EOF' You are on the bastion host in the PUBLIC subnet. From here you can SSH to other hosts in the lab: - ssh # Connect to web/api/database servers + ssh # Connect to web/api/database servers (key preinstalled) Useful commands: ip addr # Show network interfaces @@ -63,3 +66,4 @@ Useful commands: EOF echo "Bastion host setup complete" +touch /var/lib/netlab-startup-complete diff --git a/aws/terraform/modules/compute/templates/database-init.sh b/aws/terraform/modules/compute/templates/database-init.sh index c93b67f..80364f7 100644 --- a/aws/terraform/modules/compute/templates/database-init.sh +++ b/aws/terraform/modules/compute/templates/database-init.sh @@ -1,6 +1,7 @@ #!/bin/bash -# Database server initialization script +# Database server initialization script (runs via cloud-init user data) set -e +export DEBIAN_FRONTEND=noninteractive admin_username="${admin_username}" ssh_public_key="${ssh_public_key}" @@ -19,42 +20,34 @@ SSHKEY chmod 600 /home/${admin_username}/.ssh/authorized_keys chown -R ${admin_username}:${admin_username} /home/${admin_username}/.ssh -# Create simple DB listener (offline) -mkdir -p /opt/db -cat > /opt/db/db-listener.py << 'PYAPP' -import socket +# Install PostgreSQL and diagnostic tools (through the database subnet's NAT route) +apt-get -o DPkg::Lock::Timeout=600 update +apt-get -o DPkg::Lock::Timeout=600 install -y \ + python3 \ + iputils-ping \ + postgresql \ + postgresql-contrib \ + net-tools \ + dnsutils \ + traceroute \ + netcat-openbsd \ + curl \ + jq \ + vim -HOST = "0.0.0.0" -PORT = 5432 +# Listen on all interfaces; security groups and the network ACL decide who can connect. +sed -i "s/#listen_addresses = 'localhost'/listen_addresses = '*'/" /etc/postgresql/*/main/postgresql.conf +echo "host all all 10.0.0.0/16 md5" >> /etc/postgresql/*/main/pg_hba.conf -with socket.socket(socket.AF_INET, socket.SOCK_STREAM) as s: - s.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1) - s.bind((HOST, PORT)) - s.listen(5) - while True: - conn, _ = s.accept() - conn.close() -PYAPP +systemctl restart postgresql +systemctl enable postgresql -cat > /etc/systemd/system/db-listener.service << 'SVCFILE' -[Unit] -Description=Networking Lab DB Listener -After=network.target - -[Service] -Type=simple -User=root -WorkingDirectory=/opt/db -ExecStart=/usr/bin/python3 /opt/db/db-listener.py -Restart=always - -[Install] -WantedBy=multi-user.target -SVCFILE - -systemctl daemon-reload -systemctl enable db-listener -systemctl start db-listener +# Create a test database and user +sudo -u postgres psql << 'SQLCMD' +CREATE USER labuser WITH PASSWORD 'labpassword'; +CREATE DATABASE labdb OWNER labuser; +GRANT ALL PRIVILEGES ON DATABASE labdb TO labuser; +SQLCMD # Create MOTD cat > /etc/motd << 'EOF' @@ -63,13 +56,21 @@ cat > /etc/motd << 'EOF' ============================================================ You are on the database server in the DATABASE subnet. -This server runs a simple TCP listener on port 5432. +This server runs PostgreSQL on port 5432. + +Check PostgreSQL: + sudo systemctl status postgresql + pg_isready -h 127.0.0.1 -p 5432 -U labuser -d labdb -t 3 -Check listener: - sudo systemctl status db-listener - nc -zv localhost 5432 +Connection info: + Host: db.internal.test (after DNS is fixed) + Port: 5432 + User: labuser + Password: labpassword + Database: labdb ============================================================ EOF echo "Database server setup complete" +touch /var/lib/netlab-startup-complete diff --git a/aws/terraform/modules/compute/templates/web-init.sh b/aws/terraform/modules/compute/templates/web-init.sh index b0cc4cc..a36d0a1 100644 --- a/aws/terraform/modules/compute/templates/web-init.sh +++ b/aws/terraform/modules/compute/templates/web-init.sh @@ -1,6 +1,7 @@ #!/bin/bash # Web server initialization script set -e +export DEBIAN_FRONTEND=noninteractive admin_username="${admin_username}" ssh_public_key="${ssh_public_key}" @@ -20,8 +21,10 @@ chmod 600 /home/${admin_username}/.ssh/authorized_keys chown -R ${admin_username}:${admin_username} /home/${admin_username}/.ssh # Install nginx and tools -apt-get update -apt-get install -y \ +apt-get -o DPkg::Lock::Timeout=600 update +apt-get -o DPkg::Lock::Timeout=600 install -y \ + python3 \ + iputils-ping \ nginx \ openssl \ net-tools \ @@ -85,7 +88,7 @@ cat > /etc/motd << 'EOF' NETWORKING LAB - WEB SERVER ============================================================ -You are on the web server in the PRIVATE subnet. +You are on the web server in the PUBLIC subnet (it has a public IP). This server runs nginx on ports 80 (HTTP) and 443 (HTTPS). Check nginx: @@ -98,3 +101,4 @@ Check nginx: EOF echo "Web server setup complete" +touch /var/lib/netlab-startup-complete diff --git a/aws/terraform/modules/compute/variables.tf b/aws/terraform/modules/compute/variables.tf index 9d419be..29ed711 100644 --- a/aws/terraform/modules/compute/variables.tf +++ b/aws/terraform/modules/compute/variables.tf @@ -41,3 +41,8 @@ variable "db_sg_id" { variable "vpc_id" { type = string } + +variable "ami_id" { + description = "AMI for all lab instances (looked up in the root module so re-runs do not replace instances)" + type = string +} diff --git a/aws/terraform/modules/dns/main.tf b/aws/terraform/modules/dns/main.tf index 52878b6..dd05cd0 100644 --- a/aws/terraform/modules/dns/main.tf +++ b/aws/terraform/modules/dns/main.tf @@ -1,4 +1,4 @@ -# DNS Module +# Private hosted zone associated with the lab VPC. INC-4522 service records are absent. data "aws_region" "current" {} @@ -14,3 +14,4 @@ resource "aws_route53_zone" "internal" { project = "networking-lab" } } + diff --git a/aws/terraform/modules/dns/records.tf b/aws/terraform/modules/dns/records.tf index 09d85ed..8541661 100644 --- a/aws/terraform/modules/dns/records.tf +++ b/aws/terraform/modules/dns/records.tf @@ -1,2 +1,2 @@ -# DNS Records placeholder -# A records for internal services should be configured here +# INC-4522: A records for web, api, and db are intentionally absent. +# Students create them with the AWS CLI; setup does not add them. diff --git a/aws/terraform/modules/network/main.tf b/aws/terraform/modules/network/main.tf index 33deead..a627026 100644 --- a/aws/terraform/modules/network/main.tf +++ b/aws/terraform/modules/network/main.tf @@ -60,7 +60,10 @@ resource "aws_subnet" "database" { } # ============================================================================= -# NAT GATEWAY (intentionally missing private route table route) +# NAT GATEWAY +# The private route table starts with a working default route so that EC2 +# bootstrap (package installation) succeeds. setup.sh deletes that route after +# initialization to prepare INC-4521. # ============================================================================= resource "aws_eip" "nat" { @@ -105,7 +108,8 @@ resource "aws_route_table" "public" { resource "aws_route_table" "private" { vpc_id = aws_vpc.main.id - # Intentional break: missing 0.0.0.0/0 -> NAT Gateway route + # Routes are managed by aws_route resources so setup.sh can remove the + # default route after bootstrap without Terraform re-adding it silently. tags = { Name = "rt-private-${var.deployment_id}" @@ -113,6 +117,14 @@ resource "aws_route_table" "private" { } } +# INC-4521: setup.sh deletes this route once cloud-init has finished on every +# instance. Rerunning setup recreates it and prepares the fault again. +resource "aws_route" "private_nat" { + route_table_id = aws_route_table.private.id + destination_cidr_block = "0.0.0.0/0" + nat_gateway_id = aws_nat_gateway.main.id +} + resource "aws_route_table" "database" { vpc_id = aws_vpc.main.id @@ -144,7 +156,8 @@ resource "aws_route_table_association" "database" { } # ============================================================================= -# DATABASE NETWORK ACL (intentional deny for 5432) +# DATABASE NETWORK ACL (INC-4523: stateless deny for inbound TCP 5432) +# Rule order matters: the deny at 100 is evaluated before the allow at 200. # ============================================================================= resource "aws_network_acl" "database" { diff --git a/aws/terraform/modules/network/outputs.tf b/aws/terraform/modules/network/outputs.tf index d4d1ae0..11a2558 100644 --- a/aws/terraform/modules/network/outputs.tf +++ b/aws/terraform/modules/network/outputs.tf @@ -29,3 +29,27 @@ output "api_sg_id" { output "db_sg_id" { value = aws_security_group.database.id } + +output "public_route_table_id" { + value = aws_route_table.public.id +} + +output "private_route_table_id" { + value = aws_route_table.private.id +} + +output "database_route_table_id" { + value = aws_route_table.database.id +} + +output "nat_gateway_id" { + value = aws_nat_gateway.main.id +} + +output "internet_gateway_id" { + value = aws_internet_gateway.main.id +} + +output "database_network_acl_id" { + value = aws_network_acl.database.id +} diff --git a/aws/terraform/modules/network/security_groups.tf b/aws/terraform/modules/network/security_groups.tf index cc4e855..90499bf 100644 --- a/aws/terraform/modules/network/security_groups.tf +++ b/aws/terraform/modules/network/security_groups.tf @@ -1,4 +1,8 @@ # Security Groups (intentional misconfigurations for learning) +# +# Security groups are stateful and allow-only: every attached group's rules are +# additive, and there are no deny rules or priorities. INC-4523 and INC-4524 are +# repaired with the AWS CLI, not by editing these definitions. resource "aws_security_group" "bastion" { name = "netlab-bastion-${var.deployment_id}" @@ -12,7 +16,7 @@ resource "aws_security_group" "bastion" { } ingress { - description = "SSH from anywhere (intentional)" + description = "SSH from anywhere (INC-4524)" from_port = 22 to_port = 22 protocol = "tcp" @@ -60,7 +64,7 @@ resource "aws_security_group" "web" { } ingress { - description = "SSH from anywhere (intentional)" + description = "SSH from anywhere (INC-4524)" from_port = 22 to_port = 22 protocol = "tcp" @@ -68,13 +72,14 @@ resource "aws_security_group" "web" { } ingress { - description = "ICMP from anywhere (intentional)" + description = "ICMP from anywhere (INC-4524)" from_port = -1 to_port = -1 protocol = "icmp" cidr_blocks = ["0.0.0.0/0"] } + # INC-4523: no egress to the API on TCP 8080. egress { description = "Allow outbound web only" from_port = 80 @@ -109,14 +114,16 @@ resource "aws_security_group" "api" { } ingress { - description = "SSH from anywhere (intentional)" + description = "SSH from anywhere (INC-4524)" from_port = 22 to_port = 22 protocol = "tcp" cidr_blocks = ["0.0.0.0/0"] } - # Intentional break: missing inbound 8080 from web + # INC-4523: no ingress on TCP 8080 from the web tier, and no egress to the + # database on TCP 5432. Package installation only needs HTTP/HTTPS egress; + # Amazon-provided DNS is not filtered by security groups. egress { description = "Allow outbound web only" @@ -152,7 +159,7 @@ resource "aws_security_group" "database" { } ingress { - description = "SSH from anywhere (intentional)" + description = "SSH from anywhere (INC-4524)" from_port = 22 to_port = 22 protocol = "tcp" @@ -160,7 +167,7 @@ resource "aws_security_group" "database" { } ingress { - description = "Postgres from anywhere (intentional)" + description = "Postgres from anywhere (INC-4524)" from_port = 5432 to_port = 5432 protocol = "tcp" diff --git a/aws/terraform/outputs.tf b/aws/terraform/outputs.tf index f48aa16..deb113c 100644 --- a/aws/terraform/outputs.tf +++ b/aws/terraform/outputs.tf @@ -13,6 +13,36 @@ output "vpc_id" { value = module.network.vpc_id } +output "public_subnet_id" { + description = "Public subnet ID (bastion, web, NAT gateway)" + value = module.network.public_subnet_id +} + +output "private_subnet_id" { + description = "Private subnet ID (API)" + value = module.network.private_subnet_id +} + +output "database_subnet_id" { + description = "Database subnet ID" + value = module.network.database_subnet_id +} + +output "private_route_table_id" { + description = "Route table associated with the private subnet" + value = module.network.private_route_table_id +} + +output "nat_gateway_id" { + description = "NAT gateway in the public subnet" + value = module.network.nat_gateway_id +} + +output "database_network_acl_id" { + description = "Network ACL associated with the database subnet" + value = module.network.database_network_acl_id +} + output "dns_zone_id" { description = "Private DNS zone ID for this lab" value = module.dns.dns_zone_id @@ -23,6 +53,11 @@ output "bastion_public_ip" { value = module.compute.bastion_public_ip } +output "bastion_private_ip" { + description = "Private IP of the bastion host" + value = module.compute.bastion_private_ip +} + output "web_server_public_ip" { description = "Public IP of the web server" value = module.compute.web_public_ip @@ -43,6 +78,26 @@ output "database_server_private_ip" { value = module.compute.db_private_ip } +output "bastion_instance_id" { + description = "Bastion EC2 instance ID" + value = module.compute.bastion_instance_id +} + +output "web_instance_id" { + description = "Web EC2 instance ID" + value = module.compute.web_instance_id +} + +output "api_instance_id" { + description = "API EC2 instance ID" + value = module.compute.api_instance_id +} + +output "database_instance_id" { + description = "Database EC2 instance ID" + value = module.compute.db_instance_id +} + output "ssh_private_key" { description = "SSH private key for VM access" value = module.compute.ssh_private_key @@ -82,22 +137,36 @@ output "connection_instructions" { NETWORKING LAB - CONNECTION INFO (AWS) ============================================ - 1. Save the SSH key (run from aws/terraform directory): + Region: ${var.aws_region} + Deployment ID: ${random_id.deployment.hex} + VPC: ${module.network.vpc_id} + + 1. Save the SSH key (setup.sh already did this): cd aws/terraform terraform output -raw ssh_private_key > ~/.ssh/netlab-key chmod 600 ~/.ssh/netlab-key - 2. Connect to bastion: + 2. Connect to the bastion: ssh -i ~/.ssh/netlab-key ${var.admin_username}@${module.compute.bastion_public_ip} - 3. From bastion, connect to internal hosts: - ssh ${var.admin_username}@${module.compute.web_private_ip} # web server - ssh ${var.admin_username}@${module.compute.api_private_ip} # api server - ssh ${var.admin_username}@${module.compute.db_private_ip} # database server + 3. From the bastion, connect to internal hosts: + ssh ${module.compute.web_private_ip} # web server (public subnet) + ssh ${module.compute.api_private_ip} # API server (private subnet) + ssh ${module.compute.db_private_ip} # database server (database subnet) - 4. Test web endpoint: + 4. Test the public web endpoint from your machine: curl -I http://${module.compute.web_public_ip} + Useful IDs for the AWS CLI: + Private route table: ${module.network.private_route_table_id} + NAT gateway: ${module.network.nat_gateway_id} + Database network ACL: ${module.network.database_network_acl_id} + Route 53 zone: ${module.dns.dns_zone_id} + Security groups: bastion ${module.network.bastion_sg_id} + web ${module.network.web_sg_id} + api ${module.network.api_sg_id} + db ${module.network.db_sg_id} + ============================================ EOT }