You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Add a guarded development website schema rebuild and redeploy
Status: pending
Tags: operations, infra, security, enhancement, P0
Depends on: #438 accepted and merged at f143a19b7d5ef7b713c7b8139d01619f333961b6; owner authorization in the linked comment
Blocks: restoring the development website after failed Deploy Dev run 36522190231
Reporter context and authorization
The accepted #438 source passed CI but the development migration task failed, before normal service promotion. The owner explicitly authorized nuking/recreating the DTC website development database, while forbidding an AISL reset: owner direction. Apply that authorization only to the dtc_website_dev logical database's public schema as database role website_dev in AWS account 387546586013, region eu-west-1. The app role lacks CREATEDB; the recovery uses its existing website-only DATABASE_URL and bounded schema ownership. The shared RDS instance and dtc_relay_dev logical database, AISL, and every production database remain untouched. The authorization does not mean the reset has occurred.
Normative references
_docs/PROCESS.md and _docs/specs/07-security-privacy-operations.md, especially backup/recovery and bounded failure behavior.
_docs/runbooks/release-deployments.md for the supported controller, exact-SHA CI verdict, receipts, promotion, and deployed smoke. The normal path is .github/workflows/deploy-dev.yml → deploy/deploy_dev.sh → deploy/deploy_website.sh; the older deploy/cli.py path is not the supported controller.
deploy/deployment_targets.py and deploy/task_definitions.py for the reviewed website-development target, task families, and the website-only DATABASE_URL secret.
_docs/ci/change-selective-ci.md and _docs/specs/10-verification-strategy.md for versioned verification.
_docs/runbooks/development-course-content-bootstrap.md and _docs/runbooks/data-ingest.md for separate, authorized repopulation after an empty-schema deploy. A green deployment does not mean prior development content has returned.
Scope and implementation path
Add a manual, default-off, plainly named reset confirmation input to the existing Deploy Dev workflow. Only workflow_dispatch of current main on github.run_attempt == 1, with an exact confirmation phrase for dtc_website_dev, may take the reset branch. Push events, normal manual dispatches, reruns, non-main refs, and production entry points must refuse reset before any AWS or database mutation. The existing deploy-dev concurrency group, exact-source ci-gate verdict, immutable image publish, OIDC development deployer, and success artifact remain in force. The recovery dispatch deploys the new reviewed main source containing this capability and Consolidate homework answer encryption on the shared package #438; it does not replay the failed run's old workflow revision.
Extend only the development branch of the supported controller, using a narrowly named dev reset helper invoked in the website-dev-migration one-off task. No generic SQL argument, caller-selectable database, environment, role, cluster, schema, or arbitrary task family. Keep the production wrapper unable to invoke reset. Use the image's Django/PostgreSQL driver; the image has no psql dependency.
Before draining or SQL, validate the AWS caller account/region and reviewed development target, exact ECS cluster/service/task identities, current main source and dispatch context, website-only migration task secret reference, and an allowlisted recovery receipt destination. Capture redacted prior web/worker task definitions and counts. Require no competing deployment in the serialized workflow.
Set only the reviewed dev website web and worker services to desired count zero, wait for both to reach zero running/pending tasks, and refuse any concurrent dev website writer/migration task before opening the schema transaction. Treat inability to prove quiescence as a stop. Keep Relay and every other service untouched. Because dev currently has web desired 1 and worker desired 0, still verify both rather than assuming the worker cannot write.
In the one-off migration task, inspect the actual connection and require DTC_ENVIRONMENT=development, development Django settings, exact database dtc_website_dev, current role website_dev, schema public owned by website_dev, and the privileges needed to recreate that schema. Refuse a different database, role, owner, schema/search path, missing privilege, or unresolved active website connection before destructive SQL. Execute only the bounded public schema drop/recreate within one transaction, preserving the database and its other schemas; preserve the grants needed for normal migration. No database drop/create, RDS-level action, cross-database SQL, DROP OWNED, or widened role grant. If preflight shows schema recreation is unavailable, stop and return for a reviewed fix before deleting anything.
Continue the existing migration task's migrate --noinput, sync_relay_schedules, and import_mail_templates from the freshly empty schema, then use the existing exact-image web/worker promotion, worker self-check, /api/health/, /health/ready, and homepage gates. Record success only after those gates. A successful deploy proves a functioning new schema, not restored users, courses, registrations, jobs, or content.
Adjust the controller's recovery state explicitly: before any reset attempt, a failed drain/preflight may restore the captured prior service counts only when the helper proves no schema mutation began. From the moment schema reset may begin, any reset, migration, promotion, or smoke failure attempts to stop both exact dev website services and verifies their observed state. The run stays red and uploads a redacted receipt with stage, exact source/image identity, observed service counts, and whether the stop succeeded. If AWS cannot verify the stop, report STOP NOT VERIFIED and require manual service-state recovery before retrying through a fresh normal main dispatch after diagnosis. Never automatically restore incompatible old web/worker images against the reset schema. Never issue a dev-release success artifact on failure. A GitHub rerun cannot repeat the destructive reset.
Update _docs/runbooks/release-deployments.md with the operator trigger, exact target/authorization checklist, one-shot behavior, observed-service failure state, safe retry path, and content-repopulation separation. Do not place secret URLs, tokens, database rows, personal data, or unsanitized SQL/task output in logs, issues, receipts, or artifacts.
Non-goals
No production, AISL, Relay, shared RDS instance, or other logical database reset; no backup deletion or infrastructure recreation.
No automatic reset on a push or ordinary deploy, generalized reset framework, arbitrary SQL console, new database superuser, or broad IAM/DB grant.
No import of production or protected CMP data, automatic content restoration, or change to existing product pages/API behavior.
No bypass of the DTC engineer → independent tester → PM acceptance → merge/push → on-call process. No live reset during implementation or test review.
Acceptance criteria
Only an exact-confirmed first-attempt manual dispatch from current main can select the reset path; push, missing/wrong confirmation, rerun, non-main and production paths are refused before mutation. A normal dev deployment still runs unchanged.
Preflight binds AWS account/region, reviewed dev target, exact website ECS resources, migration secret, actual PostgreSQL database/role/schema owner/search path/privileges; every mismatched or unavailable value stops before destructive SQL. No Relay or production credential/target can satisfy the guard.
Both dev website services and other website writer tasks are provably quiescent before reset. A failed quiescence check performs no schema SQL and has an explicit, receipt-backed service recovery outcome.
Reset is transactional and affects only dtc_website_dev.public; a forced SQL failure leaves that schema intact. A successful reset allows a complete fresh Django migration. The dtc_relay_dev database and unrelated schemas/objects remain unchanged in an isolated PostgreSQL integration test.
Once reset may have begun, every failure attempts to stop both exact dev website services and verifies the observed state. The run remains red and the redacted receipt records the stage, observed counts, and whether stopping succeeded; if AWS cannot verify the stop, it reports STOP NOT VERIFIED and requires manual service-state recovery. No old image is auto-promoted and no successful dev-release record exists. After diagnosis and service-state recovery, a fresh normal dispatch is the retry path.
The explicit recovery path retains the same exact-SHA CI, immutable-image, migration, worker, readiness, and smoke gates as normal deployment. A successful run records only the newly proven release.
Focused guard, transaction/negative-target, drain/failure-stage, and normal-deploy regression tests pass. The versioned verification plan and evidence checker pass with every component classified, exact digests/counts, and the graph-selected browser tier. Screenshots are not_applicable only if the graph confirms no render impact.
The runbook names the exact operator trigger and manual post-reset data/content handling without treating a green empty-schema deploy as restored prior development data.
After implementation is accepted and merged, the delegated operator reviews the exact bounded target and reset command under the existing owner authorization. First observe the ordinary development deployment triggered by the push. If it succeeds, record that the destructive live path was unexercised and close with the green migration/deploy and read-only health evidence. If it still fails and bounded recovery is warranted, dispatch the one-shot reset from current main; on-call alone observes the run and supplies CI/deploy evidence. Close after a green migration/deploy and receipt plus read-only health evidence for the expected target. If the run is red, keep the issue open and document the observed service state, including STOP NOT VERIFIED when applicable. No additional owner approval is required.
Browser scenario
No page render or navigation contract changes. After a successful recovery, an operator can load https://dev.datatalks.club/, which returns the normal page and current release identity; /health/ready reports the rebuilt schema ready. Content absent because the development database was reset is handled by the separate approved bootstrap/runbooks, not by a page fallback.
Repository and operations scenarios
Wrong database, role, schema owner, account, region, task family, or confirmation: fail before schema mutation, without printing credentials.
Website services fail to drain, or a website writer is present: fail without SQL and record the service recovery state.
Reset/migration/smoke fails after reset may have begun: attempt to stop both website services; the red receipt identifies the stage, observed service counts, stop verification outcome, and safe fresh-dispatch path after any required manual service-state recovery, with no old-image rollback or success artifact.
Exact target succeeds: migrate empty website schema, promote the verified image, pass worker self-check and read-only smoke, and publish a dev-proven release record.
Incident follow-up: bounded development service drain
The ordinary Deploy Dev run 36582025304 passed CI and image publish but migration exited 21. The single owner-authorized reset dispatch 36583816114 failed during service drain before any SQL (mutation_may_have_begun=false); no schema reset occurred. The on-call evidence shows the web task entering STOPPING at 14:37:45 UTC, ECS reporting the service steady at 14:37:55, the controller refusing drain at 14:38:05, and that task reaching STOPPED at 14:39:12. The controller erases the underlying exception, so task-state timing is a strong inference, not a captured exception string. Restoration then exposed a separate missing pinned image digest on web task definition :22; dev web remained desired 1/running 0 and returned HTTP 503. The original migration exception remains unknown because the existing investigator role still lacks logs:GetLogEvents on that stream.
The AWS Gate now permits the on-call investigator's bounded ECS reads. Defer the proposed separate read-only diagnostic workflow and helper: its unmerged candidate remains local only and supplies no needed permission or live capability at this point. Do not merge or dispatch it as part of this follow-up. The historical prior source 4e7d218 had a red exact-source CI run, although its old Deploy Dev run was green under the previous workflow without the current exact-source gate; do not replay that deployment as recovery.
Additional implementation scope
In deploy/dev_reset_ecs.py, capture the exact service-owned website web and worker task ARNs before setting desired count to zero, using the validated development service identities. After the existing ECS service-stable wait, give only those captured draining tasks a bounded, explicit wait to reach actual lastStatus=STOPPED, querying only those validated captured service-owned task ARNs for their observed status through the reviewed ECS target. Service stability and desired/running/pending zero alone are insufficient proof that a task has stopped.
After the bounded wait, run the authoritative cluster-wide website writer check. A website migration task, any newly appeared website task, a foreign or unrecognized website writer, incomplete ECS result, inconsistent task identity, or timeout still refuses the reset before schema SQL. Do not force-stop or mutate any foreign task to make the check pass. Preserve the existing exact target, one-shot dispatch, transactional schema guard, receipts, and failure-stage behavior.
Report a safe, specific drain outcome in the redacted receipt or controller result: captured task still stopping at timeout, competing writer, incomplete ECS observation, or another normalized failure. Do not log raw task/environment/secret data. Do not retry the prior reset or silently dispatch a second reset; this code correction can be reviewed and tested independently of a future, separately reviewed live operation.
Update the development deployment runbook to distinguish ECS service stability from task termination and to state that the current dev outage also requires a pullable, schema-compatible web image before service restoration. The old web image digest is missing; desired-count changes alone cannot restore it. The original migration exit 21 still needs separate diagnosis.
Additional acceptance criteria
Focused tests model service counts reaching zero while a captured web task remains STOPPING, then transitions to STOPPED: the drain waits within a fixed deadline and succeeds only after the final authoritative website writer/quiescence check passes.
Focused tests prove timeout, disappeared/incomplete task details, malformed identity, and AWS read failure stop before SQL with a safe categorized outcome. They do not lead to an unbounded wait or force-stop action.
Focused tests prove a competing migration task, newly appeared website writer, or other active website task is still refused by the final cluster-wide check even if the captured service task has stopped. The new captured-task wait queries only validated development website service-owned task ARNs and never mutates other tasks. Retain the existing final authoritative cluster-wide website writer scanner without expanding or narrowing it in this correction.
Normal Deploy Dev and wrong-target/confirmation/attempt guards retain their existing behavior. The versioned verification plan and independent tester/PM gates pass under _docs/PROCESS.md; screenshots are not_applicable only if the impact graph supports that disposition.
On-call records the exact post-push CI verdict. No live schema reset, deployment replay, or image promotion is implied by accepting this source correction; any future cloud recovery has its own reviewed target, command, and evidence.
Operations scenario
A drained ECS service reports zero counts while its old task is still stopping. The controller waits a bounded time for that captured task to stop, then checks all development website writers before SQL. If a task remains active or evidence is incomplete, it fails closed with a redacted, actionable reason. The current dev web outage remains an independent recovery item until a reviewed, pullable image is available.
Add a guarded development website schema rebuild and redeploy
Status: pending
Tags:
operations,infra,security,enhancement,P0Depends on: #438 accepted and merged at
f143a19b7d5ef7b713c7b8139d01619f333961b6; owner authorization in the linked commentBlocks: restoring the development website after failed Deploy Dev run 36522190231
Reporter context and authorization
The accepted #438 source passed CI but the development migration task failed, before normal service promotion. The owner explicitly authorized nuking/recreating the DTC website development database, while forbidding an AISL reset: owner direction. Apply that authorization only to the
dtc_website_devlogical database'spublicschema as database rolewebsite_devin AWS account387546586013, regioneu-west-1. The app role lacksCREATEDB; the recovery uses its existing website-onlyDATABASE_URLand bounded schema ownership. The shared RDS instance anddtc_relay_devlogical database, AISL, and every production database remain untouched. The authorization does not mean the reset has occurred.Normative references
_docs/PROCESS.mdand_docs/specs/07-security-privacy-operations.md, especially backup/recovery and bounded failure behavior._docs/runbooks/release-deployments.mdfor the supported controller, exact-SHA CI verdict, receipts, promotion, and deployed smoke. The normal path is.github/workflows/deploy-dev.yml→deploy/deploy_dev.sh→deploy/deploy_website.sh; the olderdeploy/cli.pypath is not the supported controller.deploy/deployment_targets.pyanddeploy/task_definitions.pyfor the reviewedwebsite-developmenttarget, task families, and the website-onlyDATABASE_URLsecret._docs/ci/change-selective-ci.mdand_docs/specs/10-verification-strategy.mdfor versioned verification._docs/runbooks/development-course-content-bootstrap.mdand_docs/runbooks/data-ingest.mdfor separate, authorized repopulation after an empty-schema deploy. A green deployment does not mean prior development content has returned.Scope and implementation path
Deploy Devworkflow. Onlyworkflow_dispatchof currentmainongithub.run_attempt == 1, with an exact confirmation phrase fordtc_website_dev, may take the reset branch. Push events, normal manual dispatches, reruns, non-main refs, and production entry points must refuse reset before any AWS or database mutation. The existingdeploy-devconcurrency group, exact-sourceci-gateverdict, immutable image publish, OIDC development deployer, and success artifact remain in force. The recovery dispatch deploys the new reviewedmainsource containing this capability and Consolidate homework answer encryption on the shared package #438; it does not replay the failed run's old workflow revision.website-dev-migrationone-off task. No generic SQL argument, caller-selectable database, environment, role, cluster, schema, or arbitrary task family. Keep the production wrapper unable to invoke reset. Use the image's Django/PostgreSQL driver; the image has nopsqldependency.mainsource and dispatch context, website-only migration task secret reference, and an allowlisted recovery receipt destination. Capture redacted prior web/worker task definitions and counts. Require no competing deployment in the serialized workflow.DTC_ENVIRONMENT=development, development Django settings, exact databasedtc_website_dev, current rolewebsite_dev, schemapublicowned bywebsite_dev, and the privileges needed to recreate that schema. Refuse a different database, role, owner, schema/search path, missing privilege, or unresolved active website connection before destructive SQL. Execute only the boundedpublicschema drop/recreate within one transaction, preserving the database and its other schemas; preserve the grants needed for normal migration. No database drop/create, RDS-level action, cross-database SQL,DROP OWNED, or widened role grant. If preflight shows schema recreation is unavailable, stop and return for a reviewed fix before deleting anything.migrate --noinput,sync_relay_schedules, andimport_mail_templatesfrom the freshly empty schema, then use the existing exact-image web/worker promotion, worker self-check,/api/health/,/health/ready, and homepage gates. Record success only after those gates. A successful deploy proves a functioning new schema, not restored users, courses, registrations, jobs, or content.STOP NOT VERIFIEDand require manual service-state recovery before retrying through a fresh normalmaindispatch after diagnosis. Never automatically restore incompatible old web/worker images against the reset schema. Never issue adev-releasesuccess artifact on failure. A GitHub rerun cannot repeat the destructive reset._docs/runbooks/release-deployments.mdwith the operator trigger, exact target/authorization checklist, one-shot behavior, observed-service failure state, safe retry path, and content-repopulation separation. Do not place secret URLs, tokens, database rows, personal data, or unsanitized SQL/task output in logs, issues, receipts, or artifacts.Non-goals
Acceptance criteria
maincan select the reset path; push, missing/wrong confirmation, rerun, non-main and production paths are refused before mutation. A normal dev deployment still runs unchanged.dtc_website_dev.public; a forced SQL failure leaves that schema intact. A successful reset allows a complete fresh Django migration. Thedtc_relay_devdatabase and unrelated schemas/objects remain unchanged in an isolated PostgreSQL integration test.STOP NOT VERIFIEDand requires manual service-state recovery. No old image is auto-promoted and no successful dev-release record exists. After diagnosis and service-state recovery, a fresh normal dispatch is the retry path.not_applicableonly if the graph confirms no render impact.main; on-call alone observes the run and supplies CI/deploy evidence. Close after a green migration/deploy and receipt plus read-only health evidence for the expected target. If the run is red, keep the issue open and document the observed service state, includingSTOP NOT VERIFIEDwhen applicable. No additional owner approval is required.Browser scenario
No page render or navigation contract changes. After a successful recovery, an operator can load
https://dev.datatalks.club/, which returns the normal page and current release identity;/health/readyreports the rebuilt schema ready. Content absent because the development database was reset is handled by the separate approved bootstrap/runbooks, not by a page fallback.Repository and operations scenarios
Incident follow-up: bounded development service drain
The ordinary Deploy Dev run 36582025304 passed CI and image publish but migration exited 21. The single owner-authorized reset dispatch 36583816114 failed during service drain before any SQL (
mutation_may_have_begun=false); no schema reset occurred. The on-call evidence shows the web task enteringSTOPPINGat 14:37:45 UTC, ECS reporting the service steady at 14:37:55, the controller refusing drain at 14:38:05, and that task reachingSTOPPEDat 14:39:12. The controller erases the underlying exception, so task-state timing is a strong inference, not a captured exception string. Restoration then exposed a separate missing pinned image digest on web task definition :22; dev web remained desired 1/running 0 and returned HTTP 503. The original migration exception remains unknown because the existing investigator role still lackslogs:GetLogEventson that stream.The AWS Gate now permits the on-call investigator's bounded ECS reads. Defer the proposed separate read-only diagnostic workflow and helper: its unmerged candidate remains local only and supplies no needed permission or live capability at this point. Do not merge or dispatch it as part of this follow-up. The historical prior source
4e7d218had a red exact-source CI run, although its old Deploy Dev run was green under the previous workflow without the current exact-source gate; do not replay that deployment as recovery.Additional implementation scope
deploy/dev_reset_ecs.py, capture the exact service-owned website web and worker task ARNs before setting desired count to zero, using the validated development service identities. After the existing ECS service-stable wait, give only those captured draining tasks a bounded, explicit wait to reach actuallastStatus=STOPPED, querying only those validated captured service-owned task ARNs for their observed status through the reviewed ECS target. Service stability and desired/running/pending zero alone are insufficient proof that a task has stopped.Additional acceptance criteria
STOPPING, then transitions toSTOPPED: the drain waits within a fixed deadline and succeeds only after the final authoritative website writer/quiescence check passes._docs/PROCESS.md; screenshots arenot_applicableonly if the impact graph supports that disposition.Operations scenario
A drained ECS service reports zero counts while its old task is still stopping. The controller waits a bounded time for that captured task to stop, then checks all development website writers before SQL. If a task remains active or evidence is incomplete, it fails closed with a redacted, actionable reason. The current dev web outage remains an independent recovery item until a reviewed, pullable image is available.