diff --git a/AGENTS.md b/AGENTS.md index 372a1cc3..8cdeb340 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -90,7 +90,7 @@ docker run -p 127.0.0.1:5000:5000 ghcr.io/mlflow/mlflow mlflow server --host 0.0 python tests/test_ml_pipeline.py # Optionally restrict to a specific model type or config -python tests/test_ml_pipeline.py --model NN --config_file experiments/synapse-bella-ip2 +python tests/test_ml_pipeline.py --model NN --config_file experiments/synapse-/config.yaml ``` Dashboard validation is done manually by running the application. @@ -121,8 +121,8 @@ The only CI workflow is **CodeQL Advanced** (`.github/workflows/codeql.yml`), wh - **MongoDB** is used for persistent data from experiments and simulations. - **MLflow** is used for persistent data from ML models. -- Database access requires SSH tunneling to NERSC when running locally. -- Environment variables: `SF_DB_HOST` (dashboard), `SF_DB_READONLY_PASSWORD` (dashboard and ML training), `AM_SC_API_KEY` (dashboard and ML training, required when MLflow tracking_uri is AmSC). +- Database connection settings are read from the `database` section of the experiment's `config.yaml`. Local runs usually reach the database through an SSH tunnel, with `database.host` set to `127.0.0.1` in a local copy of `config.yaml`. +- Environment variables: `config.yaml` sets their names. `database.password_ro_env` names the read-only database password, and `mlflow.api_key_env` names the MLflow API key, which is only needed when `mlflow.tracking_uri` is AmSC. Both are used by the dashboard and by ML training. The BELLA deployment uses `SF_DB_READONLY_PASSWORD` and `AM_SC_API_KEY`. ## Common Pitfalls and Workarounds diff --git a/docs/docs.yml b/docs/docs.yml index 8efa385b..bc5e64d4 100644 --- a/docs/docs.yml +++ b/docs/docs.yml @@ -11,3 +11,4 @@ dependencies: - sphinx-autobuild - sphinx-book-theme - sphinx-copybutton + - sphinx-design diff --git a/docs/source/conf.py b/docs/source/conf.py index 85d3f562..88b0aaf1 100644 --- a/docs/source/conf.py +++ b/docs/source/conf.py @@ -19,8 +19,17 @@ # -- General configuration --------------------------------------------------- # https://www.sphinx-doc.org/en/master/usage/configuration.html#general-configuration -extensions = ["myst_parser", "sphinx.ext.extlinks", "sphinx_copybutton"] +extensions = [ + "myst_parser", + "sphinx.ext.extlinks", + "sphinx_copybutton", + "sphinx_design", +] myst_heading_anchors = 4 +# ":::" fences, used to nest the sphinx-design tab sets. All tab sets use the +# sync group "deployment" with the keys "general" (default) and "bella-nersc", +# so a selected tab applies to all pages. +myst_enable_extensions = ["colon_fence"] # Roles that turn a repository-relative path into a link to GitHub, so that # source files and directories mentioned in the docs stay navigable: diff --git a/docs/source/dashboard.md b/docs/source/dashboard.md index 880c0ff5..68694b84 100644 --- a/docs/source/dashboard.md +++ b/docs/source/dashboard.md @@ -38,26 +38,66 @@ conda-lock install --name synapse-gui environment-lock.yml #### Run the dashboard -1. Create an SSH tunnel to access the MongoDB database at NERSC (in a separate terminal): +1. If the computer running the dashboard cannot directly reach the host specified by `database.host`, create an SSH tunnel to the MongoDB database through a gateway node in a separate terminal, using `database.host` and `database.port` from your experiment's `config.yaml`: + + ::::{tab-set} + :sync-group: deployment + + :::{tab-item} General + :sync: general + + ```bash + ssh -L :: @ -N + ``` + ::: + + :::{tab-item} Project Example: BELLA @ NERSC + :sync: bella-nersc + ```bash ssh -L 27017:mongodb05.nersc.gov:27017 @dtn03.nersc.gov -N ``` + ::: + :::: + +2. If you created the SSH tunnel, set `database.host` to `127.0.0.1` in your local copy of `config.yaml`, so that the dashboard connects through it, but do not commit this change. + ```yaml + database: + host: "127.0.0.1" + ``` + +3. Move to the {repo-dir}`dashboard/` directory. + +4. Set up the read-only database password and the MLflow API key: + + ::::{tab-set} + :sync-group: deployment + + :::{tab-item} General + :sync: general + + ```bash + export ='' # Use SINGLE quotes around the password! + export ='' # Required when MLflow tracking_uri is AmSC + ``` + ::: -2. Move to the {repo-dir}`dashboard/` directory. + :::{tab-item} Project Example: BELLA @ NERSC + :sync: bella-nersc -3. Set up the database settings (read-only) and the AmSC MLflow API key: ```bash - export SF_DB_HOST='127.0.0.1' - export SF_DB_READONLY_PASSWORD='your_password_here' # Use SINGLE quotes around the password! - export AM_SC_API_KEY='your_amsc_api_key_here' # Required when MLflow tracking_uri is AmSC + export SF_DB_READONLY_PASSWORD='' # Use SINGLE quotes around the password! + export AM_SC_API_KEY='' # Required when MLflow tracking_uri is AmSC ``` + ::: + :::: -4. Activate the conda environment `synapse-gui`: +5. Activate the conda environment `synapse-gui`: ```bash conda activate synapse-gui ``` -5. Run the dashboard as a web application: +6. Run the dashboard as a web application: ```bash python -u app.py --port 8080 ``` @@ -66,28 +106,87 @@ conda-lock install --name synapse-gui environment-lock.yml #### Run the container -1. Create an SSH tunnel to access the MongoDB database at NERSC (in a separate terminal): +1. If the computer running the dashboard cannot directly reach the host specified by `database.host`, create an SSH tunnel to the MongoDB database through a gateway node in a separate terminal, using `database.host` and `database.port` from your experiment's `config.yaml`: + + ::::{tab-set} + :sync-group: deployment + + :::{tab-item} General + :sync: general + + ```bash + ssh -L :: @ -N + ``` + ::: + + :::{tab-item} Project Example: BELLA @ NERSC + :sync: bella-nersc + ```bash ssh -L 27017:mongodb05.nersc.gov:27017 @dtn03.nersc.gov -N ``` + ::: + :::: -2. Move to the root directory of the repository. +2. If you created the SSH tunnel, set `database.host` to `127.0.0.1` in your local copy of `config.yaml`, but do not commit this change. + Do this before building the image, because {repo}`dashboard.Dockerfile` copies {repo-dir}`experiments/` into it. + +3. Move to the root directory of the repository. + +4. Build the Docker image as described [below](#build-the-docker-image). + +5. Run the Docker container with the read-only database password and the MLflow API key: + + ::::{tab-set} + :sync-group: deployment + + :::{tab-item} General + :sync: general + + ```bash + docker run --network=host -v /etc/localtime:/etc/localtime -v $PWD/ml:/app/ml -e ='' -e ='' synapse-gui + ``` + For debugging, you can enter the container without starting the app: + ```bash + docker run --network=host -v /etc/localtime:/etc/localtime -v $PWD/ml:/app/ml -e ='' -e ='' -it synapse-gui bash + ``` + ::: -3. Build the Docker image as described [below](#build-the-docker-image). + :::{tab-item} Project Example: BELLA @ NERSC + :sync: bella-nersc -4. Run the Docker container: ```bash - docker run --network=host -v /etc/localtime:/etc/localtime -v $PWD/ml:/app/ml -e SF_DB_HOST='127.0.0.1' -e SF_DB_READONLY_PASSWORD='your_password_here' -e AM_SC_API_KEY='your_amsc_api_key_here' synapse-gui + docker run --network=host -v /etc/localtime:/etc/localtime -v $PWD/ml:/app/ml -e SF_DB_READONLY_PASSWORD='' -e AM_SC_API_KEY='' synapse-gui ``` For debugging, you can enter the container without starting the app: ```bash - docker run --network=host -v /etc/localtime:/etc/localtime -v $PWD/ml:/app/ml -e SF_DB_HOST='127.0.0.1' -e SF_DB_READONLY_PASSWORD='your_password_here' -e AM_SC_API_KEY='your_amsc_api_key_here' -it synapse-gui bash + docker run --network=host -v /etc/localtime:/etc/localtime -v $PWD/ml:/app/ml -e SF_DB_READONLY_PASSWORD='' -e AM_SC_API_KEY='' -it synapse-gui bash ``` + ::: + :::: + Note that `-v /etc/localtime:/etc/localtime` is necessary to synchronize the time zone in the container with the host machine. ## Run the dashboard at NERSC -Connect to the [dashboard](https://bellasuperfacility.lbl.gov/) deployed at NERSC through Spin and explore it. +Connect to the dashboard deployed for your project at NERSC through Spin and explore it: + +::::{tab-set} +:sync-group: deployment + +:::{tab-item} General +:sync: general + +Open `https:///`, the URL of your project's Spin deployment. +::: + +:::{tab-item} Project Example: BELLA @ NERSC +:sync: bella-nersc + +Open [bellasuperfacility.lbl.gov](https://bellasuperfacility.lbl.gov/). +::: +:::: + You need to upload valid Superfacility API credentials before you can launch simulations or train ML models directly from the dashboard. ## Generate Superfacility API credentials @@ -100,7 +199,24 @@ Follow the instructions at [docs.nersc.gov/services/sfapi/authentication/#client 3. Scroll down to the section "Superfacility API Clients" and click "New Client". -4. Enter a client name (e.g., "Synapse"), choose `sf558` for the user, choose "Red" security level, and select either "Your IP" or "Spin" from the "IP Presets" menu, depending on whether the key will be used from a local computer or from Spin. +4. Enter a client name (e.g., "Synapse"), choose the user that runs the jobs launched from the dashboard, choose "Red" security level, and select either "Your IP" or "Spin" from the "IP Presets" menu, depending on whether the key will be used from a local computer or from Spin. + ML training jobs expect credential files in the `$HOME` of that user, as described in [Through the dashboard](ml-training.md#through-the-dashboard). + + ::::{tab-set} + :sync-group: deployment + + :::{tab-item} General + :sync: general + + Choose your own NERSC user or a collaboration account of your NERSC project. + ::: + + :::{tab-item} Project Example: BELLA @ NERSC + :sync: bella-nersc + + Choose the collaboration account `sf558`. + ::: + :::: 5. Download the private key file in PEM format and save it as `priv_key.pem` in the root directory of the dashboard. Each time the dashboard is launched, it will automatically find the existing key file and load the corresponding credentials. @@ -135,8 +251,9 @@ The dashboard has three routes, reachable from the navigation drawer: The `Parameters` tab holds the displayed output selector, the input parameter controls, and the plot depth control. The `Optimization` tab holds the optimization controls. The `ML` tab holds the model controls and the calibration controls. -- `/hpc` ("HPC Connection"): NERSC Superfacility API credential and Perlmutter status panel. -- `/chat` ("AI Assistant"): embedded assistant route for experiment support; currently backed by [synapse-chat.lbl.gov](https://synapse-chat.lbl.gov/). +- `/hpc` ("HPC Connection"): Genesis AmSC IRI API or NERSC Superfacility API credential and HPC status panel. +- `/chat` ("AI Assistant"): embedded assistant for experiment support. + It loads [synapse-chat.lbl.gov](https://synapse-chat.lbl.gov/), which is hardcoded in {repo}`dashboard/app.py`. The experiment selector, the date range selector, and the error panel belong to the shared layout rather than to any single route, so they appear on all three. @@ -177,6 +294,8 @@ Run this workflow automatically with the Python script {repo}`publish_container. ```bash python publish_container.py --gui ``` +The script pushes to `registry.nersc.gov/m558/superfacility`, which is hardcoded for the BELLA deployment. +For other projects, follow the steps below. ```` ````{tip} @@ -206,16 +325,51 @@ docker system prune -a # Password: your NERSC password without 2FA ``` -3. Tag the Docker image: +3. Tag the Docker image with the registry path of your NERSC project: + + ::::{tab-set} + :sync-group: deployment + + :::{tab-item} General + :sync: general + + ```bash + docker tag synapse-gui:latest registry.nersc.gov//[/]synapse-gui:latest + docker tag synapse-gui:latest registry.nersc.gov//[/]synapse-gui:$(date "+%y.%m") + ``` + ::: + + :::{tab-item} Project Example: BELLA @ NERSC + :sync: bella-nersc + ```bash docker tag synapse-gui:latest registry.nersc.gov/m558/superfacility/synapse-gui:latest docker tag synapse-gui:latest registry.nersc.gov/m558/superfacility/synapse-gui:$(date "+%y.%m") ``` + ::: + :::: 4. Push the Docker image: + + ::::{tab-set} + :sync-group: deployment + + :::{tab-item} General + :sync: general + + ```bash + docker push -a registry.nersc.gov//[/]synapse-gui + ``` + ::: + + :::{tab-item} Project Example: BELLA @ NERSC + :sync: bella-nersc + ```bash docker push -a registry.nersc.gov/m558/superfacility/synapse-gui ``` + ::: + :::: ## References diff --git a/docs/source/deployment.md b/docs/source/deployment.md index 65fc35bb..73ee8071 100644 --- a/docs/source/deployment.md +++ b/docs/source/deployment.md @@ -1,6 +1,8 @@ # Deployment -Synapse is deployed using Docker images and NERSC services. +Synapse is typically deployed using Docker images, e.g., on Kubernetes. + +Below, we document our public deployment workflow (recipes currently in a [private repository](https://github.com/BLAST-AI-ML/synapse-kubernetes-nersc)) using NERSC services like [Spin](https://docs.nersc.gov/services/spin/). ## Build the dashboard image @@ -31,6 +33,29 @@ python publish_container.py --gui --ml ## NERSC deployment assumptions - Dashboard runs on Spin. -- Training and simulations run on Perlmutter through Superfacility API. -- Images are pushed to `registry.nersc.gov/m558/superfacility`. +- Training and simulations run on Perlmutter through Genesis AmSC IRI API or NERSC Superfacility API. +- Images are pushed to the registry of the deployment's NERSC project: + + ::::{tab-set} + :sync-group: deployment + + :::{tab-item} General + :sync: general + + ```text + registry.nersc.gov//[/] + ``` + ::: + + :::{tab-item} Project Example: BELLA @ NERSC + :sync: bella-nersc + + ```text + registry.nersc.gov/m558/superfacility/ + ``` + ::: + :::: + + {repo}`publish_container.py` hardcodes the BELLA path. + For other projects, tag and push the images manually, as described in [Dashboard](dashboard.md#push-the-docker-image) and [ML training](ml-training.md#push-the-docker-image). - Before publishing, validate the images locally. diff --git a/docs/source/developer-notes.md b/docs/source/developer-notes.md index a462be13..084bac2b 100644 --- a/docs/source/developer-notes.md +++ b/docs/source/developer-notes.md @@ -12,8 +12,8 @@ Ruff runs with its default rule set; there is no `pyproject.toml` or `ruff.toml` ## Conda environments -- Dashboard dependencies live in {repo}`dashboard/environment.yml`. -- ML dependencies live in {repo}`ml/environment.yml`. +- Dashboard dependencies are defined in {repo}`dashboard/environment.yml`. +- ML dependencies are defined in {repo}`ml/environment.yml`. - Regenerate the corresponding `environment-lock.yml` after dependency changes. ## Build the documentation diff --git a/docs/source/experiment-configuration.md b/docs/source/experiment-configuration.md index 22f993fb..84383de6 100644 --- a/docs/source/experiment-configuration.md +++ b/docs/source/experiment-configuration.md @@ -20,6 +20,63 @@ Each experiment should provide: - `inputs`: scalar variables with `name`, `type`, `default`, and `value_range`. - `outputs`: scalar variables with `name` and `type`. +## Database and MLflow settings + +The `database` and `mlflow` sections hold the connection settings of a deployment. +Secrets are not stored in `config.yaml`: the keys ending in `_env` name the environment variables that hold them. + +- `database.host`, `database.port`: the MongoDB server. + When you access the database through an SSH tunnel, set `database.host` to `127.0.0.1` in your local copy of `config.yaml` (see [Getting started](getting-started.md#run-the-dashboard)). +- `database.name`: the database. +- `database.auth`: the authentication database of the user. +- `database.username_ro`: the read-only user. +- `database.password_ro_env`: the environment variable that holds the password of the read-only user. +- `mlflow.tracking_uri`: the MLflow tracking server. + Without it, the dashboard cannot load models and the training script does not register them. +- `mlflow.api_key_env`: the environment variable that holds the MLflow API key, only used when `mlflow.tracking_uri` is the AmSC MLflow server, `https://mlflow.american-science-cloud.org`. + +::::{tab-set} +:sync-group: deployment + +:::{tab-item} General +:sync: general + +```yaml +experiment: "" + +database: + host: "" + port: + name: "" + auth: "" + username_ro: "" + password_ro_env: "" + +mlflow: + tracking_uri: "" + api_key_env: "" +``` +::: + +:::{tab-item} Project Example: BELLA @ NERSC +:sync: bella-nersc + +```yaml +experiment: "bella-ip2" + +database: + host: "mongodb05.nersc.gov" + port: 27017 + # name, auth, username_ro: see the experiment repository + password_ro_env: "SF_DB_READONLY_PASSWORD" + +mlflow: + tracking_uri: "https://mlflow.american-science-cloud.org" + api_key_env: "AM_SC_API_KEY" +``` +::: +:::: + ## Simulation calibration `simulation_calibration` maps simulation variable names to experimental variable names: diff --git a/docs/source/getting-started.md b/docs/source/getting-started.md index 75b31e33..dd08454d 100644 --- a/docs/source/getting-started.md +++ b/docs/source/getting-started.md @@ -2,43 +2,143 @@ For a reproducible installation, use the pinned `environment-lock.yml` of {repo-dir}`dashboard/` or {repo-dir}`ml/` rather than the unpinned `environment.yml`. +## Conventions + +Synapse is set up per project: each deployment brings its own experiment configurations, MongoDB database, MLflow server, and NERSC project. +Where a command or a value depends on the deployment, these docs show it in two tabs: + +- **General**: the command with placeholders for the values of your deployment. +- **Project Example: BELLA @ NERSC**: the same command with the values of the BELLA deployment at NERSC. + +The tab you select applies to all pages while you browse in the same browser tab. + +Placeholders use the following notation: + +- `<...>`: a value to replace, for example ``. +- `[...]`: an optional part, to replace or to omit, for example `[/]`. +- ``: the value of a key in your experiment's `config.yaml`, for example `` for the `host` key of the `database` section. + See [Experiment configuration](experiment-configuration.md). + +Names such as the NERSC project `m558`, the collaboration account `sf558`, and the host `mongodb05.nersc.gov` are specific to the BELLA deployment. +Other projects use their own. + ## Run the dashboard -From {repo-dir}`dashboard/`, launch {repo}`app.py `: +The dashboard reads the MongoDB connection settings from the `database` section of your experiment's `config.yaml`. +If the computer running the dashboard cannot directly reach the host specified by `database.host`, open an SSH tunnel to the MongoDB database through a gateway node in a separate terminal: + +::::{tab-set} +:sync-group: deployment + +:::{tab-item} General +:sync: general + +```bash +ssh -L :: @ -N +``` +::: + +:::{tab-item} Project Example: BELLA @ NERSC +:sync: bella-nersc + +```bash +ssh -L 27017:mongodb05.nersc.gov:27017 @dtn03.nersc.gov -N +``` +::: +:::: + +If you created the SSH tunnel, set `database.host` to `127.0.0.1` in your local copy of `config.yaml`, but do not commit this change. + +Assuming `conda-lock` is installed in your conda `base` environment, from {repo-dir}`dashboard/` launch {repo}`app.py `: + +::::{tab-set} +:sync-group: deployment + +:::{tab-item} General +:sync: general ```bash +conda activate base conda-lock install --name synapse-gui environment-lock.yml conda activate synapse-gui -export SF_DB_HOST='127.0.0.1' -export SF_DB_READONLY_PASSWORD='...' -export AM_SC_API_KEY='...' +export ='' # Use SINGLE quotes around the password! +export ='...' python -u app.py --port 8080 ``` +::: -For local MongoDB access, open a tunnel first: +:::{tab-item} Project Example: BELLA @ NERSC +:sync: bella-nersc ```bash -ssh -L 27017:mongodb05.nersc.gov:27017 @dtn03.nersc.gov -N +conda activate base +conda-lock install --name synapse-gui environment-lock.yml +conda activate synapse-gui +export SF_DB_READONLY_PASSWORD='' # Use SINGLE quotes around the password! +export AM_SC_API_KEY='...' +python -u app.py --port 8080 ``` +::: +:::: ## Train a model Training requires an experiment configuration. -Experiment configs are not part of this repository: clone the private repository for your experiment into {repo-dir}`experiments/` first, so that `experiments/synapse-/config.yaml` exists. +Experiment configurations are not part of this repository: clone the private repository for your experiment into {repo-dir}`experiments/` first, so that `experiments/synapse-/config.yaml` exists. See [Experiment configuration](experiment-configuration.md) for the expected layout. +If `database.port` is not `27017` and you access the database through a gateway, open a separate tunnel for training with local port `27017`: `ssh -L 27017:: @ -N`. +Unlike the dashboard, {repo}`train_model.py ` does not read `database.port` yet and always connects to the default MongoDB port. + +Assuming `conda-lock` is installed in your conda `base` environment, from {repo-dir}`ml/` run {repo}`train_model.py `: + +::::{tab-set} +:sync-group: deployment + +:::{tab-item} General +:sync: general + +```bash +conda activate base +conda-lock install --name synapse-ml environment-lock.yml +conda activate synapse-ml +export ='...' +export ='...' +python train_model.py --test --config_file ../experiments/synapse-/config.yaml --model NN +``` +::: -From {repo-dir}`ml/`, run {repo}`train_model.py `: +:::{tab-item} Project Example: BELLA @ NERSC +:sync: bella-nersc ```bash +conda activate base conda-lock install --name synapse-ml environment-lock.yml conda activate synapse-ml -export SF_DB_READONLY_PASSWORD='...' +export SF_DB_READONLY_PASSWORD='' # Use SINGLE quotes around the password! export AM_SC_API_KEY='...' python train_model.py --test --config_file ../experiments/synapse-/config.yaml --model NN ``` +::: +:::: ## Required environment variables -- `SF_DB_HOST`: MongoDB host for the dashboard. +The experiment's `config.yaml` names the environment variables that hold the credentials, so that no secret is stored in the configuration file: + +::::{tab-set} +:sync-group: deployment + +:::{tab-item} General +:sync: general + +- ``: read-only MongoDB password. +- ``: MLflow API key, only needed when `mlflow.tracking_uri` is the American Science Cloud (AmSC) MLflow server, `https://mlflow.american-science-cloud.org`. +::: + +:::{tab-item} Project Example: BELLA @ NERSC +:sync: bella-nersc + - `SF_DB_READONLY_PASSWORD`: read-only MongoDB password. -- `AM_SC_API_KEY`: American Science Cloud MLflow API key when the config uses that service. +- `AM_SC_API_KEY`: AmSC MLflow API key. +::: +:::: diff --git a/docs/source/ml-training.md b/docs/source/ml-training.md index 826ac29b..992762d8 100644 --- a/docs/source/ml-training.md +++ b/docs/source/ml-training.md @@ -36,25 +36,70 @@ This section describes how to train ML models locally. #### Run the training -1. Create an SSH tunnel to access the MongoDB database at NERSC (in a separate terminal): +1. If the computer running the training cannot directly reach the host specified by `database.host`, create an SSH tunnel to the MongoDB database through a gateway node in a separate terminal, using `database.host` and `database.port` from your experiment's `config.yaml`: + + ::::{tab-set} + :sync-group: deployment + + :::{tab-item} General + :sync: general + + ```bash + ssh -L 27017:: @ -N + ``` + ::: + + :::{tab-item} Project Example: BELLA @ NERSC + :sync: bella-nersc + ```bash ssh -L 27017:mongodb05.nersc.gov:27017 @dtn03.nersc.gov -N ``` + ::: + :::: -2. Move to the {repo-dir}`ml/` directory. + ```{note} + The local port is 27017 because {repo}`train_model.py ` does not read `database.port` yet and always connects to the default MongoDB port. + ``` + +2. If you created the SSH tunnel, set `database.host` to `127.0.0.1` in your local copy of `config.yaml`, so that the training connects through it, but do not commit this change. + ```yaml + database: + host: "127.0.0.1" + ``` + +3. Move to the {repo-dir}`ml/` directory. + +4. Set up the read-only database password and the MLflow API key: + + ::::{tab-set} + :sync-group: deployment + + :::{tab-item} General + :sync: general -3. Set up the database settings (read-only) and the AmSC MLflow API key: ```bash - export SF_DB_READONLY_PASSWORD='your_password_here' # Use SINGLE quotes around the password! - export AM_SC_API_KEY='your_amsc_api_key_here' # Required when MLflow tracking_uri is AmSC + export ='' # Use SINGLE quotes around the password! + export ='' # Required when MLflow tracking_uri is AmSC ``` + ::: + + :::{tab-item} Project Example: BELLA @ NERSC + :sync: bella-nersc -4. Activate the conda environment `synapse-ml`: + ```bash + export SF_DB_READONLY_PASSWORD='' # Use SINGLE quotes around the password! + export AM_SC_API_KEY='' # Required when MLflow tracking_uri is AmSC + ``` + ::: + :::: + +5. Activate the conda environment `synapse-ml`: ```bash conda activate synapse-ml ``` -5. Run the ML training script in test mode: +6. Run the ML training script in test mode: ```bash python train_model.py --test --model --config_file ``` @@ -75,9 +120,26 @@ It requires a local, empty MLflow server so it does not touch a production serve ``` Optionally, restrict to a specific model type or config file: + + ::::{tab-set} + :sync-group: deployment + + :::{tab-item} General + :sync: general + + ```bash + python tests/test_ml_pipeline.py --model NN --config_file experiments/synapse-/config.yaml + ``` + ::: + + :::{tab-item} Project Example: BELLA @ NERSC + :sync: bella-nersc + ```bash - python tests/test_ml_pipeline.py --model NN --config_file experiments/synapse-bella-ip2/config.yaml + python tests/test_ml_pipeline.py --model NN --config_file experiments/synapse-/config.yaml ``` + ::: + :::: If your MLflow server is running on a different port (e.g. 5001 instead of 5000), pass it explicitly: ```bash @@ -118,11 +180,29 @@ This section describes how to train ML models at NERSC. 1. Move to the {repo-dir}`ml/` directory. -2. Set up the database settings (read-only) and the AmSC MLflow API key: +2. Set up the read-only database password and the MLflow API key: + + ::::{tab-set} + :sync-group: deployment + + :::{tab-item} General + :sync: general + ```bash - export SF_DB_READONLY_PASSWORD='your_password_here' # Use SINGLE quotes around the password! - export AM_SC_API_KEY='your_amsc_api_key_here' # Required when MLflow tracking_uri is AmSC + export ='' # Use SINGLE quotes around the password! + export ='' # Required when MLflow tracking_uri is AmSC ``` + ::: + + :::{tab-item} Project Example: BELLA @ NERSC + :sync: bella-nersc + + ```bash + export SF_DB_READONLY_PASSWORD='' # Use SINGLE quotes around the password! + export AM_SC_API_KEY='' # Required when MLflow tracking_uri is AmSC + ``` + ::: + :::: 3. Activate the conda environment `synapse-ml`: ```bash @@ -146,35 +226,125 @@ The Docker image is pulled from the [NERSC registry](https://registry.nersc.gov) ssh perlmutter-p1.nersc.gov ``` -2. Ensure the file `$HOME/db.profile` contains the read-only database password and the AmSC MLflow API key: `export SF_DB_READONLY_PASSWORD='your_password_here'` and `export AM_SC_API_KEY='your_amsc_api_key_here'`. +2. Ensure the file `$HOME/db-podman.profile` contains the read-only database password and the MLflow API key. + The file is passed to the container with `--env-file`, so write one `NAME=value` per line, without `export` and without quotes, which would become part of the value: + + ::::{tab-set} + :sync-group: deployment + + :::{tab-item} General + :sync: general + + ```text + =your_password_here + =your_api_key_here + ``` + ::: + + :::{tab-item} Project Example: BELLA @ NERSC + :sync: bella-nersc + + ```text + SF_DB_READONLY_PASSWORD= + AM_SC_API_KEY=your_amsc_api_key_here + ``` + ::: + :::: + + Run `chmod 600 $HOME/db-podman.profile` so that only you can read it. + +3. Pull the Docker image from the registry path of your NERSC project: + + ::::{tab-set} + :sync-group: deployment + + :::{tab-item} General + :sync: general + + ```bash + podman-hpc login --username $USER registry.nersc.gov + # Password: your NERSC password without 2FA + podman-hpc pull registry.nersc.gov//[/]synapse-ml:latest + ``` + ::: + + :::{tab-item} Project Example: BELLA @ NERSC + :sync: bella-nersc -3. Pull the Docker image: ```bash podman-hpc login --username $USER registry.nersc.gov # Password: your NERSC password without 2FA podman-hpc pull registry.nersc.gov/m558/superfacility/synapse-ml:latest ``` + ::: + :::: + +4. Allocate a GPU node, charged to your NERSC project, and run the container: + + ::::{tab-set} + :sync-group: deployment + + :::{tab-item} General + :sync: general + + ```bash + salloc -N 1 --ntasks-per-node=1 -t 1:00:00 -q interactive -C gpu --gpu-bind=single:1 -c 32 -G 1 -A + podman-hpc run --gpu -v /etc/localtime:/etc/localtime --env-file $HOME/db-podman.profile -v :/app/ml/config.yaml --rm -it registry.nersc.gov//[/]synapse-ml:latest python -u /app/ml/train_model.py --test --config_file /app/ml/config.yaml --model NN + ``` + ::: + + :::{tab-item} Project Example: BELLA @ NERSC + :sync: bella-nersc -4. Allocate a GPU node and run the container: ```bash salloc -N 1 --ntasks-per-node=1 -t 1:00:00 -q interactive -C gpu --gpu-bind=single:1 -c 32 -G 1 -A m558 - podman-hpc run --gpu -v /etc/localtime:/etc/localtime -v $HOME/db.profile:/root/db.profile -v /path/to/config.yaml:/app/ml/config.yaml --rm -it registry.nersc.gov/m558/superfacility/synapse-ml:latest python -u /app/ml/train_model.py --test --config_file /app/ml/config.yaml --model NN + podman-hpc run --gpu -v /etc/localtime:/etc/localtime --env-file $HOME/db-podman.profile -v :/app/ml/config.yaml --rm -it registry.nersc.gov/m558/superfacility/synapse-ml:latest python -u /app/ml/train_model.py --test --config_file /app/ml/config.yaml --model NN ``` + ::: + :::: + Note that `-v /etc/localtime:/etc/localtime` is necessary to synchronize the time zone in the container with the host machine. ### Through the dashboard ````{warning} -When ML models are trained through the dashboard, Synapse uses NERSC's Superfacility API with the collaboration account `sf558`. -Because this is a non-interactive, non-user account, Synapse also uses a custom user to pull the image from the [NERSC registry](https://registry.nersc.gov) to Perlmutter. -The registry login credentials need to be prepared (only once) in the `$HOME` of user `sf558` (`/global/homes/s/sf558/`), in a file named `registry.profile` with the following content: -```bash -export REGISTRY_USER="robot\$m558+perlmutter-nersc-gov" -export REGISTRY_PASSWORD="..." -``` +When ML models are trained through the dashboard, Synapse submits the batch script {repo}`ml/training_pm.sbatch` through NERSC's Superfacility API, as the user of the Superfacility API client (see [Generate Superfacility API credentials](dashboard.md#generate-superfacility-api-credentials)). +The batch job reads two files from the `$HOME` of that user, which need to be prepared once: + +- `db-podman.profile`: the read-only database password and the MLflow API key, in the format described in [Manually with Docker](#manually-with-docker). +- `registry.profile`: the login credentials used to pull the image from the [NERSC registry](https://registry.nersc.gov) to Perlmutter. + A collaboration account cannot log in to the registry interactively, so use a robot account of your registry project: + + ::::{tab-set} + :sync-group: deployment + + :::{tab-item} General + :sync: general + + ```bash + export REGISTRY_USER="robot\$+" + export REGISTRY_PASSWORD="..." + ``` + ::: + + :::{tab-item} Project Example: BELLA @ NERSC + :sync: bella-nersc + + In the `$HOME` of the collaboration account `sf558`, `/global/homes/s/sf558/`: + ```bash + export REGISTRY_USER="robot\$m558+perlmutter-nersc-gov" + export REGISTRY_PASSWORD="..." + ``` + ::: + :::: ```` -Connect to the [dashboard](https://bellasuperfacility.lbl.gov/) deployed at NERSC through Spin and click the `Train` button in the `ML` panel. +```{note} +{repo}`ml/training_pm.sbatch` and {repo}`dashboard/model_manager.py` hardcode values of the BELLA deployment: the NERSC project `m558`, the `realtime` QoS, the image `registry.nersc.gov/m558/superfacility/synapse-ml`, and the directory `/global/cfs/cdirs/m558/superfacility/model_training/` for the configuration file and the job logs. +Other projects need to adapt these values before training ML models through the dashboard. +``` + +Connect to the dashboard deployed for your project at NERSC through Spin (see [Run the dashboard at NERSC](dashboard.md#run-the-dashboard-at-nersc)) and click the `Train` button in the `ML` panel. You need to upload valid Superfacility API credentials before you can launch simulations or train ML models directly from the dashboard. ## Model types @@ -256,6 +426,8 @@ Run this workflow automatically with the Python script {repo}`publish_container. ```bash python publish_container.py --ml ``` +The script pushes to `registry.nersc.gov/m558/superfacility`, which is hardcoded for the BELLA deployment. +For other projects, follow the steps below. ```` ````{tip} @@ -292,16 +464,51 @@ docker --version # Password: your NERSC password without 2FA ``` -3. Tag the Docker image: +3. Tag the Docker image with the registry path of your NERSC project: + + ::::{tab-set} + :sync-group: deployment + + :::{tab-item} General + :sync: general + + ```bash + docker tag synapse-ml:latest registry.nersc.gov//[/]synapse-ml:latest + docker tag synapse-ml:latest registry.nersc.gov//[/]synapse-ml:$(date "+%y.%m") + ``` + ::: + + :::{tab-item} Project Example: BELLA @ NERSC + :sync: bella-nersc + ```bash docker tag synapse-ml:latest registry.nersc.gov/m558/superfacility/synapse-ml:latest docker tag synapse-ml:latest registry.nersc.gov/m558/superfacility/synapse-ml:$(date "+%y.%m") ``` + ::: + :::: 4. Push the Docker image: + + ::::{tab-set} + :sync-group: deployment + + :::{tab-item} General + :sync: general + + ```bash + docker push -a registry.nersc.gov//[/]synapse-ml + ``` + ::: + + :::{tab-item} Project Example: BELLA @ NERSC + :sync: bella-nersc + ```bash docker push -a registry.nersc.gov/m558/superfacility/synapse-ml ``` + ::: + :::: ## References diff --git a/docs/source/overview.md b/docs/source/overview.md index 531bec3b..5435c2d8 100644 --- a/docs/source/overview.md +++ b/docs/source/overview.md @@ -23,22 +23,22 @@ To display ML predictions, the application requires the following: - **Simulation and experimental data points**: Each data point consists of values for the scalar inputs and outputs defined in the experiment configuration file. Data points are stored in a [MongoDB](https://www.mongodb.com/) database, with each experiment represented by a separate collection. Experimental and simulation data points are stored in the same collection and distinguished by the `experiment_flag` attribute. -- **ML models**: Machine learning models that interpolate between data points and are stored in [MLflow](https://mlflow.org/). +- **ML models**: Machine learning models that interpolate between data points, stored in [MLflow](https://mlflow.org/). - **Simulation movies** (optional): For certain experiments, users can click on simulation data points to visualize simulation movies. - The corresponding MP4 files are stored in the Perlmutter shared file system at `/global/cfs/cdirs/m558/superfacility/simulation_data`. - This directory is mounted on the container image running on Spin. + The corresponding MP4 files are stored in a directory named `simulation_data` on the Perlmutter shared file system, for example `/global/cfs/cdirs/m558/superfacility/simulation_data` for the BELLA deployment. + This directory is mounted on the container image running on Spin (see [Simulation outputs](simulations.md#simulation-outputs)). ## Launching ML training at NERSC ML models can be trained by launching jobs on Perlmutter from the GUI, through the [NERSC Superfacility API](https://docs.nersc.gov/services/sfapi/). The application requires the following: -- **Superfacility API credential file**: Instructions on generating and uploading the credential file from the GUI are in [Dashboard](dashboard.md). +- **Superfacility API credential file**: Instructions on generating and uploading the credential file from the GUI are in [Dashboard](dashboard.md#generate-superfacility-api-credentials). - **Submission script**: The batch script {repo}`ml/training_pm.sbatch` is copied into the container image pushed to the NERSC registry and deployed through Spin (see {repo}`dashboard.Dockerfile`). It serves as a template for Superfacility API job submission when users launch model training from the GUI. - **Python scripts and configuration files**: These include {repo}`ml/train_model.py`, {repo}`ml/Neural_Net_Classes.py`, and the experiment configuration file `config.yaml`. The Python scripts are copied into the ML container image pushed to the NERSC registry (see {repo}`ml.Dockerfile`), and the Superfacility API job runs them from inside that image on Perlmutter, at `/app/ml/`. - When users launch model training from the GUI, only `config.yaml` is copied to the Perlmutter shared file system, at `/global/cfs/cdirs/m558/superfacility/model_training/`, where the batch job mounts it into the container. + When users launch model training from the GUI, only `config.yaml` is copied to the Perlmutter shared file system, for example `/global/cfs/cdirs/m558/superfacility/model_training/` for the BELLA deployment, where the batch job mounts it into the container. The `config.yaml` file is automatically populated with the configuration values specified in the GUI before being copied to the shared file system. ## Workflow diff --git a/docs/source/simulations.md b/docs/source/simulations.md index 3e461561..b00b41f2 100644 --- a/docs/source/simulations.md +++ b/docs/source/simulations.md @@ -28,6 +28,11 @@ Before submission, it writes the current dashboard parameters to `single_simulat /global/cfs/cdirs/m558/superfacility/simulation_running//templates ``` + ```{note} + This path is hardcoded for the BELLA deployment in {repo}`dashboard/parameters_manager.py` and {repo}`dashboard/model_manager.py`. + Other projects need to adapt it before launching simulations from the dashboard. + ``` + 5. The dashboard reads `submission_script_single` and submits it through Superfacility API. 6. Job status is polled until a terminal state, such as completed, failed, or cancelled. @@ -41,6 +46,28 @@ These scripts are experiment-specific and are usually run manually on Perlmutter Simulation records should be written to the experiment's MongoDB collection with `experiment_flag: 0`. Field names should match either the experiment config outputs or the configured simulation calibration variable names. -When a simulation record includes a `data_directory` under `/global/cfs/cdirs/m558/superfacility/simulation_data`, the dashboard can link the record to a plot file in that directory's `plots/` subdirectory. +When a simulation record includes a `data_directory` inside a directory named `simulation_data`, the dashboard can link the record to a plot file in that directory's `plots/` subdirectory. It prefers a single MP4 file and otherwise falls back to the last PNG file whose name contains `iteration`. This support is optional because the experiment's simulation scripts must create the record and its corresponding files. + +The dashboard looks for the `simulation_data` directory at `/app/simulation_data/` in its container, so the deployment mounts it there from the Perlmutter shared file system: + +::::{tab-set} +:sync-group: deployment + +:::{tab-item} General +:sync: general + +```text +/global/cfs/cdirs//[/]simulation_data +``` +::: + +:::{tab-item} Project Example: BELLA @ NERSC +:sync: bella-nersc + +```text +/global/cfs/cdirs/m558/superfacility/simulation_data +``` +::: +::::