From 9b61dc4f922c26265af106191c0152690908896e Mon Sep 17 00:00:00 2001 From: Axel Huebl Date: Fri, 18 Sep 2026 10:42:10 -0700 Subject: [PATCH 01/13] Add Sphinx-Design (for Tabs) --- docs/docs.yml | 1 + docs/source/conf.py | 11 ++++++++++- 2 files changed, 11 insertions(+), 1 deletion(-) diff --git a/docs/docs.yml b/docs/docs.yml index fb6df6a6..e645b9f9 100644 --- a/docs/docs.yml +++ b/docs/docs.yml @@ -10,3 +10,4 @@ dependencies: - sphinx-autobuild - sphinx-book-theme - sphinx-copybutton + - sphinx-design diff --git a/docs/source/conf.py b/docs/source/conf.py index 85d3f562..88b0aaf1 100644 --- a/docs/source/conf.py +++ b/docs/source/conf.py @@ -19,8 +19,17 @@ # -- General configuration --------------------------------------------------- # https://www.sphinx-doc.org/en/master/usage/configuration.html#general-configuration -extensions = ["myst_parser", "sphinx.ext.extlinks", "sphinx_copybutton"] +extensions = [ + "myst_parser", + "sphinx.ext.extlinks", + "sphinx_copybutton", + "sphinx_design", +] myst_heading_anchors = 4 +# ":::" fences, used to nest the sphinx-design tab sets. All tab sets use the +# sync group "deployment" with the keys "general" (default) and "bella-nersc", +# so a selected tab applies to all pages. +myst_enable_extensions = ["colon_fence"] # Roles that turn a repository-relative path into a link to GitHub, so that # source files and directories mentioned in the docs stay navigable: From 53ad856372e9b5d50744162670058a74fd3df7b8 Mon Sep 17 00:00:00 2001 From: Axel Huebl Date: Fri, 18 Sep 2026 10:46:13 -0700 Subject: [PATCH 02/13] Update AGENTS.md --- AGENTS.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 372a1cc3..7d86c8be 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -121,8 +121,8 @@ The only CI workflow is **CodeQL Advanced** (`.github/workflows/codeql.yml`), wh - **MongoDB** is used for persistent data from experiments and simulations. - **MLflow** is used for persistent data from ML models. -- Database access requires SSH tunneling to NERSC when running locally. -- Environment variables: `SF_DB_HOST` (dashboard), `SF_DB_READONLY_PASSWORD` (dashboard and ML training), `AM_SC_API_KEY` (dashboard and ML training, required when MLflow tracking_uri is AmSC). +- Database connection settings are read from the `database` section of the experiment's `config.yaml`. Local runs usually reach the database through an SSH tunnel, with `database.host` set to `127.0.0.1` in a local copy of `config.yaml`. +- Environment variables: `config.yaml` sets their names. `database.password_ro_env` names the read-only database password, and `mlflow.api_key_env` names the MLflow API key, which is only needed when `mlflow.tracking_uri` is AmSC. Both are used by the dashboard and by ML training. The BELLA deployment uses `SF_DB_READONLY_PASSWORD` and `AM_SC_API_KEY`. ## Common Pitfalls and Workarounds From beb8f89adb9aa6ab9a19b173ac97570c3f10a897 Mon Sep 17 00:00:00 2001 From: Axel Huebl Date: Fri, 18 Sep 2026 10:46:23 -0700 Subject: [PATCH 03/13] Docs: Generalize with Tabs --- docs/source/dashboard.md | 188 ++++++++++++++++-- docs/source/deployment.md | 25 ++- docs/source/experiment-configuration.md | 57 ++++++ docs/source/getting-started.md | 110 ++++++++++- docs/source/ml-training.md | 246 ++++++++++++++++++++++-- docs/source/overview.md | 5 +- docs/source/simulations.md | 29 ++- 7 files changed, 613 insertions(+), 47 deletions(-) diff --git a/docs/source/dashboard.md b/docs/source/dashboard.md index 880c0ff5..da388f54 100644 --- a/docs/source/dashboard.md +++ b/docs/source/dashboard.md @@ -38,26 +38,67 @@ conda-lock install --name synapse-gui environment-lock.yml #### Run the dashboard -1. Create an SSH tunnel to access the MongoDB database at NERSC (in a separate terminal): +1. In a separate terminal, create an SSH tunnel to the MongoDB database through a gateway node, using `database.host` and `database.port` from your experiment's `config.yaml`: + + ::::{tab-set} + :sync-group: deployment + + :::{tab-item} General + :sync: general + + ```bash + ssh -L :: @ -N + ``` + ::: + + :::{tab-item} Project Example: BELLA @ NERSC + :sync: bella-nersc + ```bash ssh -L 27017:mongodb05.nersc.gov:27017 @dtn03.nersc.gov -N ``` + ::: + :::: + +2. Set `database.host` to `127.0.0.1` in your local copy of `config.yaml`, so that the dashboard connects through the tunnel. + Do not commit this change. + ```yaml + database: + host: "127.0.0.1" + ``` + +3. Move to the {repo-dir}`dashboard/` directory. + +4. Set up the read-only database password and the MLflow API key: + + ::::{tab-set} + :sync-group: deployment + + :::{tab-item} General + :sync: general + + ```bash + export ='your_password_here' # Use SINGLE quotes around the password! + export ='your_api_key_here' # Required when MLflow tracking_uri is AmSC + ``` + ::: -2. Move to the {repo-dir}`dashboard/` directory. + :::{tab-item} Project Example: BELLA @ NERSC + :sync: bella-nersc -3. Set up the database settings (read-only) and the AmSC MLflow API key: ```bash - export SF_DB_HOST='127.0.0.1' export SF_DB_READONLY_PASSWORD='your_password_here' # Use SINGLE quotes around the password! export AM_SC_API_KEY='your_amsc_api_key_here' # Required when MLflow tracking_uri is AmSC ``` + ::: + :::: -4. Activate the conda environment `synapse-gui`: +5. Activate the conda environment `synapse-gui`: ```bash conda activate synapse-gui ``` -5. Run the dashboard as a web application: +6. Run the dashboard as a web application: ```bash python -u app.py --port 8080 ``` @@ -66,28 +107,88 @@ conda-lock install --name synapse-gui environment-lock.yml #### Run the container -1. Create an SSH tunnel to access the MongoDB database at NERSC (in a separate terminal): +1. In a separate terminal, create an SSH tunnel to the MongoDB database through a gateway node, using `database.host` and `database.port` from your experiment's `config.yaml`: + + ::::{tab-set} + :sync-group: deployment + + :::{tab-item} General + :sync: general + + ```bash + ssh -L :: @ -N + ``` + ::: + + :::{tab-item} Project Example: BELLA @ NERSC + :sync: bella-nersc + ```bash ssh -L 27017:mongodb05.nersc.gov:27017 @dtn03.nersc.gov -N ``` + ::: + :::: + +2. Set `database.host` to `127.0.0.1` in your local copy of `config.yaml`. + Do this before building the image, because {repo}`dashboard.Dockerfile` copies {repo-dir}`experiments/` into it. + Do not commit this change. -2. Move to the root directory of the repository. +3. Move to the root directory of the repository. -3. Build the Docker image as described [below](#build-the-docker-image). +4. Build the Docker image as described [below](#build-the-docker-image). + +5. Run the Docker container with the read-only database password and the MLflow API key: + + ::::{tab-set} + :sync-group: deployment + + :::{tab-item} General + :sync: general -4. Run the Docker container: ```bash - docker run --network=host -v /etc/localtime:/etc/localtime -v $PWD/ml:/app/ml -e SF_DB_HOST='127.0.0.1' -e SF_DB_READONLY_PASSWORD='your_password_here' -e AM_SC_API_KEY='your_amsc_api_key_here' synapse-gui + docker run --network=host -v /etc/localtime:/etc/localtime -v $PWD/ml:/app/ml -e ='your_password_here' -e ='your_api_key_here' synapse-gui ``` For debugging, you can enter the container without starting the app: ```bash - docker run --network=host -v /etc/localtime:/etc/localtime -v $PWD/ml:/app/ml -e SF_DB_HOST='127.0.0.1' -e SF_DB_READONLY_PASSWORD='your_password_here' -e AM_SC_API_KEY='your_amsc_api_key_here' -it synapse-gui bash + docker run --network=host -v /etc/localtime:/etc/localtime -v $PWD/ml:/app/ml -e ='your_password_here' -e ='your_api_key_here' -it synapse-gui bash + ``` + ::: + + :::{tab-item} Project Example: BELLA @ NERSC + :sync: bella-nersc + + ```bash + docker run --network=host -v /etc/localtime:/etc/localtime -v $PWD/ml:/app/ml -e SF_DB_READONLY_PASSWORD='your_password_here' -e AM_SC_API_KEY='your_amsc_api_key_here' synapse-gui ``` + For debugging, you can enter the container without starting the app: + ```bash + docker run --network=host -v /etc/localtime:/etc/localtime -v $PWD/ml:/app/ml -e SF_DB_READONLY_PASSWORD='your_password_here' -e AM_SC_API_KEY='your_amsc_api_key_here' -it synapse-gui bash + ``` + ::: + :::: + Note that `-v /etc/localtime:/etc/localtime` is necessary to synchronize the time zone in the container with the host machine. ## Run the dashboard at NERSC -Connect to the [dashboard](https://bellasuperfacility.lbl.gov/) deployed at NERSC through Spin and explore it. +Connect to the dashboard deployed for your project at NERSC through Spin and explore it: + +::::{tab-set} +:sync-group: deployment + +:::{tab-item} General +:sync: general + +Open `https:///`, the URL of your project's Spin deployment. +::: + +:::{tab-item} Project Example: BELLA @ NERSC +:sync: bella-nersc + +Open [bellasuperfacility.lbl.gov](https://bellasuperfacility.lbl.gov/). +::: +:::: + You need to upload valid Superfacility API credentials before you can launch simulations or train ML models directly from the dashboard. ## Generate Superfacility API credentials @@ -100,7 +201,24 @@ Follow the instructions at [docs.nersc.gov/services/sfapi/authentication/#client 3. Scroll down to the section "Superfacility API Clients" and click "New Client". -4. Enter a client name (e.g., "Synapse"), choose `sf558` for the user, choose "Red" security level, and select either "Your IP" or "Spin" from the "IP Presets" menu, depending on whether the key will be used from a local computer or from Spin. +4. Enter a client name (e.g., "Synapse"), choose the user that runs the jobs launched from the dashboard, choose "Red" security level, and select either "Your IP" or "Spin" from the "IP Presets" menu, depending on whether the key will be used from a local computer or from Spin. + ML training jobs expect credential files in the `$HOME` of that user, as described in [Through the dashboard](ml-training.md#through-the-dashboard). + + ::::{tab-set} + :sync-group: deployment + + :::{tab-item} General + :sync: general + + Choose your own NERSC user or a collaboration account of your NERSC project. + ::: + + :::{tab-item} Project Example: BELLA @ NERSC + :sync: bella-nersc + + Choose the collaboration account `sf558`. + ::: + :::: 5. Download the private key file in PEM format and save it as `priv_key.pem` in the root directory of the dashboard. Each time the dashboard is launched, it will automatically find the existing key file and load the corresponding credentials. @@ -136,7 +254,8 @@ The dashboard has three routes, reachable from the navigation drawer: The `Optimization` tab holds the optimization controls. The `ML` tab holds the model controls and the calibration controls. - `/hpc` ("HPC Connection"): NERSC Superfacility API credential and Perlmutter status panel. -- `/chat` ("AI Assistant"): embedded assistant route for experiment support; currently backed by [synapse-chat.lbl.gov](https://synapse-chat.lbl.gov/). +- `/chat` ("AI Assistant"): embedded assistant for experiment support. + It loads [synapse-chat.lbl.gov](https://synapse-chat.lbl.gov/), which is hardcoded in {repo}`dashboard/app.py`. The experiment selector, the date range selector, and the error panel belong to the shared layout rather than to any single route, so they appear on all three. @@ -177,6 +296,8 @@ Run this workflow automatically with the Python script {repo}`publish_container. ```bash python publish_container.py --gui ``` +The script pushes to `registry.nersc.gov/m558/superfacility`, which is hardcoded for the BELLA deployment. +For other projects, follow the steps below. ```` ````{tip} @@ -206,16 +327,51 @@ docker system prune -a # Password: your NERSC password without 2FA ``` -3. Tag the Docker image: +3. Tag the Docker image with the registry path of your NERSC project: + + ::::{tab-set} + :sync-group: deployment + + :::{tab-item} General + :sync: general + + ```bash + docker tag synapse-gui:latest registry.nersc.gov//[/]synapse-gui:latest + docker tag synapse-gui:latest registry.nersc.gov//[/]synapse-gui:$(date "+%y.%m") + ``` + ::: + + :::{tab-item} Project Example: BELLA @ NERSC + :sync: bella-nersc + ```bash docker tag synapse-gui:latest registry.nersc.gov/m558/superfacility/synapse-gui:latest docker tag synapse-gui:latest registry.nersc.gov/m558/superfacility/synapse-gui:$(date "+%y.%m") ``` + ::: + :::: 4. Push the Docker image: + + ::::{tab-set} + :sync-group: deployment + + :::{tab-item} General + :sync: general + + ```bash + docker push -a registry.nersc.gov//[/]synapse-gui + ``` + ::: + + :::{tab-item} Project Example: BELLA @ NERSC + :sync: bella-nersc + ```bash docker push -a registry.nersc.gov/m558/superfacility/synapse-gui ``` + ::: + :::: ## References diff --git a/docs/source/deployment.md b/docs/source/deployment.md index 65fc35bb..17f7d6bc 100644 --- a/docs/source/deployment.md +++ b/docs/source/deployment.md @@ -32,5 +32,28 @@ python publish_container.py --gui --ml - Dashboard runs on Spin. - Training and simulations run on Perlmutter through Superfacility API. -- Images are pushed to `registry.nersc.gov/m558/superfacility`. +- Images are pushed to the registry of the deployment's NERSC project: + + ::::{tab-set} + :sync-group: deployment + + :::{tab-item} General + :sync: general + + ```text + registry.nersc.gov//[/] + ``` + ::: + + :::{tab-item} Project Example: BELLA @ NERSC + :sync: bella-nersc + + ```text + registry.nersc.gov/m558/superfacility/ + ``` + ::: + :::: + + {repo}`publish_container.py` hardcodes the BELLA path. + For other projects, tag and push the images manually, as described in [Dashboard](dashboard.md#push-the-docker-image) and [ML training](ml-training.md#push-the-docker-image). - Before publishing, validate the images locally. diff --git a/docs/source/experiment-configuration.md b/docs/source/experiment-configuration.md index 22f993fb..84383de6 100644 --- a/docs/source/experiment-configuration.md +++ b/docs/source/experiment-configuration.md @@ -20,6 +20,63 @@ Each experiment should provide: - `inputs`: scalar variables with `name`, `type`, `default`, and `value_range`. - `outputs`: scalar variables with `name` and `type`. +## Database and MLflow settings + +The `database` and `mlflow` sections hold the connection settings of a deployment. +Secrets are not stored in `config.yaml`: the keys ending in `_env` name the environment variables that hold them. + +- `database.host`, `database.port`: the MongoDB server. + When you access the database through an SSH tunnel, set `database.host` to `127.0.0.1` in your local copy of `config.yaml` (see [Getting started](getting-started.md#run-the-dashboard)). +- `database.name`: the database. +- `database.auth`: the authentication database of the user. +- `database.username_ro`: the read-only user. +- `database.password_ro_env`: the environment variable that holds the password of the read-only user. +- `mlflow.tracking_uri`: the MLflow tracking server. + Without it, the dashboard cannot load models and the training script does not register them. +- `mlflow.api_key_env`: the environment variable that holds the MLflow API key, only used when `mlflow.tracking_uri` is the AmSC MLflow server, `https://mlflow.american-science-cloud.org`. + +::::{tab-set} +:sync-group: deployment + +:::{tab-item} General +:sync: general + +```yaml +experiment: "" + +database: + host: "" + port: + name: "" + auth: "" + username_ro: "" + password_ro_env: "" + +mlflow: + tracking_uri: "" + api_key_env: "" +``` +::: + +:::{tab-item} Project Example: BELLA @ NERSC +:sync: bella-nersc + +```yaml +experiment: "bella-ip2" + +database: + host: "mongodb05.nersc.gov" + port: 27017 + # name, auth, username_ro: see the experiment repository + password_ro_env: "SF_DB_READONLY_PASSWORD" + +mlflow: + tracking_uri: "https://mlflow.american-science-cloud.org" + api_key_env: "AM_SC_API_KEY" +``` +::: +:::: + ## Simulation calibration `simulation_calibration` maps simulation variable names to experimental variable names: diff --git a/docs/source/getting-started.md b/docs/source/getting-started.md index 75b31e33..42966d5b 100644 --- a/docs/source/getting-started.md +++ b/docs/source/getting-started.md @@ -2,24 +2,82 @@ For a reproducible installation, use the pinned `environment-lock.yml` of {repo-dir}`dashboard/` or {repo-dir}`ml/` rather than the unpinned `environment.yml`. +## Conventions + +Synapse is set up per project: each deployment brings its own experiment configurations, MongoDB database, MLflow server, and NERSC project. +Where a command or a value depends on the deployment, these docs show it in two tabs: + +- **General**: the command with placeholders for the values of your deployment. +- **Project Example: BELLA @ NERSC**: the same command with the values of the BELLA deployment at NERSC. + +The tab you select applies to all pages while you browse in the same browser tab. + +Placeholders use the following notation: + +- `<...>`: a value to replace, for example ``. +- `[...]`: an optional part, to replace or to omit, for example `[/]`. +- ``: the value of a key in your experiment's `config.yaml`, for example `` for the `host` key of the `database` section. + See [Experiment configuration](experiment-configuration.md). + +Names such as the NERSC project `m558`, the collaboration account `sf558`, and the host `mongodb05.nersc.gov` are specific to the BELLA deployment. +Other projects use their own. + ## Run the dashboard +The dashboard reads the MongoDB connection settings from the `database` section of your experiment's `config.yaml`. +If the database is only reachable through a gateway node, first open an SSH tunnel in a separate terminal: + +::::{tab-set} +:sync-group: deployment + +:::{tab-item} General +:sync: general + +```bash +ssh -L :: @ -N +``` +::: + +:::{tab-item} Project Example: BELLA @ NERSC +:sync: bella-nersc + +```bash +ssh -L 27017:mongodb05.nersc.gov:27017 @dtn03.nersc.gov -N +``` +::: +:::: + +Then set `database.host` to `127.0.0.1` in your local copy of `config.yaml`, and do not commit this change. + From {repo-dir}`dashboard/`, launch {repo}`app.py `: +::::{tab-set} +:sync-group: deployment + +:::{tab-item} General +:sync: general + ```bash conda-lock install --name synapse-gui environment-lock.yml conda activate synapse-gui -export SF_DB_HOST='127.0.0.1' -export SF_DB_READONLY_PASSWORD='...' -export AM_SC_API_KEY='...' +export ='...' +export ='...' python -u app.py --port 8080 ``` +::: -For local MongoDB access, open a tunnel first: +:::{tab-item} Project Example: BELLA @ NERSC +:sync: bella-nersc ```bash -ssh -L 27017:mongodb05.nersc.gov:27017 @dtn03.nersc.gov -N +conda-lock install --name synapse-gui environment-lock.yml +conda activate synapse-gui +export SF_DB_READONLY_PASSWORD='...' +export AM_SC_API_KEY='...' +python -u app.py --port 8080 ``` +::: +:::: ## Train a model @@ -29,16 +87,52 @@ See [Experiment configuration](experiment-configuration.md) for the expected lay From {repo-dir}`ml/`, run {repo}`train_model.py `: +::::{tab-set} +:sync-group: deployment + +:::{tab-item} General +:sync: general + +```bash +conda-lock install --name synapse-ml environment-lock.yml +conda activate synapse-ml +export ='...' +export ='...' +python train_model.py --test --config_file ../experiments/synapse-/config.yaml --model NN +``` +::: + +:::{tab-item} Project Example: BELLA @ NERSC +:sync: bella-nersc + ```bash conda-lock install --name synapse-ml environment-lock.yml conda activate synapse-ml export SF_DB_READONLY_PASSWORD='...' export AM_SC_API_KEY='...' -python train_model.py --test --config_file ../experiments/synapse-/config.yaml --model NN +python train_model.py --test --config_file ../experiments/synapse-bella-ip2/config.yaml --model NN ``` +::: +:::: ## Required environment variables -- `SF_DB_HOST`: MongoDB host for the dashboard. +The experiment's `config.yaml` names the environment variables that hold the credentials, so that no secret is stored in the configuration file: + +::::{tab-set} +:sync-group: deployment + +:::{tab-item} General +:sync: general + +- ``: read-only MongoDB password. +- ``: MLflow API key, only needed when `mlflow.tracking_uri` is the American Science Cloud (AmSC) MLflow server, `https://mlflow.american-science-cloud.org`. +::: + +:::{tab-item} Project Example: BELLA @ NERSC +:sync: bella-nersc + - `SF_DB_READONLY_PASSWORD`: read-only MongoDB password. -- `AM_SC_API_KEY`: American Science Cloud MLflow API key when the config uses that service. +- `AM_SC_API_KEY`: AmSC MLflow API key. +::: +:::: diff --git a/docs/source/ml-training.md b/docs/source/ml-training.md index 826ac29b..de98736a 100644 --- a/docs/source/ml-training.md +++ b/docs/source/ml-training.md @@ -36,25 +36,71 @@ This section describes how to train ML models locally. #### Run the training -1. Create an SSH tunnel to access the MongoDB database at NERSC (in a separate terminal): +1. In a separate terminal, create an SSH tunnel to the MongoDB database through a gateway node, using `database.host` and `database.port` from your experiment's `config.yaml`: + + ::::{tab-set} + :sync-group: deployment + + :::{tab-item} General + :sync: general + + ```bash + ssh -L 27017:: @ -N + ``` + ::: + + :::{tab-item} Project Example: BELLA @ NERSC + :sync: bella-nersc + ```bash ssh -L 27017:mongodb05.nersc.gov:27017 @dtn03.nersc.gov -N ``` + ::: + :::: -2. Move to the {repo-dir}`ml/` directory. + ```{note} + The local port is 27017 because {repo}`train_model.py ` does not read `database.port` yet and always connects to the default MongoDB port. + ``` + +2. Set `database.host` to `127.0.0.1` in your local copy of `config.yaml`, so that the training connects through the tunnel. + Do not commit this change. + ```yaml + database: + host: "127.0.0.1" + ``` + +3. Move to the {repo-dir}`ml/` directory. + +4. Set up the read-only database password and the MLflow API key: + + ::::{tab-set} + :sync-group: deployment + + :::{tab-item} General + :sync: general + + ```bash + export ='your_password_here' # Use SINGLE quotes around the password! + export ='your_api_key_here' # Required when MLflow tracking_uri is AmSC + ``` + ::: + + :::{tab-item} Project Example: BELLA @ NERSC + :sync: bella-nersc -3. Set up the database settings (read-only) and the AmSC MLflow API key: ```bash export SF_DB_READONLY_PASSWORD='your_password_here' # Use SINGLE quotes around the password! export AM_SC_API_KEY='your_amsc_api_key_here' # Required when MLflow tracking_uri is AmSC ``` + ::: + :::: -4. Activate the conda environment `synapse-ml`: +5. Activate the conda environment `synapse-ml`: ```bash conda activate synapse-ml ``` -5. Run the ML training script in test mode: +6. Run the ML training script in test mode: ```bash python train_model.py --test --model --config_file ``` @@ -75,9 +121,26 @@ It requires a local, empty MLflow server so it does not touch a production serve ``` Optionally, restrict to a specific model type or config file: + + ::::{tab-set} + :sync-group: deployment + + :::{tab-item} General + :sync: general + + ```bash + python tests/test_ml_pipeline.py --model NN --config_file experiments/synapse-/config.yaml + ``` + ::: + + :::{tab-item} Project Example: BELLA @ NERSC + :sync: bella-nersc + ```bash python tests/test_ml_pipeline.py --model NN --config_file experiments/synapse-bella-ip2/config.yaml ``` + ::: + :::: If your MLflow server is running on a different port (e.g. 5001 instead of 5000), pass it explicitly: ```bash @@ -118,11 +181,29 @@ This section describes how to train ML models at NERSC. 1. Move to the {repo-dir}`ml/` directory. -2. Set up the database settings (read-only) and the AmSC MLflow API key: +2. Set up the read-only database password and the MLflow API key: + + ::::{tab-set} + :sync-group: deployment + + :::{tab-item} General + :sync: general + + ```bash + export ='your_password_here' # Use SINGLE quotes around the password! + export ='your_api_key_here' # Required when MLflow tracking_uri is AmSC + ``` + ::: + + :::{tab-item} Project Example: BELLA @ NERSC + :sync: bella-nersc + ```bash export SF_DB_READONLY_PASSWORD='your_password_here' # Use SINGLE quotes around the password! export AM_SC_API_KEY='your_amsc_api_key_here' # Required when MLflow tracking_uri is AmSC ``` + ::: + :::: 3. Activate the conda environment `synapse-ml`: ```bash @@ -146,35 +227,125 @@ The Docker image is pulled from the [NERSC registry](https://registry.nersc.gov) ssh perlmutter-p1.nersc.gov ``` -2. Ensure the file `$HOME/db.profile` contains the read-only database password and the AmSC MLflow API key: `export SF_DB_READONLY_PASSWORD='your_password_here'` and `export AM_SC_API_KEY='your_amsc_api_key_here'`. +2. Ensure the file `$HOME/db-podman.profile` contains the read-only database password and the MLflow API key. + The file is passed to the container with `--env-file`, so write one `NAME=value` per line, without `export` and without quotes, which would become part of the value: + + ::::{tab-set} + :sync-group: deployment + + :::{tab-item} General + :sync: general + + ```text + =your_password_here + =your_api_key_here + ``` + ::: + + :::{tab-item} Project Example: BELLA @ NERSC + :sync: bella-nersc + + ```text + SF_DB_READONLY_PASSWORD=your_password_here + AM_SC_API_KEY=your_amsc_api_key_here + ``` + ::: + :::: + + Run `chmod 600 $HOME/db-podman.profile` so that only you can read it. + +3. Pull the Docker image from the registry path of your NERSC project: + + ::::{tab-set} + :sync-group: deployment + + :::{tab-item} General + :sync: general + + ```bash + podman-hpc login --username $USER registry.nersc.gov + # Password: your NERSC password without 2FA + podman-hpc pull registry.nersc.gov//[/]synapse-ml:latest + ``` + ::: + + :::{tab-item} Project Example: BELLA @ NERSC + :sync: bella-nersc -3. Pull the Docker image: ```bash podman-hpc login --username $USER registry.nersc.gov # Password: your NERSC password without 2FA podman-hpc pull registry.nersc.gov/m558/superfacility/synapse-ml:latest ``` + ::: + :::: + +4. Allocate a GPU node, charged to your NERSC project, and run the container: + + ::::{tab-set} + :sync-group: deployment + + :::{tab-item} General + :sync: general + + ```bash + salloc -N 1 --ntasks-per-node=1 -t 1:00:00 -q interactive -C gpu --gpu-bind=single:1 -c 32 -G 1 -A + podman-hpc run --gpu -v /etc/localtime:/etc/localtime --env-file $HOME/db-podman.profile -v :/app/ml/config.yaml --rm -it registry.nersc.gov//[/]synapse-ml:latest python -u /app/ml/train_model.py --test --config_file /app/ml/config.yaml --model NN + ``` + ::: + + :::{tab-item} Project Example: BELLA @ NERSC + :sync: bella-nersc -4. Allocate a GPU node and run the container: ```bash salloc -N 1 --ntasks-per-node=1 -t 1:00:00 -q interactive -C gpu --gpu-bind=single:1 -c 32 -G 1 -A m558 - podman-hpc run --gpu -v /etc/localtime:/etc/localtime -v $HOME/db.profile:/root/db.profile -v /path/to/config.yaml:/app/ml/config.yaml --rm -it registry.nersc.gov/m558/superfacility/synapse-ml:latest python -u /app/ml/train_model.py --test --config_file /app/ml/config.yaml --model NN + podman-hpc run --gpu -v /etc/localtime:/etc/localtime --env-file $HOME/db-podman.profile -v :/app/ml/config.yaml --rm -it registry.nersc.gov/m558/superfacility/synapse-ml:latest python -u /app/ml/train_model.py --test --config_file /app/ml/config.yaml --model NN ``` + ::: + :::: + Note that `-v /etc/localtime:/etc/localtime` is necessary to synchronize the time zone in the container with the host machine. ### Through the dashboard ````{warning} -When ML models are trained through the dashboard, Synapse uses NERSC's Superfacility API with the collaboration account `sf558`. -Because this is a non-interactive, non-user account, Synapse also uses a custom user to pull the image from the [NERSC registry](https://registry.nersc.gov) to Perlmutter. -The registry login credentials need to be prepared (only once) in the `$HOME` of user `sf558` (`/global/homes/s/sf558/`), in a file named `registry.profile` with the following content: -```bash -export REGISTRY_USER="robot\$m558+perlmutter-nersc-gov" -export REGISTRY_PASSWORD="..." -``` +When ML models are trained through the dashboard, Synapse submits the batch script {repo}`ml/training_pm.sbatch` through NERSC's Superfacility API, as the user of the Superfacility API client (see [Generate Superfacility API credentials](dashboard.md#generate-superfacility-api-credentials)). +The batch job reads two files from the `$HOME` of that user, which need to be prepared once: + +- `db-podman.profile`: the read-only database password and the MLflow API key, in the format described in [Manually with Docker](#manually-with-docker). +- `registry.profile`: the login credentials used to pull the image from the [NERSC registry](https://registry.nersc.gov) to Perlmutter. + A collaboration account cannot log in to the registry interactively, so use a robot account of your registry project: + + ::::{tab-set} + :sync-group: deployment + + :::{tab-item} General + :sync: general + + ```bash + export REGISTRY_USER="robot\$+" + export REGISTRY_PASSWORD="..." + ``` + ::: + + :::{tab-item} Project Example: BELLA @ NERSC + :sync: bella-nersc + + In the `$HOME` of the collaboration account `sf558`, `/global/homes/s/sf558/`: + ```bash + export REGISTRY_USER="robot\$m558+perlmutter-nersc-gov" + export REGISTRY_PASSWORD="..." + ``` + ::: + :::: ```` -Connect to the [dashboard](https://bellasuperfacility.lbl.gov/) deployed at NERSC through Spin and click the `Train` button in the `ML` panel. +```{note} +{repo}`ml/training_pm.sbatch` and {repo}`dashboard/model_manager.py` hardcode values of the BELLA deployment: the NERSC project `m558`, the image `registry.nersc.gov/m558/superfacility/synapse-ml`, and the directory `/global/cfs/cdirs/m558/superfacility/model_training/` for the configuration file and the job logs. +Other projects need to adapt these values before training ML models through the dashboard. +``` + +Connect to the dashboard deployed for your project at NERSC through Spin (see [Run the dashboard at NERSC](dashboard.md#run-the-dashboard-at-nersc)) and click the `Train` button in the `ML` panel. You need to upload valid Superfacility API credentials before you can launch simulations or train ML models directly from the dashboard. ## Model types @@ -256,6 +427,8 @@ Run this workflow automatically with the Python script {repo}`publish_container. ```bash python publish_container.py --ml ``` +The script pushes to `registry.nersc.gov/m558/superfacility`, which is hardcoded for the BELLA deployment. +For other projects, follow the steps below. ```` ````{tip} @@ -292,16 +465,51 @@ docker --version # Password: your NERSC password without 2FA ``` -3. Tag the Docker image: +3. Tag the Docker image with the registry path of your NERSC project: + + ::::{tab-set} + :sync-group: deployment + + :::{tab-item} General + :sync: general + + ```bash + docker tag synapse-ml:latest registry.nersc.gov//[/]synapse-ml:latest + docker tag synapse-ml:latest registry.nersc.gov//[/]synapse-ml:$(date "+%y.%m") + ``` + ::: + + :::{tab-item} Project Example: BELLA @ NERSC + :sync: bella-nersc + ```bash docker tag synapse-ml:latest registry.nersc.gov/m558/superfacility/synapse-ml:latest docker tag synapse-ml:latest registry.nersc.gov/m558/superfacility/synapse-ml:$(date "+%y.%m") ``` + ::: + :::: 4. Push the Docker image: + + ::::{tab-set} + :sync-group: deployment + + :::{tab-item} General + :sync: general + + ```bash + docker push -a registry.nersc.gov//[/]synapse-ml + ``` + ::: + + :::{tab-item} Project Example: BELLA @ NERSC + :sync: bella-nersc + ```bash docker push -a registry.nersc.gov/m558/superfacility/synapse-ml ``` + ::: + :::: ## References diff --git a/docs/source/overview.md b/docs/source/overview.md index 531bec3b..ba21c381 100644 --- a/docs/source/overview.md +++ b/docs/source/overview.md @@ -25,8 +25,8 @@ To display ML predictions, the application requires the following: Experimental and simulation data points are stored in the same collection and distinguished by the `experiment_flag` attribute. - **ML models**: Machine learning models that interpolate between data points and are stored in [MLflow](https://mlflow.org/). - **Simulation movies** (optional): For certain experiments, users can click on simulation data points to visualize simulation movies. - The corresponding MP4 files are stored in the Perlmutter shared file system at `/global/cfs/cdirs/m558/superfacility/simulation_data`. - This directory is mounted on the container image running on Spin. + The corresponding MP4 files are stored in a directory named `simulation_data` on the Perlmutter shared file system, for example `/global/cfs/cdirs/m558/superfacility/simulation_data` for the BELLA deployment. + This directory is mounted on the container image running on Spin (see [Simulation outputs](simulations.md#simulation-outputs)). ## Launching ML training at NERSC @@ -39,6 +39,7 @@ The application requires the following: - **Python scripts and configuration files**: These include {repo}`ml/train_model.py`, {repo}`ml/Neural_Net_Classes.py`, and the experiment configuration file `config.yaml`. The Python scripts are copied into the ML container image pushed to the NERSC registry (see {repo}`ml.Dockerfile`), and the Superfacility API job runs them from inside that image on Perlmutter, at `/app/ml/`. When users launch model training from the GUI, only `config.yaml` is copied to the Perlmutter shared file system, at `/global/cfs/cdirs/m558/superfacility/model_training/`, where the batch job mounts it into the container. + This path is hardcoded for the BELLA deployment (see [Through the dashboard](ml-training.md#through-the-dashboard)). The `config.yaml` file is automatically populated with the configuration values specified in the GUI before being copied to the shared file system. ## Workflow diff --git a/docs/source/simulations.md b/docs/source/simulations.md index 3e461561..91749c57 100644 --- a/docs/source/simulations.md +++ b/docs/source/simulations.md @@ -28,6 +28,11 @@ Before submission, it writes the current dashboard parameters to `single_simulat /global/cfs/cdirs/m558/superfacility/simulation_running//templates ``` + ```{note} + This path is hardcoded for the BELLA deployment in {repo}`dashboard/parameters_manager.py`. + Other projects need to adapt it before launching simulations from the dashboard. + ``` + 5. The dashboard reads `submission_script_single` and submits it through Superfacility API. 6. Job status is polled until a terminal state, such as completed, failed, or cancelled. @@ -41,6 +46,28 @@ These scripts are experiment-specific and are usually run manually on Perlmutter Simulation records should be written to the experiment's MongoDB collection with `experiment_flag: 0`. Field names should match either the experiment config outputs or the configured simulation calibration variable names. -When a simulation record includes a `data_directory` under `/global/cfs/cdirs/m558/superfacility/simulation_data`, the dashboard can link the record to a plot file in that directory's `plots/` subdirectory. +When a simulation record includes a `data_directory` inside a directory named `simulation_data`, the dashboard can link the record to a plot file in that directory's `plots/` subdirectory. It prefers a single MP4 file and otherwise falls back to the last PNG file whose name contains `iteration`. This support is optional because the experiment's simulation scripts must create the record and its corresponding files. + +The dashboard looks for the `simulation_data` directory at `/app/simulation_data/` in its container, so the deployment mounts it there from the Perlmutter shared file system: + +::::{tab-set} +:sync-group: deployment + +:::{tab-item} General +:sync: general + +```text +/global/cfs/cdirs//[/]simulation_data +``` +::: + +:::{tab-item} Project Example: BELLA @ NERSC +:sync: bella-nersc + +```text +/global/cfs/cdirs/m558/superfacility/simulation_data +``` +::: +:::: From 35e9b90acfc62400996daefdce822a0070eeaa87 Mon Sep 17 00:00:00 2001 From: Edoardo Zoni Date: Fri, 18 Sep 2026 14:26:48 -0700 Subject: [PATCH 04/13] Fix three issues: - The generic dashboard-training guidance misses the hard-coded realtime QoS - The getting-started tunnel is incompatible with training when MongoDB uses a non-default port - A remaining BELLA-specific command is both ungeneralized and invalid --- AGENTS.md | 2 +- docs/source/getting-started.md | 2 ++ docs/source/ml-training.md | 2 +- 3 files changed, 4 insertions(+), 2 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 7d86c8be..8cdeb340 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -90,7 +90,7 @@ docker run -p 127.0.0.1:5000:5000 ghcr.io/mlflow/mlflow mlflow server --host 0.0 python tests/test_ml_pipeline.py # Optionally restrict to a specific model type or config -python tests/test_ml_pipeline.py --model NN --config_file experiments/synapse-bella-ip2 +python tests/test_ml_pipeline.py --model NN --config_file experiments/synapse-/config.yaml ``` Dashboard validation is done manually by running the application. diff --git a/docs/source/getting-started.md b/docs/source/getting-started.md index 42966d5b..0b917b36 100644 --- a/docs/source/getting-started.md +++ b/docs/source/getting-started.md @@ -84,6 +84,8 @@ python -u app.py --port 8080 Training requires an experiment configuration. Experiment configs are not part of this repository: clone the private repository for your experiment into {repo-dir}`experiments/` first, so that `experiments/synapse-/config.yaml` exists. See [Experiment configuration](experiment-configuration.md) for the expected layout. +If `database.port` is not `27017` and you access the database through a gateway, open a separate tunnel for training with local port `27017`: `ssh -L 27017:: @ -N`. +Unlike the dashboard, {repo}`train_model.py ` does not read `database.port` yet and always connects to the default MongoDB port. From {repo-dir}`ml/`, run {repo}`train_model.py `: diff --git a/docs/source/ml-training.md b/docs/source/ml-training.md index de98736a..2bed0569 100644 --- a/docs/source/ml-training.md +++ b/docs/source/ml-training.md @@ -341,7 +341,7 @@ The batch job reads two files from the `$HOME` of that user, which need to be pr ```` ```{note} -{repo}`ml/training_pm.sbatch` and {repo}`dashboard/model_manager.py` hardcode values of the BELLA deployment: the NERSC project `m558`, the image `registry.nersc.gov/m558/superfacility/synapse-ml`, and the directory `/global/cfs/cdirs/m558/superfacility/model_training/` for the configuration file and the job logs. +{repo}`ml/training_pm.sbatch` and {repo}`dashboard/model_manager.py` hardcode values of the BELLA deployment: the NERSC project `m558`, the `realtime` QoS, the image `registry.nersc.gov/m558/superfacility/synapse-ml`, and the directory `/global/cfs/cdirs/m558/superfacility/model_training/` for the configuration file and the job logs. Other projects need to adapt these values before training ML models through the dashboard. ``` From 1e4e6f3fb9279b2ebb819dac30ebf79051fe335d Mon Sep 17 00:00:00 2001 From: Edoardo Zoni Date: Fri, 18 Sep 2026 15:45:54 -0700 Subject: [PATCH 05/13] Fix two remaining commands specific to BELLA iP2 --- docs/source/getting-started.md | 2 +- docs/source/ml-training.md | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/source/getting-started.md b/docs/source/getting-started.md index 0b917b36..c9d1c253 100644 --- a/docs/source/getting-started.md +++ b/docs/source/getting-started.md @@ -112,7 +112,7 @@ conda-lock install --name synapse-ml environment-lock.yml conda activate synapse-ml export SF_DB_READONLY_PASSWORD='...' export AM_SC_API_KEY='...' -python train_model.py --test --config_file ../experiments/synapse-bella-ip2/config.yaml --model NN +python train_model.py --test --config_file ../experiments/synapse-/config.yaml --model NN ``` ::: :::: diff --git a/docs/source/ml-training.md b/docs/source/ml-training.md index 2bed0569..adf8c2d5 100644 --- a/docs/source/ml-training.md +++ b/docs/source/ml-training.md @@ -137,7 +137,7 @@ It requires a local, empty MLflow server so it does not touch a production serve :sync: bella-nersc ```bash - python tests/test_ml_pipeline.py --model NN --config_file experiments/synapse-bella-ip2/config.yaml + python tests/test_ml_pipeline.py --model NN --config_file experiments/synapse-/config.yaml ``` ::: :::: From 712bf27ad6eb276f80bb96b81c4a65f2be0d3105 Mon Sep 17 00:00:00 2001 From: Edoardo Zoni <59625522+EZoni@users.noreply.github.com> Date: Mon, 21 Sep 2026 13:23:07 -0700 Subject: [PATCH 06/13] Add `conda activate base` to code snippets in "Getting started" Co-authored-by: Edoardo Zoni <59625522+EZoni@users.noreply.github.com> --- docs/source/getting-started.md | 4 ++++ 1 file changed, 4 insertions(+) diff --git a/docs/source/getting-started.md b/docs/source/getting-started.md index c9d1c253..43c8b171 100644 --- a/docs/source/getting-started.md +++ b/docs/source/getting-started.md @@ -58,6 +58,7 @@ From {repo-dir}`dashboard/`, launch {repo}`app.py `: :sync: general ```bash +conda activate base conda-lock install --name synapse-gui environment-lock.yml conda activate synapse-gui export ='...' @@ -70,6 +71,7 @@ python -u app.py --port 8080 :sync: bella-nersc ```bash +conda activate base conda-lock install --name synapse-gui environment-lock.yml conda activate synapse-gui export SF_DB_READONLY_PASSWORD='...' @@ -96,6 +98,7 @@ From {repo-dir}`ml/`, run {repo}`train_model.py `: :sync: general ```bash +conda activate base conda-lock install --name synapse-ml environment-lock.yml conda activate synapse-ml export ='...' @@ -108,6 +111,7 @@ python train_model.py --test --config_file ../experiments/synapse-/c :sync: bella-nersc ```bash +conda activate base conda-lock install --name synapse-ml environment-lock.yml conda activate synapse-ml export SF_DB_READONLY_PASSWORD='...' From af97f05bab2ec6394e32f5baac7061036a49df15 Mon Sep 17 00:00:00 2001 From: Edoardo Zoni Date: Mon, 21 Sep 2026 15:50:08 -0700 Subject: [PATCH 07/13] Miscellaneous fixes --- docs/source/deployment.md | 2 +- docs/source/developer-notes.md | 4 ++-- docs/source/getting-started.md | 6 +++--- docs/source/overview.md | 7 +++---- 4 files changed, 9 insertions(+), 10 deletions(-) diff --git a/docs/source/deployment.md b/docs/source/deployment.md index 17f7d6bc..8d069fca 100644 --- a/docs/source/deployment.md +++ b/docs/source/deployment.md @@ -1,6 +1,6 @@ # Deployment -Synapse is deployed using Docker images and NERSC services. +Synapse is currently deployed using Docker images and NERSC services. ## Build the dashboard image diff --git a/docs/source/developer-notes.md b/docs/source/developer-notes.md index a462be13..084bac2b 100644 --- a/docs/source/developer-notes.md +++ b/docs/source/developer-notes.md @@ -12,8 +12,8 @@ Ruff runs with its default rule set; there is no `pyproject.toml` or `ruff.toml` ## Conda environments -- Dashboard dependencies live in {repo}`dashboard/environment.yml`. -- ML dependencies live in {repo}`ml/environment.yml`. +- Dashboard dependencies are defined in {repo}`dashboard/environment.yml`. +- ML dependencies are defined in {repo}`ml/environment.yml`. - Regenerate the corresponding `environment-lock.yml` after dependency changes. ## Build the documentation diff --git a/docs/source/getting-started.md b/docs/source/getting-started.md index 43c8b171..53f952d4 100644 --- a/docs/source/getting-started.md +++ b/docs/source/getting-started.md @@ -49,7 +49,7 @@ ssh -L 27017:mongodb05.nersc.gov:27017 @dtn03.nersc.gov -N Then set `database.host` to `127.0.0.1` in your local copy of `config.yaml`, and do not commit this change. -From {repo-dir}`dashboard/`, launch {repo}`app.py `: +Assuming `conda-lock` is installed in your conda `base` environment, from {repo-dir}`dashboard/` launch {repo}`app.py `: ::::{tab-set} :sync-group: deployment @@ -84,12 +84,12 @@ python -u app.py --port 8080 ## Train a model Training requires an experiment configuration. -Experiment configs are not part of this repository: clone the private repository for your experiment into {repo-dir}`experiments/` first, so that `experiments/synapse-/config.yaml` exists. +Experiment configurations are not part of this repository: clone the private repository for your experiment into {repo-dir}`experiments/` first, so that `experiments/synapse-/config.yaml` exists. See [Experiment configuration](experiment-configuration.md) for the expected layout. If `database.port` is not `27017` and you access the database through a gateway, open a separate tunnel for training with local port `27017`: `ssh -L 27017:: @ -N`. Unlike the dashboard, {repo}`train_model.py ` does not read `database.port` yet and always connects to the default MongoDB port. -From {repo-dir}`ml/`, run {repo}`train_model.py `: +Assuming `conda-lock` is installed in your conda `base` environment, from {repo-dir}`ml/` run {repo}`train_model.py `: ::::{tab-set} :sync-group: deployment diff --git a/docs/source/overview.md b/docs/source/overview.md index ba21c381..5435c2d8 100644 --- a/docs/source/overview.md +++ b/docs/source/overview.md @@ -23,7 +23,7 @@ To display ML predictions, the application requires the following: - **Simulation and experimental data points**: Each data point consists of values for the scalar inputs and outputs defined in the experiment configuration file. Data points are stored in a [MongoDB](https://www.mongodb.com/) database, with each experiment represented by a separate collection. Experimental and simulation data points are stored in the same collection and distinguished by the `experiment_flag` attribute. -- **ML models**: Machine learning models that interpolate between data points and are stored in [MLflow](https://mlflow.org/). +- **ML models**: Machine learning models that interpolate between data points, stored in [MLflow](https://mlflow.org/). - **Simulation movies** (optional): For certain experiments, users can click on simulation data points to visualize simulation movies. The corresponding MP4 files are stored in a directory named `simulation_data` on the Perlmutter shared file system, for example `/global/cfs/cdirs/m558/superfacility/simulation_data` for the BELLA deployment. This directory is mounted on the container image running on Spin (see [Simulation outputs](simulations.md#simulation-outputs)). @@ -33,13 +33,12 @@ To display ML predictions, the application requires the following: ML models can be trained by launching jobs on Perlmutter from the GUI, through the [NERSC Superfacility API](https://docs.nersc.gov/services/sfapi/). The application requires the following: -- **Superfacility API credential file**: Instructions on generating and uploading the credential file from the GUI are in [Dashboard](dashboard.md). +- **Superfacility API credential file**: Instructions on generating and uploading the credential file from the GUI are in [Dashboard](dashboard.md#generate-superfacility-api-credentials). - **Submission script**: The batch script {repo}`ml/training_pm.sbatch` is copied into the container image pushed to the NERSC registry and deployed through Spin (see {repo}`dashboard.Dockerfile`). It serves as a template for Superfacility API job submission when users launch model training from the GUI. - **Python scripts and configuration files**: These include {repo}`ml/train_model.py`, {repo}`ml/Neural_Net_Classes.py`, and the experiment configuration file `config.yaml`. The Python scripts are copied into the ML container image pushed to the NERSC registry (see {repo}`ml.Dockerfile`), and the Superfacility API job runs them from inside that image on Perlmutter, at `/app/ml/`. - When users launch model training from the GUI, only `config.yaml` is copied to the Perlmutter shared file system, at `/global/cfs/cdirs/m558/superfacility/model_training/`, where the batch job mounts it into the container. - This path is hardcoded for the BELLA deployment (see [Through the dashboard](ml-training.md#through-the-dashboard)). + When users launch model training from the GUI, only `config.yaml` is copied to the Perlmutter shared file system, for example `/global/cfs/cdirs/m558/superfacility/model_training/` for the BELLA deployment, where the batch job mounts it into the container. The `config.yaml` file is automatically populated with the configuration values specified in the GUI before being copied to the shared file system. ## Workflow From 4cebc86669c9ed93e07d53fce09ff111aba95bb7 Mon Sep 17 00:00:00 2001 From: Edoardo Zoni Date: Tue, 22 Sep 2026 11:58:41 -0700 Subject: [PATCH 08/13] Reword instructions about SSH tunnel and database.host --- docs/source/dashboard.md | 10 ++++------ docs/source/getting-started.md | 4 ++-- docs/source/ml-training.md | 5 ++--- 3 files changed, 8 insertions(+), 11 deletions(-) diff --git a/docs/source/dashboard.md b/docs/source/dashboard.md index da388f54..d190bbf5 100644 --- a/docs/source/dashboard.md +++ b/docs/source/dashboard.md @@ -38,7 +38,7 @@ conda-lock install --name synapse-gui environment-lock.yml #### Run the dashboard -1. In a separate terminal, create an SSH tunnel to the MongoDB database through a gateway node, using `database.host` and `database.port` from your experiment's `config.yaml`: +1. If the computer running the dashboard cannot directly reach the host specified by `database.host`, create an SSH tunnel to the MongoDB database through a gateway node in a separate terminal, using `database.host` and `database.port` from your experiment's `config.yaml`: ::::{tab-set} :sync-group: deployment @@ -60,8 +60,7 @@ conda-lock install --name synapse-gui environment-lock.yml ::: :::: -2. Set `database.host` to `127.0.0.1` in your local copy of `config.yaml`, so that the dashboard connects through the tunnel. - Do not commit this change. +2. If you created the SSH tunnel, set `database.host` to `127.0.0.1` in your local copy of `config.yaml`, so that the dashboard connects through it, but do not commit this change. ```yaml database: host: "127.0.0.1" @@ -107,7 +106,7 @@ conda-lock install --name synapse-gui environment-lock.yml #### Run the container -1. In a separate terminal, create an SSH tunnel to the MongoDB database through a gateway node, using `database.host` and `database.port` from your experiment's `config.yaml`: +1. If the computer running the dashboard cannot directly reach the host specified by `database.host`, create an SSH tunnel to the MongoDB database through a gateway node in a separate terminal, using `database.host` and `database.port` from your experiment's `config.yaml`: ::::{tab-set} :sync-group: deployment @@ -129,9 +128,8 @@ conda-lock install --name synapse-gui environment-lock.yml ::: :::: -2. Set `database.host` to `127.0.0.1` in your local copy of `config.yaml`. +2. If you created the SSH tunnel, set `database.host` to `127.0.0.1` in your local copy of `config.yaml`, but do not commit this change. Do this before building the image, because {repo}`dashboard.Dockerfile` copies {repo-dir}`experiments/` into it. - Do not commit this change. 3. Move to the root directory of the repository. diff --git a/docs/source/getting-started.md b/docs/source/getting-started.md index 53f952d4..8578f225 100644 --- a/docs/source/getting-started.md +++ b/docs/source/getting-started.md @@ -25,7 +25,7 @@ Other projects use their own. ## Run the dashboard The dashboard reads the MongoDB connection settings from the `database` section of your experiment's `config.yaml`. -If the database is only reachable through a gateway node, first open an SSH tunnel in a separate terminal: +If the computer running the dashboard cannot directly reach the host specified by `database.host`, open an SSH tunnel to the MongoDB database through a gateway node in a separate terminal: ::::{tab-set} :sync-group: deployment @@ -47,7 +47,7 @@ ssh -L 27017:mongodb05.nersc.gov:27017 @dtn03.nersc.gov -N ::: :::: -Then set `database.host` to `127.0.0.1` in your local copy of `config.yaml`, and do not commit this change. +If you created the SSH tunnel, set `database.host` to `127.0.0.1` in your local copy of `config.yaml`, but do not commit this change. Assuming `conda-lock` is installed in your conda `base` environment, from {repo-dir}`dashboard/` launch {repo}`app.py `: diff --git a/docs/source/ml-training.md b/docs/source/ml-training.md index adf8c2d5..9c17fa7e 100644 --- a/docs/source/ml-training.md +++ b/docs/source/ml-training.md @@ -36,7 +36,7 @@ This section describes how to train ML models locally. #### Run the training -1. In a separate terminal, create an SSH tunnel to the MongoDB database through a gateway node, using `database.host` and `database.port` from your experiment's `config.yaml`: +1. If the computer running the training cannot directly reach the host specified by `database.host`, create an SSH tunnel to the MongoDB database through a gateway node in a separate terminal, using `database.host` and `database.port` from your experiment's `config.yaml`: ::::{tab-set} :sync-group: deployment @@ -62,8 +62,7 @@ This section describes how to train ML models locally. The local port is 27017 because {repo}`train_model.py ` does not read `database.port` yet and always connects to the default MongoDB port. ``` -2. Set `database.host` to `127.0.0.1` in your local copy of `config.yaml`, so that the training connects through the tunnel. - Do not commit this change. +2. If you created the SSH tunnel, set `database.host` to `127.0.0.1` in your local copy of `config.yaml`, so that the training connects through it, but do not commit this change. ```yaml database: host: "127.0.0.1" From 1059a92ff63433ed671651e085549b54b23a3c6e Mon Sep 17 00:00:00 2001 From: Axel Huebl Date: Tue, 22 Sep 2026 13:36:58 -0700 Subject: [PATCH 09/13] Improve Deployment Intro --- docs/source/deployment.md | 4 +++- 1 file changed, 3 insertions(+), 1 deletion(-) diff --git a/docs/source/deployment.md b/docs/source/deployment.md index 8d069fca..f3f8cd69 100644 --- a/docs/source/deployment.md +++ b/docs/source/deployment.md @@ -1,6 +1,8 @@ # Deployment -Synapse is currently deployed using Docker images and NERSC services. +Synapse is typically deployed using Docker images, e.g., on Kubernetes. + +Below, we document our public deployment workflow (recipes currently in a `private repo `__) using NERSC services like `Spin `__. ## Build the dashboard image From dba19901a245f15c0300531f2ab229166ed325b5 Mon Sep 17 00:00:00 2001 From: Axel Huebl Date: Tue, 22 Sep 2026 13:53:04 -0700 Subject: [PATCH 10/13] AmSC IRI API Mention Co-authored-by: Axel Huebl --- docs/source/dashboard.md | 2 +- docs/source/deployment.md | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/source/dashboard.md b/docs/source/dashboard.md index d190bbf5..200135a8 100644 --- a/docs/source/dashboard.md +++ b/docs/source/dashboard.md @@ -251,7 +251,7 @@ The dashboard has three routes, reachable from the navigation drawer: The `Parameters` tab holds the displayed output selector, the input parameter controls, and the plot depth control. The `Optimization` tab holds the optimization controls. The `ML` tab holds the model controls and the calibration controls. -- `/hpc` ("HPC Connection"): NERSC Superfacility API credential and Perlmutter status panel. +- `/hpc` ("HPC Connection"): Genesis AmSC IRI API or NERSC Superfacility API credential and HPC status panel. - `/chat` ("AI Assistant"): embedded assistant for experiment support. It loads [synapse-chat.lbl.gov](https://synapse-chat.lbl.gov/), which is hardcoded in {repo}`dashboard/app.py`. diff --git a/docs/source/deployment.md b/docs/source/deployment.md index f3f8cd69..686e77bf 100644 --- a/docs/source/deployment.md +++ b/docs/source/deployment.md @@ -33,7 +33,7 @@ python publish_container.py --gui --ml ## NERSC deployment assumptions - Dashboard runs on Spin. -- Training and simulations run on Perlmutter through Superfacility API. +- Training and simulations run on Perlmutter through Genesis AmSC IRI API or NERSC Superfacility API. - Images are pushed to the registry of the deployment's NERSC project: ::::{tab-set} From ab5d9cc1c83e337ef91894845fcd21ab7d937452 Mon Sep 17 00:00:00 2001 From: Edoardo Zoni <59625522+EZoni@users.noreply.github.com> Date: Wed, 7 Oct 2026 14:49:52 -0700 Subject: [PATCH 11/13] Apply suggestions from code review Co-authored-by: Edoardo Zoni <59625522+EZoni@users.noreply.github.com> --- docs/source/dashboard.md | 8 ++++---- docs/source/getting-started.md | 6 +++--- docs/source/ml-training.md | 6 +++--- docs/source/simulations.md | 2 +- 4 files changed, 11 insertions(+), 11 deletions(-) diff --git a/docs/source/dashboard.md b/docs/source/dashboard.md index 200135a8..d3733517 100644 --- a/docs/source/dashboard.md +++ b/docs/source/dashboard.md @@ -86,8 +86,8 @@ conda-lock install --name synapse-gui environment-lock.yml :sync: bella-nersc ```bash - export SF_DB_READONLY_PASSWORD='your_password_here' # Use SINGLE quotes around the password! - export AM_SC_API_KEY='your_amsc_api_key_here' # Required when MLflow tracking_uri is AmSC + export SF_DB_READONLY_PASSWORD='' # Use SINGLE quotes around the password! + export AM_SC_API_KEY='' # Required when MLflow tracking_uri is AmSC ``` ::: :::: @@ -156,11 +156,11 @@ conda-lock install --name synapse-gui environment-lock.yml :sync: bella-nersc ```bash - docker run --network=host -v /etc/localtime:/etc/localtime -v $PWD/ml:/app/ml -e SF_DB_READONLY_PASSWORD='your_password_here' -e AM_SC_API_KEY='your_amsc_api_key_here' synapse-gui + docker run --network=host -v /etc/localtime:/etc/localtime -v $PWD/ml:/app/ml -e SF_DB_READONLY_PASSWORD='' -e AM_SC_API_KEY='' synapse-gui ``` For debugging, you can enter the container without starting the app: ```bash - docker run --network=host -v /etc/localtime:/etc/localtime -v $PWD/ml:/app/ml -e SF_DB_READONLY_PASSWORD='your_password_here' -e AM_SC_API_KEY='your_amsc_api_key_here' -it synapse-gui bash + docker run --network=host -v /etc/localtime:/etc/localtime -v $PWD/ml:/app/ml -e SF_DB_READONLY_PASSWORD='' -e AM_SC_API_KEY='' -it synapse-gui bash ``` ::: :::: diff --git a/docs/source/getting-started.md b/docs/source/getting-started.md index 8578f225..dd08454d 100644 --- a/docs/source/getting-started.md +++ b/docs/source/getting-started.md @@ -61,7 +61,7 @@ Assuming `conda-lock` is installed in your conda `base` environment, from {repo- conda activate base conda-lock install --name synapse-gui environment-lock.yml conda activate synapse-gui -export ='...' +export ='' # Use SINGLE quotes around the password! export ='...' python -u app.py --port 8080 ``` @@ -74,7 +74,7 @@ python -u app.py --port 8080 conda activate base conda-lock install --name synapse-gui environment-lock.yml conda activate synapse-gui -export SF_DB_READONLY_PASSWORD='...' +export SF_DB_READONLY_PASSWORD='' # Use SINGLE quotes around the password! export AM_SC_API_KEY='...' python -u app.py --port 8080 ``` @@ -114,7 +114,7 @@ python train_model.py --test --config_file ../experiments/synapse-/c conda activate base conda-lock install --name synapse-ml environment-lock.yml conda activate synapse-ml -export SF_DB_READONLY_PASSWORD='...' +export SF_DB_READONLY_PASSWORD='' # Use SINGLE quotes around the password! export AM_SC_API_KEY='...' python train_model.py --test --config_file ../experiments/synapse-/config.yaml --model NN ``` diff --git a/docs/source/ml-training.md b/docs/source/ml-training.md index 9c17fa7e..3610bc43 100644 --- a/docs/source/ml-training.md +++ b/docs/source/ml-training.md @@ -88,7 +88,7 @@ This section describes how to train ML models locally. :sync: bella-nersc ```bash - export SF_DB_READONLY_PASSWORD='your_password_here' # Use SINGLE quotes around the password! + export SF_DB_READONLY_PASSWORD='' # Use SINGLE quotes around the password! export AM_SC_API_KEY='your_amsc_api_key_here' # Required when MLflow tracking_uri is AmSC ``` ::: @@ -198,7 +198,7 @@ This section describes how to train ML models at NERSC. :sync: bella-nersc ```bash - export SF_DB_READONLY_PASSWORD='your_password_here' # Use SINGLE quotes around the password! + export SF_DB_READONLY_PASSWORD='' # Use SINGLE quotes around the password! export AM_SC_API_KEY='your_amsc_api_key_here' # Required when MLflow tracking_uri is AmSC ``` ::: @@ -245,7 +245,7 @@ The Docker image is pulled from the [NERSC registry](https://registry.nersc.gov) :sync: bella-nersc ```text - SF_DB_READONLY_PASSWORD=your_password_here + SF_DB_READONLY_PASSWORD= AM_SC_API_KEY=your_amsc_api_key_here ``` ::: diff --git a/docs/source/simulations.md b/docs/source/simulations.md index 91749c57..b00b41f2 100644 --- a/docs/source/simulations.md +++ b/docs/source/simulations.md @@ -29,7 +29,7 @@ Before submission, it writes the current dashboard parameters to `single_simulat ``` ```{note} - This path is hardcoded for the BELLA deployment in {repo}`dashboard/parameters_manager.py`. + This path is hardcoded for the BELLA deployment in {repo}`dashboard/parameters_manager.py` and {repo}`dashboard/model_manager.py`. Other projects need to adapt it before launching simulations from the dashboard. ``` From 211f79e8467c1bfe497a6be4cb9f22c4934c54e4 Mon Sep 17 00:00:00 2001 From: Edoardo Zoni <59625522+EZoni@users.noreply.github.com> Date: Thu, 8 Oct 2026 11:15:51 -0700 Subject: [PATCH 12/13] Fix links in deployment.md --- docs/source/deployment.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/source/deployment.md b/docs/source/deployment.md index 686e77bf..73ee8071 100644 --- a/docs/source/deployment.md +++ b/docs/source/deployment.md @@ -2,7 +2,7 @@ Synapse is typically deployed using Docker images, e.g., on Kubernetes. -Below, we document our public deployment workflow (recipes currently in a `private repo `__) using NERSC services like `Spin `__. +Below, we document our public deployment workflow (recipes currently in a [private repository](https://github.com/BLAST-AI-ML/synapse-kubernetes-nersc)) using NERSC services like [Spin](https://docs.nersc.gov/services/spin/). ## Build the dashboard image From 4df995f4787d7f84d7f6417cf685f0a32a858188 Mon Sep 17 00:00:00 2001 From: Edoardo Zoni Date: Thu, 8 Oct 2026 11:20:22 -0700 Subject: [PATCH 13/13] Consistent placeholder style --- docs/source/dashboard.md | 8 ++++---- docs/source/ml-training.md | 12 ++++++------ 2 files changed, 10 insertions(+), 10 deletions(-) diff --git a/docs/source/dashboard.md b/docs/source/dashboard.md index d3733517..68694b84 100644 --- a/docs/source/dashboard.md +++ b/docs/source/dashboard.md @@ -77,8 +77,8 @@ conda-lock install --name synapse-gui environment-lock.yml :sync: general ```bash - export ='your_password_here' # Use SINGLE quotes around the password! - export ='your_api_key_here' # Required when MLflow tracking_uri is AmSC + export ='' # Use SINGLE quotes around the password! + export ='' # Required when MLflow tracking_uri is AmSC ``` ::: @@ -144,11 +144,11 @@ conda-lock install --name synapse-gui environment-lock.yml :sync: general ```bash - docker run --network=host -v /etc/localtime:/etc/localtime -v $PWD/ml:/app/ml -e ='your_password_here' -e ='your_api_key_here' synapse-gui + docker run --network=host -v /etc/localtime:/etc/localtime -v $PWD/ml:/app/ml -e ='' -e ='' synapse-gui ``` For debugging, you can enter the container without starting the app: ```bash - docker run --network=host -v /etc/localtime:/etc/localtime -v $PWD/ml:/app/ml -e ='your_password_here' -e ='your_api_key_here' -it synapse-gui bash + docker run --network=host -v /etc/localtime:/etc/localtime -v $PWD/ml:/app/ml -e ='' -e ='' -it synapse-gui bash ``` ::: diff --git a/docs/source/ml-training.md b/docs/source/ml-training.md index 3610bc43..992762d8 100644 --- a/docs/source/ml-training.md +++ b/docs/source/ml-training.md @@ -79,8 +79,8 @@ This section describes how to train ML models locally. :sync: general ```bash - export ='your_password_here' # Use SINGLE quotes around the password! - export ='your_api_key_here' # Required when MLflow tracking_uri is AmSC + export ='' # Use SINGLE quotes around the password! + export ='' # Required when MLflow tracking_uri is AmSC ``` ::: @@ -89,7 +89,7 @@ This section describes how to train ML models locally. ```bash export SF_DB_READONLY_PASSWORD='' # Use SINGLE quotes around the password! - export AM_SC_API_KEY='your_amsc_api_key_here' # Required when MLflow tracking_uri is AmSC + export AM_SC_API_KEY='' # Required when MLflow tracking_uri is AmSC ``` ::: :::: @@ -189,8 +189,8 @@ This section describes how to train ML models at NERSC. :sync: general ```bash - export ='your_password_here' # Use SINGLE quotes around the password! - export ='your_api_key_here' # Required when MLflow tracking_uri is AmSC + export ='' # Use SINGLE quotes around the password! + export ='' # Required when MLflow tracking_uri is AmSC ``` ::: @@ -199,7 +199,7 @@ This section describes how to train ML models at NERSC. ```bash export SF_DB_READONLY_PASSWORD='' # Use SINGLE quotes around the password! - export AM_SC_API_KEY='your_amsc_api_key_here' # Required when MLflow tracking_uri is AmSC + export AM_SC_API_KEY='' # Required when MLflow tracking_uri is AmSC ``` ::: ::::