Skip to content
Merged
6 changes: 3 additions & 3 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -90,7 +90,7 @@ docker run -p 127.0.0.1:5000:5000 ghcr.io/mlflow/mlflow mlflow server --host 0.0
python tests/test_ml_pipeline.py

# Optionally restrict to a specific model type or config
python tests/test_ml_pipeline.py --model NN --config_file experiments/synapse-bella-ip2
python tests/test_ml_pipeline.py --model NN --config_file experiments/synapse-<experiment>/config.yaml
```

Dashboard validation is done manually by running the application.
Expand Down Expand Up @@ -121,8 +121,8 @@ The only CI workflow is **CodeQL Advanced** (`.github/workflows/codeql.yml`), wh

- **MongoDB** is used for persistent data from experiments and simulations.
- **MLflow** is used for persistent data from ML models.
- Database access requires SSH tunneling to NERSC when running locally.
- Environment variables: `SF_DB_HOST` (dashboard), `SF_DB_READONLY_PASSWORD` (dashboard and ML training), `AM_SC_API_KEY` (dashboard and ML training, required when MLflow tracking_uri is AmSC).
- Database connection settings are read from the `database` section of the experiment's `config.yaml`. Local runs usually reach the database through an SSH tunnel, with `database.host` set to `127.0.0.1` in a local copy of `config.yaml`.
- Environment variables: `config.yaml` sets their names. `database.password_ro_env` names the read-only database password, and `mlflow.api_key_env` names the MLflow API key, which is only needed when `mlflow.tracking_uri` is AmSC. Both are used by the dashboard and by ML training. The BELLA deployment uses `SF_DB_READONLY_PASSWORD` and `AM_SC_API_KEY`.

## Common Pitfalls and Workarounds

Expand Down
1 change: 1 addition & 0 deletions docs/docs.yml
Original file line number Diff line number Diff line change
Expand Up @@ -11,3 +11,4 @@ dependencies:
- sphinx-autobuild
- sphinx-book-theme
- sphinx-copybutton
- sphinx-design
11 changes: 10 additions & 1 deletion docs/source/conf.py
Original file line number Diff line number Diff line change
Expand Up @@ -19,8 +19,17 @@
# -- General configuration ---------------------------------------------------
# https://www.sphinx-doc.org/en/master/usage/configuration.html#general-configuration

extensions = ["myst_parser", "sphinx.ext.extlinks", "sphinx_copybutton"]
extensions = [
"myst_parser",
"sphinx.ext.extlinks",
"sphinx_copybutton",
"sphinx_design",
]
myst_heading_anchors = 4
# ":::" fences, used to nest the sphinx-design tab sets. All tab sets use the
# sync group "deployment" with the keys "general" (default) and "bella-nersc",
# so a selected tab applies to all pages.
myst_enable_extensions = ["colon_fence"]

# Roles that turn a repository-relative path into a link to GitHub, so that
# source files and directories mentioned in the docs stay navigable:
Expand Down
192 changes: 173 additions & 19 deletions docs/source/dashboard.md
Original file line number Diff line number Diff line change
Expand Up @@ -38,26 +38,66 @@ conda-lock install --name synapse-gui environment-lock.yml

#### Run the dashboard

1. Create an SSH tunnel to access the MongoDB database at NERSC (in a separate terminal):
1. If the computer running the dashboard cannot directly reach the host specified by `database.host`, create an SSH tunnel to the MongoDB database through a gateway node in a separate terminal, using `database.host` and `database.port` from your experiment's `config.yaml`:

::::{tab-set}
:sync-group: deployment

:::{tab-item} General
:sync: general

```bash
ssh -L <database.port>:<database.host>:<database.port> <username>@<gateway_host> -N
```
:::

:::{tab-item} Project Example: BELLA @ NERSC
:sync: bella-nersc

```bash
ssh -L 27017:mongodb05.nersc.gov:27017 <username>@dtn03.nersc.gov -N
```
:::
::::

2. If you created the SSH tunnel, set `database.host` to `127.0.0.1` in your local copy of `config.yaml`, so that the dashboard connects through it, but do not commit this change.
```yaml
database:
host: "127.0.0.1"
```

3. Move to the {repo-dir}`dashboard/` directory.

4. Set up the read-only database password and the MLflow API key:

::::{tab-set}
:sync-group: deployment

:::{tab-item} General
:sync: general

```bash
export <database.password_ro_env>='<your_password_here>' # Use SINGLE quotes around the password!
export <mlflow.api_key_env>='<your_api_key_here>' # Required when MLflow tracking_uri is AmSC
```
:::

2. Move to the {repo-dir}`dashboard/` directory.
:::{tab-item} Project Example: BELLA @ NERSC
:sync: bella-nersc

3. Set up the database settings (read-only) and the AmSC MLflow API key:
```bash
export SF_DB_HOST='127.0.0.1'

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Removed: went away in a recent PR.
We still need to remove references to it in a few other locations: https://github.com/search?q=repo%3ABLAST-AI-ML%2Fsynapse+SF_DB_HOST&type=code

export SF_DB_READONLY_PASSWORD='your_password_here' # Use SINGLE quotes around the password!
export AM_SC_API_KEY='your_amsc_api_key_here' # Required when MLflow tracking_uri is AmSC
export SF_DB_READONLY_PASSWORD='<your_password_here>' # Use SINGLE quotes around the password!
export AM_SC_API_KEY='<your_amsc_api_key_here>' # Required when MLflow tracking_uri is AmSC
```
:::
::::

4. Activate the conda environment `synapse-gui`:
5. Activate the conda environment `synapse-gui`:
```bash
conda activate synapse-gui
```

5. Run the dashboard as a web application:
6. Run the dashboard as a web application:
```bash
python -u app.py --port 8080
```
Expand All @@ -66,28 +106,87 @@ conda-lock install --name synapse-gui environment-lock.yml

#### Run the container

1. Create an SSH tunnel to access the MongoDB database at NERSC (in a separate terminal):
1. If the computer running the dashboard cannot directly reach the host specified by `database.host`, create an SSH tunnel to the MongoDB database through a gateway node in a separate terminal, using `database.host` and `database.port` from your experiment's `config.yaml`:

::::{tab-set}
:sync-group: deployment

:::{tab-item} General
:sync: general

```bash
ssh -L <database.port>:<database.host>:<database.port> <username>@<gateway_host> -N
```
:::

:::{tab-item} Project Example: BELLA @ NERSC
:sync: bella-nersc

```bash
ssh -L 27017:mongodb05.nersc.gov:27017 <username>@dtn03.nersc.gov -N
```
:::
::::

2. Move to the root directory of the repository.
2. If you created the SSH tunnel, set `database.host` to `127.0.0.1` in your local copy of `config.yaml`, but do not commit this change.
Do this before building the image, because {repo}`dashboard.Dockerfile` copies {repo-dir}`experiments/` into it.

3. Move to the root directory of the repository.

4. Build the Docker image as described [below](#build-the-docker-image).

5. Run the Docker container with the read-only database password and the MLflow API key:

::::{tab-set}
:sync-group: deployment

:::{tab-item} General
:sync: general

```bash
docker run --network=host -v /etc/localtime:/etc/localtime -v $PWD/ml:/app/ml -e <database.password_ro_env>='<your_password_here>' -e <mlflow.api_key_env>='<your_api_key_here>' synapse-gui
```
For debugging, you can enter the container without starting the app:
```bash
docker run --network=host -v /etc/localtime:/etc/localtime -v $PWD/ml:/app/ml -e <database.password_ro_env>='<your_password_here>' -e <mlflow.api_key_env>='<your_api_key_here>' -it synapse-gui bash
```
:::

3. Build the Docker image as described [below](#build-the-docker-image).
:::{tab-item} Project Example: BELLA @ NERSC
:sync: bella-nersc

4. Run the Docker container:
```bash
docker run --network=host -v /etc/localtime:/etc/localtime -v $PWD/ml:/app/ml -e SF_DB_HOST='127.0.0.1' -e SF_DB_READONLY_PASSWORD='your_password_here' -e AM_SC_API_KEY='your_amsc_api_key_here' synapse-gui
docker run --network=host -v /etc/localtime:/etc/localtime -v $PWD/ml:/app/ml -e SF_DB_READONLY_PASSWORD='<your_password_here>' -e AM_SC_API_KEY='<your_amsc_api_key_here>' synapse-gui
```
For debugging, you can enter the container without starting the app:
```bash
docker run --network=host -v /etc/localtime:/etc/localtime -v $PWD/ml:/app/ml -e SF_DB_HOST='127.0.0.1' -e SF_DB_READONLY_PASSWORD='your_password_here' -e AM_SC_API_KEY='your_amsc_api_key_here' -it synapse-gui bash
docker run --network=host -v /etc/localtime:/etc/localtime -v $PWD/ml:/app/ml -e SF_DB_READONLY_PASSWORD='<your_password_here>' -e AM_SC_API_KEY='<your_amsc_api_key_here>' -it synapse-gui bash
```
:::
::::

Note that `-v /etc/localtime:/etc/localtime` is necessary to synchronize the time zone in the container with the host machine.

## Run the dashboard at NERSC

Connect to the [dashboard](https://bellasuperfacility.lbl.gov/) deployed at NERSC through Spin and explore it.
Connect to the dashboard deployed for your project at NERSC through Spin and explore it:

::::{tab-set}
:sync-group: deployment

:::{tab-item} General
:sync: general

Open `https://<dashboard_url>/`, the URL of your project's Spin deployment.
:::

:::{tab-item} Project Example: BELLA @ NERSC
:sync: bella-nersc

Open [bellasuperfacility.lbl.gov](https://bellasuperfacility.lbl.gov/).
:::
::::

You need to upload valid Superfacility API credentials before you can launch simulations or train ML models directly from the dashboard.

## Generate Superfacility API credentials
Expand All @@ -100,7 +199,24 @@ Follow the instructions at [docs.nersc.gov/services/sfapi/authentication/#client

3. Scroll down to the section "Superfacility API Clients" and click "New Client".

4. Enter a client name (e.g., "Synapse"), choose `sf558` for the user, choose "Red" security level, and select either "Your IP" or "Spin" from the "IP Presets" menu, depending on whether the key will be used from a local computer or from Spin.
4. Enter a client name (e.g., "Synapse"), choose the user that runs the jobs launched from the dashboard, choose "Red" security level, and select either "Your IP" or "Spin" from the "IP Presets" menu, depending on whether the key will be used from a local computer or from Spin.
ML training jobs expect credential files in the `$HOME` of that user, as described in [Through the dashboard](ml-training.md#through-the-dashboard).

::::{tab-set}
:sync-group: deployment

:::{tab-item} General
:sync: general

Choose your own NERSC user or a collaboration account of your NERSC project.
:::

:::{tab-item} Project Example: BELLA @ NERSC
:sync: bella-nersc

Choose the collaboration account `sf558`.
:::
::::

5. Download the private key file in PEM format and save it as `priv_key.pem` in the root directory of the dashboard.
Each time the dashboard is launched, it will automatically find the existing key file and load the corresponding credentials.
Expand Down Expand Up @@ -135,8 +251,9 @@ The dashboard has three routes, reachable from the navigation drawer:
The `Parameters` tab holds the displayed output selector, the input parameter controls, and the plot depth control.
The `Optimization` tab holds the optimization controls.
The `ML` tab holds the model controls and the calibration controls.
- `/hpc` ("HPC Connection"): NERSC Superfacility API credential and Perlmutter status panel.
- `/chat` ("AI Assistant"): embedded assistant route for experiment support; currently backed by [synapse-chat.lbl.gov](https://synapse-chat.lbl.gov/).
- `/hpc` ("HPC Connection"): Genesis AmSC IRI API or NERSC Superfacility API credential and HPC status panel.
- `/chat` ("AI Assistant"): embedded assistant for experiment support.
It loads [synapse-chat.lbl.gov](https://synapse-chat.lbl.gov/), which is hardcoded in {repo}`dashboard/app.py`.

The experiment selector, the date range selector, and the error panel belong to the shared layout rather than to any single route, so they appear on all three.

Expand Down Expand Up @@ -177,6 +294,8 @@ Run this workflow automatically with the Python script {repo}`publish_container.
```bash
python publish_container.py --gui
```
The script pushes to `registry.nersc.gov/m558/superfacility`, which is hardcoded for the BELLA deployment.
For other projects, follow the steps below.
````

````{tip}
Expand Down Expand Up @@ -206,16 +325,51 @@ docker system prune -a
# Password: your NERSC password without 2FA
```

3. Tag the Docker image:
3. Tag the Docker image with the registry path of your NERSC project:

::::{tab-set}
:sync-group: deployment

:::{tab-item} General
:sync: general

```bash
docker tag synapse-gui:latest registry.nersc.gov/<nersc_project>/[<namespace>/]synapse-gui:latest
docker tag synapse-gui:latest registry.nersc.gov/<nersc_project>/[<namespace>/]synapse-gui:$(date "+%y.%m")
```
:::

:::{tab-item} Project Example: BELLA @ NERSC
:sync: bella-nersc

```bash
docker tag synapse-gui:latest registry.nersc.gov/m558/superfacility/synapse-gui:latest
docker tag synapse-gui:latest registry.nersc.gov/m558/superfacility/synapse-gui:$(date "+%y.%m")
```
:::
::::

4. Push the Docker image:

::::{tab-set}
:sync-group: deployment

:::{tab-item} General
:sync: general

```bash
docker push -a registry.nersc.gov/<nersc_project>/[<namespace>/]synapse-gui
```
:::

:::{tab-item} Project Example: BELLA @ NERSC
:sync: bella-nersc

```bash
docker push -a registry.nersc.gov/m558/superfacility/synapse-gui
```
:::
::::

## References

Expand Down
31 changes: 28 additions & 3 deletions docs/source/deployment.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,8 @@
# Deployment

Synapse is deployed using Docker images and NERSC services.
Synapse is typically deployed using Docker images, e.g., on Kubernetes.

Below, we document our public deployment workflow (recipes currently in a [private repository](https://github.com/BLAST-AI-ML/synapse-kubernetes-nersc)) using NERSC services like [Spin](https://docs.nersc.gov/services/spin/).

## Build the dashboard image

Expand Down Expand Up @@ -31,6 +33,29 @@ python publish_container.py --gui --ml
## NERSC deployment assumptions

- Dashboard runs on Spin.
- Training and simulations run on Perlmutter through Superfacility API.
- Images are pushed to `registry.nersc.gov/m558/superfacility`.
- Training and simulations run on Perlmutter through Genesis AmSC IRI API or NERSC Superfacility API.
- Images are pushed to the registry of the deployment's NERSC project:

::::{tab-set}
:sync-group: deployment

:::{tab-item} General
:sync: general

```text
registry.nersc.gov/<nersc_project>/[<namespace>/]
```
:::

:::{tab-item} Project Example: BELLA @ NERSC
:sync: bella-nersc

```text
registry.nersc.gov/m558/superfacility/
```
:::
::::

{repo}`publish_container.py` hardcodes the BELLA path.
For other projects, tag and push the images manually, as described in [Dashboard](dashboard.md#push-the-docker-image) and [ML training](ml-training.md#push-the-docker-image).
- Before publishing, validate the images locally.
4 changes: 2 additions & 2 deletions docs/source/developer-notes.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,8 +12,8 @@ Ruff runs with its default rule set; there is no `pyproject.toml` or `ruff.toml`

## Conda environments

- Dashboard dependencies live in {repo}`dashboard/environment.yml`.
- ML dependencies live in {repo}`ml/environment.yml`.
- Dashboard dependencies are defined in {repo}`dashboard/environment.yml`.
- ML dependencies are defined in {repo}`ml/environment.yml`.
- Regenerate the corresponding `environment-lock.yml` after dependency changes.

## Build the documentation
Expand Down
Loading
Loading