Skip to content

Online Deployment Tutorial

Will Benoit edited this page Mar 3, 2026 · 16 revisions

Aframe Account

Logging In

First, ensure that your local ssh key has been added to your ssh agent:

ssh-add <path/to/sshkey>

Next, login to a node on CIT with ssh agent forwarding enabled:

ssh -A marie.curie@ldas-grid.ligo.caltech.edu

Once logged in, ssh to the aframe node with as the aframe user:

ssh aframe.online@aframe.ldas.cit

Directory Structure

Once logged in, you can ls to see the directory structure:

(base) [aframe.online@aframe ~]$ pwd
/home/aframe.online
(base) [aframe.online@aframe ~]$ ls
aframe  cilogon_cert  dev  images  miniconda3  production  public_html  robot  software
(base) [aframe.online@aframe ~]$ 

Here's a highlevel summary of the directory contents:

  • ~/robot: Aframes kerberos keytab is located for passwordless refreshing of scitokens
  • ~/images: Apptainer images for aframe and amplfi
  • ~/dev: Development analyses (i.e. not being used to upload triggers to gracedb.ligo.org)
  • ~/production: Production analyses
  • ~/software: Github repositories for aframe and amplfi

Additionally, the scitoken used for authentication during the search is stored at /local/aframe/scitoken. The scitoken is automatically refreshed by the online pipeline, but its current status can be examined with httokendecode -H /local/aframe/scitoken.

Initializing an Online Run

The easiest way to initialize a new online run is with the aframe-init command line tool that is installed in the base aframe environment:

(base) [aframe.online@aframe ~]$ pwd
/home/aframe.online
(base) [aframe.online@aframe ~]$ cd software/aframe/
(base) [aframe.online@aframe aframe]$ uv run aframe-init online --help
usage: aframe-init [options] online [-h] [-d DIRECTORY] [--weights-dir WEIGHTS_DIR]

optional arguments:
  -h, --help            Show this help message and exit.
  -d DIRECTORY, --directory DIRECTORY
                        (required, type: <class 'Path'>)
  --weights-dir WEIGHTS_DIR
                        (type: <class 'Path'>, default: null)
(base) [aframe.online@aframe aframe]$ uv run aframe-init online -d /home/aframe.online/dev/online-tutorial
(base) [aframe.online@aframe aframe]$ cd /home/aframe.online/dev/online-tutorial
(base) [aframe.online@aframe online-tutorial]]$ ls
config.yaml  crontab  prior.yaml  run.sh

This will create a new directory for you run at the specified path. You should notice 4 files in the directory:

  • config.yaml: Main configuration file for the run
  • prior.yaml: Prior for filtering parameters produced by AMPLFI (this is referenced within config.yaml)
  • run.sh: Bash script for launching the run
  • crontab: Cron file for automatically refreshing credentials, and relaunching pipeline when system reboots

Setting Up the Run Directory

Running aframe-init creates the run directory and some of the necessary files, but other necessary components are not (yet) automatically created.

Model weights

The paths to the model weights need to be specified in the run.sh file. Traditionally, we've placed the weights in a directory named models and named them aframe.pt, amplfi-hl.ckpt, and amplfi-hlv.ckpt. These correspond to the AFRAME_WEIGHTS, AMPLFI_HL_WEIGHTS, and AMPLFI_HLV_WEIGHTS environment variables.

Background and foreground data

The files containing the background distribution, foreground distribution, and rejected testing waveform parameters must be specified in the run.sh file. These are usually placed in a data directory and named background.hdf5, foreground.hdf5, and rejected-parameters.hdf5. These correspond to the ONLINE_BACKGROUND_FILE, ONLINE_FOREGROUND_FILE, and ONLINE_REJECTED_FILE environment variables. It's crucial that these files were generated by the aframe.pt model mentioned above.

Completing the run.sh file

In addition to the model and data paths described above, there are a few more environment variables that must be specified in the run.sh file:

  • BEARER_TOKEN_FILE: Location of scitoken for authentication. This should be somewhere in /local/aframe.online; typically we use /local/aframe.online/scitoken
  • SCITOKEN_FILE: Must match BEARER_TOKEN_FILE
  • AFRAME_ONLINE_OUTDIR: Location of logs and events. This must be <run_dir>/output; e.g. /home/aframe/dev/o4_llpic_playground/output. In the future, this will be done automatically.
  • AFRAME_CONTAINER_ROOT: Path to directory with the container for the online search. Currently, the containers are stored in /home/aframe/images/aframe.
  • CUDA_VISIBLE_DEVICES: There are two GPUs available on the aframe node. GPU 0 is reserved for Aframe, but GPU 1 is part of the condor pool and may be used by other jobs. Any LLPIC runs (and in the future, production runs) should have this variable set to 0. Other search instances should set this to 1

Updating the config.yaml file

The config.yaml file contains the model configuration parameters for the online search and should be updated to reflect how the models were trained/tested. In particular, loading the AMPLFI models requires specifying the exact architecture, and this may change from model to model. Most other parameters can remain unchanged, but here are some important ones to verify before running:

  • channels: The Hanford, Livingston, and Virgo data channel to read from each frame file. This will change depending on whether we're analyzing MDC, LLPIC, or live data.
  • ifo_suffix: This value controls which data directory within /dev/shm/kakfa to read from, and changes depending on whether we're analyzing MDC, LLPIC, or live data.
  • far_threshold: The false alarm rate, in units of inverse days, below which we'll upload an event. This controls how frequently we upload events to GraceDB on average, and should typically be set to 12, but for testing purposes it can be useful to set a higher rate.
  • emails: A list of emails to send notifications to. Recipients will be alerted when the process starts, when an event is detected below email_far_threshold, and when an error occurs and the pipeline crashes.
  • server: The GraceDB server to which to upload events. Currently, this should be either playground, test, test01, or local, with playground being reserved for LLPIC uploads. Setting to local will just write out data products locally and is useful for testing.

Verifying Pipeline Activity

a. Viewing the run summary page

Summary pages for individual events and the search as a whole are created automatically during the run, and are written to the directory specified by the MONITOR_OUTDIR environment variable in the run.sh file, e.g.:

export MONITOR_OUTDIR=/home/aframe.online/public_html/o4_llpic_playground/

Visiting https://ldas-jobs.ligo.caltech.edu/~aframe.online/ and navigating to the directory given by MONITOR_OUTDIR allows for viewing these pages in a browser. At the top level of this directory is a summary.html page that indicates whether the search is active and whether the data is in analysis-ready mode. Also note that the UTC time of the last update is listed at the top of the page. If the page hasn't been refreshed in a while, the status information may be out of date.

image

Summary pages for individual events can be found in the events directory and will be discussed below.

b. Check for Active tmux Session

If the summary hasn't been updated, it could be that the search has crashed, but it's possible that only the monitoring process failed. You can connect to the aframe and look to see whether the search is still running in the tmux session in which it was started.

To view active tmux sessions:

(base) [aframe.online@aframe ~]$ tmux ls
llpic-test: 1 windows (created Tue Jul  8 07:49:32 2025) [272x86]

To attach to a session:

tmux attach -t llpic-test

Once connected, it should be clear whether the process is running as output logs are continuously being written to STDOUT:

2025-07-09 17:57:07,221 (17:57:07,221) - root - DEBUG - Read successful
2025-07-09 17:57:07,227 (17:57:07,228) - root - DEBUG - Updating input buffer
2025-07-09 17:57:07,228 (17:57:07,228) - root - DEBUG - Updating snapshotter
2025-07-09 17:57:07,228 (17:57:07,228) - root - DEBUG - Whitening data
2025-07-09 17:57:07,230 (17:57:07,230) - root - DEBUG - Performing inference
2025-07-09 17:57:07,234 (17:57:07,234) - root - DEBUG - Updating output buffer
2025-07-09 17:57:07,310 (17:57:07,310) - root - DEBUG - Searching for event...
2025-07-09 17:57:07,310 (17:57:07,311) - root - DEBUG - Reading frames from timestamp 1436119045

c. Check Running Processes

If there are no active tmux sessions, check whether there are the expected number of online processes

(base) [aframe.online@aframe ~]$ pgrep online | wc -l
6

If the process count is 0, the search is not running, and again be restarted; see the section below on how to do so. If the process count is more than 0 but less than 6, it likely means that something is going wrong and the processes should be killed and the pipeline restarted.

Log Files

There are multiple log files that capture the output of different parts of the search. In the top level of the run directory, there are monitoring.log and summary_pages.log files. Additionally, in outputs/logs, there are directories that correspond to each instance of the search being started. These directories contain daily log files, with the active file always being online.log.

(base) [aframe.online@aframe logs]$ pwd
/home/aframe/dev/o4_llpic_playground/output/logs
(base) [aframe.online@aframe logs]$ ls
2025-07-08T15:03:21  2025-07-08T15:09:16  2025-07-08T21:08:18
(base) [aframe.online@aframe logs]$ cd 2025-07-08T21:08:18
(base) [aframe.online@aframe 2025-07-08T21:08:18]$ ls
online.log  online.log.2025-07-08
(base) [aframe.online@aframe 2025-07-08T21:08:18]$

Log file descriptions

Log file name Description When to examine
monitoring.log Tracks pipeline restarts and start-up errors Search failing to start
summary_pages.log Logging from the summary html pages process Summary pages failing
online.log Main search logging statements Search crash, investigation of search behavior for an event
online.log.<date> Historical contents from online.log Investigation of search behavior for an event

Summary pages

The output directory of the monitoring process is structured as follows:

outdir/
├── event_data.hdf5
├── summary.html
├── monitor.log
├── plots/
│   ├── aframe_latency.png
│   └── ...
└── events/
    ├── event_1436032137/
    │   ├── event.html
    │   └── plots/
    │       ├── H1_qtransform.png
    │       └── ...
    ├── event_1436033156/
    └── ...

At the top level, there is a summary.html page that contains information about the status of the run, as discussed above. Additionally, there is a histogram of event reporting latency. Typical latencies for MDCs should be less than 5 seconds, and less than 15 seconds for production runs. Events with latencies greater than this are worth investigating. The event_data.hdf5 file information on a per-event basis and can be a good place to look to get started. Also on the summary page are event rate plots. When running on non-injected data streams, we want to see that the event rate roughly stays at the expected level of 1 per 2 hours. All of the individual plots can be found in the plots directory at this level.

Inside the events directory is a directory for each detected event. Each event has an event.html page which shows Q-plots for each active detector, Aframe's detection statistic timeseries, AMPLFI's skymaps and corner plot, and the ASD from each active detector. At the top of each event page is a link to the event in GraceDB. All of the individual plots are again located in the plots directory for the event, and these can be a good place to start to get an overview of how the algorithms performed.

image

Re-examining events offline

If Aframe is offline when an event occurs, or if there's a reason to re-analyze an event with a different configuration, buoy can be used to do so. In /home/aframe.online/buoy is a set of config files and models that match the current online run. To run buoy over a particular event, specify the event identifier (either the S-event, G-event, or GW-label) in the events parameter in the buoy.yaml file. Multiple events can be given as strings in a list.

The default config has to_html set to True and the outdir located in the /home/aframe.online/public_html/buoy_results directory, so the summary page for an analyzed event can be found here: https://ldas-jobs.ligo.caltech.edu/~aframe.online/buoy_results/.

Once the config files are set up as desired, run the following:

apptainer exec --nv /home/aframe.online/images/buoy/buoy.sif /opt/env/bin/buoy --config buoy.yaml