The cortex-training command and cortex_training Python package provide the
supported command-line and SDK interfaces for Cortex Training.
ct is installed as an alias for cortex-training and supports the same
commands and flags. You can replace cortex-training with ct in any example
below, such as ct login config.json or ct --job JOB_ID step.
Requires Python 3.10+. Installing the package gives you the cortex-training
CLI, the cortex-training tui log viewer, and the cortex_training Python SDK.
This project uses uv; pip works in place of
uv pip throughout if you prefer.
Install straight from the repository:
uv pip install git+https://github.com/snowflakedb/cortex-training.gitOr from a local checkout:
git clone https://github.com/snowflakedb/cortex-training.git
cd cortex-training
uv pip install . # add -e for an editable/dev installVerify the install:
cortex-training --help
cortex-training tui --helpA job is a server-side lifecycle resource for remote model workers, not a training script. Its sub-jobs define the model, GPU requirements, and training or sampling configuration. Submitting a job starts those workers; once the job is running, you send batches, optimizer steps, or generation requests to its job ID.
Jobs separate worker setup and GPU allocation from the code driving the experiment. You can reuse running workers across requests, inspect status and logs from another CLI process, and cancel the job to release capacity. A combined training and sampling job also lets an RL loop synchronize weights between its sub-jobs.
The CLI and CortexTrainingClient operate on the same jobs; they are not
alternative execution models. Use the CLI for setup, inspection, and individual
operations. Use a Python loop or a recipe for dataset
iteration, batching, rewards, evaluation, and repeated calls. The client loop
creates a job or uses an existing job ID, submits operations, and polls their
results. submit --wait waits for workers to be running; it does not upload or
execute your Python loop.
See jobs and sub-jobs and the Python SDK reference for details.
Global flags such as --config, --job, and --compact go before the
subcommand. login also accepts its own --config after the subcommand.
--job-id is an alias for --job:
cortex-training --config config.json list
cortex-training --job JOB_ID step
cortex-training checkpoints JOB_IDThe data-plane actions fwd-bwd, step, load, generate, and weight-sync
require the global --job JOB_ID option. Their --help usage lines show its
placement. Management commands such as get and checkpoints instead take a
positional JOB_ID after the subcommand.
cortex-training --connection training list # Use a Snowflake profile
cortex-training list # Use the configured default profile
cortex-training login config.json # Remember config for future commands
cortex-training login --config config.json # Equivalent login syntax
cortex-training --config config.json list # Use config for one commandcortex-training submit examples/api/training.json
cortex-training submit examples/api/sampling.json
cortex-training submit job.json --dry-run # Validate without submitting
cortex-training submit job.json --wait # Wait until running, not finished
cortex-training submit - < job.json # Read JSON from stdincortex-training list
cortex-training list --status running
cortex-training get JOB_ID
cortex-training wait JOB_ID # Wait until running, not finished
cortex-training cancel JOB_ID
cortex-training checkpoints JOB_ID
cortex-training capacity # All supported GPU types
cortex-training capacity --hardware B200cortex-training --job JOB_ID fwd-bwd examples/api/fwd-bwd.json
cortex-training --job JOB_ID step # Default learning rate: 1e-4
cortex-training --job JOB_ID step --lr 2e-5
cortex-training --job JOB_ID generate examples/api/generate.json
cortex-training --job JOB_ID weight-sync
cortex-training --job JOB_ID weight-sync --weight-format loraSee also generation payloads and weight-sync routing.
cortex-training --job JOB_ID load CHECKPOINT_ID
cortex-training --job JOB_ID load CHECKPOINT_ID --source-job-id SOURCE_JOB_ID
cortex-training --job JOB_ID load CHECKPOINT_ID --target-sub-job-id JOB_ID:training:0
cortex-training --job JOB_ID load CHECKPOINT_ID --no-pollcortex-training tui # Open job picker
cortex-training tui JOB_ID # Open job logs
cortex-training download-log JOB_ID --output-dir ./logs
cortex-training download-log JOB_ID --log-type stdout --output-dir ./logs
cortex-training download-metrics JOB_ID --output-dir ./metricsSee execution logs, persisted stdout, and GPU metrics for download paths and output fields.
cortex-training --compact list
cortex-training get JOB_ID | jq '.sub_jobs'
cortex-training --help
cortex-training fwd-bwd --help| Command | Default behavior | Alternative |
|---|---|---|
submit |
Return after submission | --wait waits until running; --dry-run validates without submitting |
wait JOB_ID |
Wait until running, not until training finishes | Use get JOB_ID to inspect current status |
fwd-bwd, generate |
Poll the submitted request until completion | Set top-level "poll": false in the input JSON |
step |
Poll until completion; learning rate 1e-4 |
Set --lr; polling cannot be disabled |
load, weight-sync |
Poll the submitted request until completion | --no-poll returns without waiting for the result |
capacity |
Query H200, B200, and B300 | Select one with --hardware |
download-log, download-metrics |
Write under the current directory | Set --output-dir |
- Connection config, login, and environment variables
- Submit, manage jobs, and GPU capacity
- Forward-backward and optimizer steps
- Load checkpoints and initialize sampling
- Generate and sync weights
- Download logs, stdout, and metrics
- Log TUI, JSON output and help, and troubleshooting
cortex-training submits and manages Cortex Training jobs through the Cortex
Training REST endpoint.
The Snowflake-native workflow uses the same connection profiles as the Python
Connector and Snowflake CLI:
# ~/.snowflake/connections.toml
[training]
account = "ORG-ACCOUNT"
host = "ACCOUNT.snowflakecomputing.com"
user = "USER"
authenticator = "programmatic_access_token"
token = "YOUR_PROGRAMMATIC_ACCESS_TOKEN"
database = "CORTEX_TRAINING_DB"
schema = "PUBLIC"Protect files containing credentials, then select the named profile:
chmod 600 ~/.snowflake/connections.toml
cortex-training --connection training listcortex-training list with no connection arguments uses the
Connector-configured default profile. Set
SNOWFLAKE_DEFAULT_CONNECTION_NAME=training, configure
default_connection_name = "training" in Snowflake's config.toml, or name
the connection [default].
Profile lookup is a fallback. An explicit or remembered legacy JSON config, or
a complete direct --base-url / --host + --pat connection, continues to
win. This preserves existing scripts.
For Snowflake PAT auth, use host for the account hostname. Do not use
base_url for Snowflake PAT auth.
{
"host": "ACCOUNT.snowflakecomputing.com",
"pat": "YOUR_PROGRAMMATIC_ACCESS_TOKEN",
"database": "CORTEX_TRAINING_DB",
"schema": "PUBLIC",
"endpoint": "cortex-training",
"poll_interval": 0.5,
"poll_timeout": 1800.0,
"verify_ssl": true
}If you prefer not to store the PAT in the file, omit pat and set it in the
shell instead:
export CORTEX_TRAINING_PAT='YOUR_PROGRAMMATIC_ACCESS_TOKEN'To target a local or otherwise compatible server instead of a Snowflake
account, use base_url with an explicit scheme. This skips PAT auth:
{
"base_url": "http://localhost:8084",
"database": "MY_DB",
"schema": "PUBLIC",
"endpoint": "cortex-training"
}Login validates the config and stores only the config path, not the config contents:
cortex-training login config.json
cortex-training login --config config.jsonProvide exactly one config path, either positionally or with --config after
login. Both forms validate and remember the same file. Login requires an
explicit path even when a global --config, CORTEX_TRAINING_CONFIG, or a
previous login is available.
The login state is written to ~/.config/cortex-training/login.json by default,
or $XDG_CONFIG_HOME/cortex-training/login.json when XDG_CONFIG_HOME is set.
You can bypass login for one command with:
cortex-training --config config.json listor by setting:
export CORTEX_TRAINING_CONFIG=/path/to/config.jsonExplicit CLI flags override config values.
cortex-training list
cortex-training list --status running
cortex-training get JOB_ID
cortex-training checkpoints JOB_ID
cortex-training cancel JOB_ID
cortex-training wait JOB_IDget fetches current job details; checkpoints lists saved checkpoints.
wait waits for the job to reach running, not for training to finish.
See Manage Jobs for the operational workflow.
Print the caller account's reserved GPU capacity and current usage, separated by hardware:
cortex-training capacity
cortex-training capacity --hardware B200The default command queries H200, B200, and B300 independently and
prints a capacity_by_hardware map. Each entry includes has_reservation,
max_total_gpus, reserved_gpus, in_use_gpus, pending_gpus, and
available_gpus.
Each hardware lookup is attempted once with at most a 10-second connect timeout
and a 30-second read timeout; a shorter SDK timeout still wins. A silent
network, proxy, or service hop therefore fails instead of leaving the command
blocked indefinitely. The three lookups run serially; use --hardware when
only one type is needed.
--hardware keeps the single-capacity response shape for one GPU type.
max_total_gpus is the canonical ceiling and supersedes the deprecated
reserved_gpus. in_use_gpus counts only GPUs the account holds; queued work
is reported separately in pending_gpus. See
REST API reference section 5.4 and
GPU hardware.
The submit command expects a CreateJob JSON body:
{
"hardware": "H200",
"sub_job_configs": [
{
"job_type": "sampling",
"model_name": "gpt2",
"inference_config": {
"max_seq_len": 128,
"n_gpus": 1
}
}
]
}Optional hardware is H200, B200, or B300. Omit it to use H200. Every
sub-job in the job uses that type. See GPU hardware.
Submit it:
cortex-training submit job.json
cortex-training submit job.json --wait
cortex-training submit job.json --dry-runWithout --wait, submission returns without waiting for the job to run.
--wait waits until running, not until training finishes. --dry-run
validates and prints the request body without sending it, applying the same
checks as a live submit — so the printed body already shows any peft_config
in its canonical form. A log_probability or log_prob sub-job is rejected
before send, including on --dry-run, with
log_probability sub-jobs are not currently supported.
The repo includes a Prime-RL/Qwen3.6 training example:
cortex-training submit examples/api/training.json
cortex-training submit examples/api/sampling.jsonThat file creates a training sub-job for Qwen/Qwen3.6-35B-A3B with
training_config.model_provider set to prime_rl.
After the training job is running, send one tokenized training batch:
cortex-training --job JOB_ID fwd-bwd examples/api/fwd-bwd.json
cortex-training --job-id JOB_ID stepThe fwd-bwd JSON is human-readable: it contains text samples, tokenizer
settings, batch size, sequence length, position_ids, and label generation
settings. The CLI tokenizes the text, builds tensor kwargs, serializes
{"args": (), "kwargs": ...} as a DSSST1 safetensors frame (see
REST API reference section 9),
submits forward_backward, and polls the request by default. Set "poll": false in
the JSON to print only the submitted request_id.
Text payloads require transformers in the client environment. You can also
provide pre-tokenized tensor data directly under payload.kwargs for fully
offline use.
Run an optimizer step after fwd-bwd with:
cortex-training --job-id JOB_ID step
cortex-training --job-id JOB_ID step --lr 2e-5When omitted, --lr defaults to 1e-4.
After a job has already been created and reached running, load a checkpoint into that existing job with:
cortex-training --job-id JOB_ID load CHECKPOINT_IDTo load from another job's checkpoint store:
cortex-training --job-id JOB_ID load CHECKPOINT_ID --source-job-id SOURCE_JOB_IDTo load into a specific training sub-job (useful for multi-sub-job sessions):
cortex-training --job-id JOB_ID load CHECKPOINT_ID --target-sub-job-id JOB_ID:training:0When --target-sub-job-id is omitted, the service routes the load to the
session's training sub-job. Sampling sub-jobs are not valid targets.
To find the training sub-job and its GPU count:
cortex-training get JOB_ID | jq '.sub_jobs[] | select(.job_type=="training") | {sub_job_id, n_gpus: .training_config.n_gpus}'get takes the job id as a positional argument, so --job-id is not used here.
The global --job-id option is only for the data-plane subcommands that have no
positional job id (fwd-bwd, step, load, generate, weight-sync).
For Python sub-job discovery, see the
runtime load API reference.
A job has at most one training sub-job, so load --target-sub-job-id can be
omitted. Use it when you want explicit control over which sub-job receives the
checkpoint rather than relying on the server's default resolution; it must name a
training sub-job. weight-sync takes its own --target-sub-job-id, which names
sampling sub-jobs and can be repeated — see
Sync Training Weights.
When changing n_gpus from the checkpoint's source job, create the target
training sub-job with "load_optimizer_states": false in its training_config.
This cannot be changed at load time. See
DP size compatibility for the constraint.
This is the runtime load path. Create-time resume still uses
source_checkpoint_info
in the submitted sub-job JSON.
load polls until the request completes by default. Pass --no-poll to return
the request metadata without waiting for the result.
Sampling requires a weights-only checkpoint and a new sampling job with
source_checkpoint_info in its submitted JSON; load targets existing training
jobs, not sampling jobs. Resumable checkpoints are not directly loadable by the
sampling runtime.
See Serve a Training Checkpoint for the recipe workflow, or Start sampling from saved weights for the Python save-and-create example.
After a sampling job is running, send readable prompts with sampling parameters:
cortex-training --job-id JOB_ID generate examples/api/generate.jsonThe generate JSON contains prompts, optional sampling_params, and optional
routing_key / strict fields. sampling_params may be one object applied to
all prompts or a list of objects/nulls aligned with prompts. A flat list of
integers such as "prompts": [1, 2, 3] is one pre-tokenized prompt, not three;
use the nested form [[1, 2], [3, 4]] for a batch. The CLI submits generate
and polls the request by default. Set "poll": false to print only the
submitted request_id.
For an RL-style job with one training and one sampling sub-job, sync training weights into sampling with:
cortex-training --job-id JOB_ID weight-syncBy default this syncs from JOB_ID:training:0 to JOB_ID:sampling:0, routes
the operation through JOB_ID:training:0, and polls for completion. Override
sub-job ids when needed:
cortex-training --job-id JOB_ID weight-sync \
--source-sub-job-id JOB_ID:training:0 \
--target-sub-job-id JOB_ID:sampling:0 \
--target-sub-job-id JOB_ID:sampling:1If a backend needs a different operation routing hint, pass
--operation-sub-job-id or --operation-sub-job-type.
Use --weight-format lora for adapter-only sync; vllm and hf are also
accepted formats. Pass --no-poll to return the request metadata without waiting
for synchronization to complete.
Pull every log file the job's experiment run produced. Each sub-job's
_logs/ artifact directory may contain multiple files (e.g.
execution.jsonl, server.log); all of them are downloaded:
cortex-training download-log JOB_ID --output-dir /path/to/dirFiles are written as <output_dir>/<sub_job_id>/<filename> so siblings
do not collide. When --output-dir is omitted, the current working
directory is used instead. The CLI also prints a JSON summary listing
each saved_path.
Programmatic access is CortexTrainingClient.fetch_execution_logs(job_id),
which returns a list of {sub_job_id, filename, artifact_uri, content} dicts.
Reconstruct each sub-job's persisted stdout/stderr chunks into
<output_dir>/<sub_job_id>/stdout.log:
cortex-training download-log JOB_ID --log-type stdout --output-dir /path/to/logs
cortex-training download-log JOB_ID --log-type stdout --output-dir /path/to/logs --resumeThe current working directory is used when --output-dir is omitted.
--resume continues a stdout download already in that directory. It does not
apply to execution logs, including download-log without --log-type stdout.
Two downloads of the same directory at once are unsupported.
Reconstruct each sub-job's GPU metric chunks into
<output_dir>/<sub_job_id>/gpu.jsonl:
cortex-training download-metrics JOB_ID --output-dir /path/to/metrics
cortex-training download-metrics JOB_ID --output-dir /path/to/metrics --resumeThe command prints the saved path, chunk count, and first/last logical artifact
URIs for each reconstructed file. --resume continues a metrics download
already in that directory. Two downloads of the same directory at once are
unsupported.
cortex-training tui is a read-only terminal UI for tailing a running job's
logs live. It uses the same connection handling and fallback order as the CLI,
including named and configured-default Snowflake profiles:
cortex-training tui # opens a job picker
cortex-training tui JOB_ID # opens that job's logs directlyWithout login state, pass connection details the same way as the CLI:
cortex-training tui JOB_ID --connection training
cortex-training tui JOB_ID --config config.json
cortex-training tui JOB_ID --host ACCOUNT.snowflakecomputing.com --pat YOUR_PAT \
--database CORTEX_TRAINING_DB --schema PUBLIC --endpoint cortex-training
cortex-training tui JOB_ID --base-url http://localhost:8084 # local/mockKeep the PAT out of your shell history by exporting it instead (omit JOB_ID
to open the job picker):
export CORTEX_TRAINING_PAT='YOUR_PROGRAMMATIC_ACCESS_TOKEN'
cortex-training tui \
--host ACCOUNT.snowflakecomputing.com \
--database CORTEX_TRAINING_DB --schema PUBLICPass --sub-job-id JOB_ID:training:0 to open one sub-job's log directly instead
of the source list.
The left panel lists the job's sub-jobs; select one to tail its logs (the
zone-manager pod is the Ray head, so a sub-job's worker output is included).
Logs are cached locally so reopening a job replays instantly without
re-fetching from the server — under ~/.cache/cortex-training/ (or
$XDG_CACHE_HOME), overridable with CORTEX_TRAINING_TUI_CACHE_DIR.
A finished job loads its saved console once. A stream error while the job is still running stays an error. The level key filters structured live lines and does not hide lines from that saved console.
The TUI also writes two files into your home directory: saved logs from the s
key (~/cortex-training-<job8>-<source>.log, where <job8> is the first eight
characters of the job id) and its own error log
(~/.cortex-training-errors.log).
Keys in the log view:
| Key | Action |
|---|---|
/ |
Filter the current source |
L |
Cycle minimum log level (INFO / WARNING / ERROR) |
p |
Pause / resume auto-scroll |
s |
Save the current source to ~/cortex-training-<job8>-<source>.log |
y |
Copy the whole log to the clipboard |
c |
Copy the current selection |
r |
Refresh the sub-job list |
[ / ] |
Narrow / widen the sources panel |
b / esc |
Back |
q |
Quit |
In the job picker, / filters by id/status/type and r refreshes. The
--poll-interval flag (default 1.0s) is the minimum interval between log
polls per source, biasing toward server reliability over freshness.
Commands other than the TUI write JSON to stdout, pretty-printed by default.
Use the global --compact flag for compact JSON, or pipe output to jq:
cortex-training --compact list
cortex-training get JOB_ID | jq '.sub_jobs'Use cortex-training --help to list commands and global flags, or
cortex-training COMMAND --help for command-specific arguments. Job-scoped
data-plane help includes the required global option:
usage: cortex-training --job JOB_ID fwd-bwd [-h] json_file
Connection values can also come from:
CORTEX_TRAINING_CONNECTION
SNOWFLAKE_DEFAULT_CONNECTION_NAME
CORTEX_TRAINING_CONFIG
CORTEX_TRAINING_BASE_URL
CORTEX_TRAINING_HOST
SNOWFLAKE_HOST
CORTEX_TRAINING_PAT
SNOWFLAKE_PAT
CORTEX_TRAINING_DATABASE
SNOWFLAKE_DATABASE
CORTEX_TRAINING_SCHEMA
SNOWFLAKE_SCHEMA
CORTEX_TRAINING_ENDPOINTCORTEX_TRAINING_DISABLE_TELEMETRY (truthy) skips OTLP client metrics on
Snowflake profile and PAT clients. CORTEX_TRAINING_ENABLE_SUCCESS_TELEMETRY
(truthy) also emits successful outcomes for essential operations; failures
are emitted by default. See the Python SDK reference.
CORTEX_TRAINING_DISABLE_TENSOR_PROMPTS (truthy) sends pre-tokenized prompts
as JSON lists inside the request frame instead of as tensors. The request body
stays a DSSST1 frame either way.
If you see provide --base-url for local/mock use, or both --host and --pat,
the CLI found a host but no PAT. Add "pat": "..." to config.json or set
CORTEX_TRAINING_PAT.
If no legacy config or complete direct connection is present, the CLI falls
back to the Snowflake Connector's configured default. If your profile is not
named default, pass --connection NAME or set
SNOWFLAKE_DEFAULT_CONNECTION_NAME.
If Snowflake rejects your PAT, the CLI keeps the original error and adds next
steps: 394400 (08001) ... Programmatic access token is invalid from a
connection profile, or a 401 on a request that sent the PAT directly
(config.json, --pat, or CORTEX_TRAINING_PAT). The most common cause is a
user with no network policy; see
Network policy requirement.
A 401 on a connection-profile request gets no hint, because the PAT was
already accepted at login.
If you see Invalid URL ... No scheme supplied, the config is using a bare
Snowflake hostname as base_url. Use host for Snowflake PAT auth, or use a
full local/mock URL such as http://localhost:8084 for base_url.
For server errors, the CLI prints any Snowflake request id and response body
returned by the service. Include those details when debugging a 500.