Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
57 changes: 55 additions & 2 deletions aalto/generative-ai-tools.rst
Original file line number Diff line number Diff line change
Expand Up @@ -99,6 +99,8 @@ of text inputs.
* This is a very cost-efficient method to use these models.
* These models will eventually be available in the Aalto AI
Assistant as options you can test.
* For a quick command-line chat with these models on Triton, see
:ref:`llm-cli-chat` below.

* OpenAI, and many other providers, also provide general API access
for a price. Without an Aalto contract, data security can not be
Expand All @@ -110,8 +112,59 @@ of text inputs.
can help you try it out on triton.


Locally installed models
------------------------
.. _llm-cli-chat:

Command-line chat with Aalto models on Triton
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~

Triton provides an ``llm`` module for quick interactive chat with the
:doc:`Aalto-hosted open-source models <llm-web-apis>` from the command
line. It does not change your Python or Conda environment.

**First-time setup**

1. On Triton, load the module::

module load llm

2. Run ``llm`` once and accept the Aalto configuration when prompted.
This sets the default model and disables local prompt/response
logging.

3. Create an API key at
`https://llm-gateway.k8s.aalto.fi/ <https://llm-gateway.k8s.aalto.fi/>`__
(Aalto network or VPN required; see :doc:`llm-web-apis`).

4. Store the key::

llm keys set aalto

**Normal use**

Ask a question::

llm "your question"

Interactive chat::

llm chat

Continue the previous conversation::

llm -c "Explain the previous answer more simply."

Pipe content into a prompt::

cat script.py | llm "Review this Python code."

Reasoning output (when the model produces it) is shown automatically.
Hide it with::

llm -R "your question"


Locally pre-downloaded models
------------------------------

Open-source models can be downloaded and run on your own computers,
which can provide the ultimate performance for large analysis. The
Expand Down
55 changes: 51 additions & 4 deletions triton/usage/ai-agents.rst
Original file line number Diff line number Diff line change
Expand Up @@ -271,7 +271,8 @@ the agent's equivalent).

You are working on the Triton HPC cluster at Aalto University.
Follow these rules and fetch the linked pages when you need
site-specific details.
site-specific details. If the search_scicomp_docs MCP tool is
available, use it before guessing Triton-specific commands or limits.

* Never read, print, move, or commit credentials.
* Do not run git commands. Suggest the exact commands and let the
Expand Down Expand Up @@ -339,8 +340,8 @@ constraints), a skill describes *how* to carry out a specific task: which
page to fetch, which commands to run, what output to check. Example skills
for Triton projects:

* **Triton documentation lookup:** Fetch the relevant Triton page, extract
the site-specific limit or command, and cite it before proposing an action.
* **Triton documentation lookup:** Fetch the relevant Triton page, extract the
site-specific limit or command, and cite it before proposing an action.
* **Slurm job preparation:** Draft or validate ``#SBATCH`` lines against this
project's conventions and the Triton tutorials; group short tasks into
arrays; require user approval before running ``sbatch``.
Expand All @@ -357,4 +358,50 @@ for Triton projects:
above. Treat every skill as executable third-party content: read all of
its instructions and scripts before adding it, and reject any skill that
asks for credentials, weakens confirmation settings, or sends project
data elsewhere.
data elsewhere.


SciComp docs MCP tool (``search_scicomp_docs``)
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~

Agents often invent plausible but wrong Slurm flags, partitions, or module
names. An MCP (`Model Context Protocol
<https://modelcontextprotocol.io>`__) tool can give the agent a grounded
search over the SciComp documentation instead of guessing from memory.

On ``code.triton.aalto.fi`` (the login node for coding agents) you can connect
directly to the hosted SciComp Docs MCP. It exposes one tool,
``search_scicomp_docs``, which ranks relevant pages and returns excerpts with
published ``https://scicomp.aalto.fi/`` URLs.

Full client notes are in the `docs MCP guide
<https://github.com/AaltoSciComp/llm-examples/blob/main/triton-mcp/docs-mcp.md>`__.

**Connect (client config).** Add this to your MCP client settings (Cursor
example: MCP servers config):

.. code-block:: json

{
"mcpServers": {
"scicomp-docs": {
"url": "https://docs.triton.aalto.fi/mcp/"
}
}
}

Restart or reload MCP if your client requires it, then confirm
``search_scicomp_docs`` appears in the tool list. With Codex you can instead
run:

.. code-block:: bash

codex mcp add scicomp-docs --url https://docs.triton.aalto.fi/mcp/

**When the tool is used.** Pose a docs question in chat; the agent should
call ``search_scicomp_docs`` and cite the returned ``https://scicomp.aalto.fi/``
URLs. For example: “How do I request a GPU on Triton?”, “Where is the Aalto
storage quota documentation?”, or “What is the policy for AI agents on
Triton?”. In agent / auto mode, if the agent needs Triton or SciComp facts
(partitions, modules, quotas, Slurm, policy, …), it may call the same tool
without you mentioning docs.
Loading