From 226f436b41654e36c0592df3c91a16edf0a34389 Mon Sep 17 00:00:00 2001 From: YuTian8328 Date: Tue, 22 Sep 2026 08:46:50 +0300 Subject: [PATCH] update ai agent on triton --- aalto/generative-ai-tools.rst | 57 +++++++++++++++++++++++++++++++++-- triton/usage/ai-agents.rst | 55 ++++++++++++++++++++++++++++++--- 2 files changed, 106 insertions(+), 6 deletions(-) diff --git a/aalto/generative-ai-tools.rst b/aalto/generative-ai-tools.rst index e33b60ba1..449c47223 100644 --- a/aalto/generative-ai-tools.rst +++ b/aalto/generative-ai-tools.rst @@ -99,6 +99,8 @@ of text inputs. * This is a very cost-efficient method to use these models. * These models will eventually be available in the Aalto AI Assistant as options you can test. + * For a quick command-line chat with these models on Triton, see + :ref:`llm-cli-chat` below. * OpenAI, and many other providers, also provide general API access for a price. Without an Aalto contract, data security can not be @@ -110,8 +112,59 @@ of text inputs. can help you try it out on triton. -Locally installed models ------------------------- +.. _llm-cli-chat: + +Command-line chat with Aalto models on Triton +~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ + +Triton provides an ``llm`` module for quick interactive chat with the +:doc:`Aalto-hosted open-source models ` from the command +line. It does not change your Python or Conda environment. + +**First-time setup** + +1. On Triton, load the module:: + + module load llm + +2. Run ``llm`` once and accept the Aalto configuration when prompted. + This sets the default model and disables local prompt/response + logging. + +3. Create an API key at + `https://llm-gateway.k8s.aalto.fi/ `__ + (Aalto network or VPN required; see :doc:`llm-web-apis`). + +4. Store the key:: + + llm keys set aalto + +**Normal use** + +Ask a question:: + + llm "your question" + +Interactive chat:: + + llm chat + +Continue the previous conversation:: + + llm -c "Explain the previous answer more simply." + +Pipe content into a prompt:: + + cat script.py | llm "Review this Python code." + +Reasoning output (when the model produces it) is shown automatically. +Hide it with:: + + llm -R "your question" + + +Locally pre-downloaded models +------------------------------ Open-source models can be downloaded and run on your own computers, which can provide the ultimate performance for large analysis. The diff --git a/triton/usage/ai-agents.rst b/triton/usage/ai-agents.rst index 276cfcbe9..97543be13 100644 --- a/triton/usage/ai-agents.rst +++ b/triton/usage/ai-agents.rst @@ -271,7 +271,8 @@ the agent's equivalent). You are working on the Triton HPC cluster at Aalto University. Follow these rules and fetch the linked pages when you need - site-specific details. + site-specific details. If the search_scicomp_docs MCP tool is + available, use it before guessing Triton-specific commands or limits. * Never read, print, move, or commit credentials. * Do not run git commands. Suggest the exact commands and let the @@ -339,8 +340,8 @@ constraints), a skill describes *how* to carry out a specific task: which page to fetch, which commands to run, what output to check. Example skills for Triton projects: -* **Triton documentation lookup:** Fetch the relevant Triton page, extract - the site-specific limit or command, and cite it before proposing an action. +* **Triton documentation lookup:** Fetch the relevant Triton page, extract the + site-specific limit or command, and cite it before proposing an action. * **Slurm job preparation:** Draft or validate ``#SBATCH`` lines against this project's conventions and the Triton tutorials; group short tasks into arrays; require user approval before running ``sbatch``. @@ -357,4 +358,50 @@ for Triton projects: above. Treat every skill as executable third-party content: read all of its instructions and scripts before adding it, and reject any skill that asks for credentials, weakens confirmation settings, or sends project - data elsewhere. + data elsewhere. + + +SciComp docs MCP tool (``search_scicomp_docs``) +~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ + +Agents often invent plausible but wrong Slurm flags, partitions, or module +names. An MCP (`Model Context Protocol +`__) tool can give the agent a grounded +search over the SciComp documentation instead of guessing from memory. + +On ``code.triton.aalto.fi`` (the login node for coding agents) you can connect +directly to the hosted SciComp Docs MCP. It exposes one tool, +``search_scicomp_docs``, which ranks relevant pages and returns excerpts with +published ``https://scicomp.aalto.fi/`` URLs. + +Full client notes are in the `docs MCP guide +`__. + +**Connect (client config).** Add this to your MCP client settings (Cursor +example: MCP servers config): + +.. code-block:: json + + { + "mcpServers": { + "scicomp-docs": { + "url": "https://docs.triton.aalto.fi/mcp/" + } + } + } + +Restart or reload MCP if your client requires it, then confirm +``search_scicomp_docs`` appears in the tool list. With Codex you can instead +run: + +.. code-block:: bash + + codex mcp add scicomp-docs --url https://docs.triton.aalto.fi/mcp/ + +**When the tool is used.** Pose a docs question in chat; the agent should +call ``search_scicomp_docs`` and cite the returned ``https://scicomp.aalto.fi/`` +URLs. For example: “How do I request a GPU on Triton?”, “Where is the Aalto +storage quota documentation?”, or “What is the policy for AI agents on +Triton?”. In agent / auto mode, if the agent needs Triton or SciComp facts +(partitions, modules, quotas, Slurm, policy, …), it may call the same tool +without you mentioning docs. \ No newline at end of file