From 5cb511929a99bfc00f7a9896d4cee9216e32e77e Mon Sep 17 00:00:00 2001 From: Bernard Ladenthin Date: Tue, 22 Sep 2026 23:48:01 +0200 Subject: [PATCH 01/36] llama-atmosphere-agent: REPL commands, approval prompt, status line, rendered answers MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Four things a terminal agent needs to be usable, built on what Atmosphere already has and without a new dependency: - SlashCommands: /help /status /tools /mode /compact /clear /exit, with aliases and a trailing argument. An unknown /command goes to the model rather than being rejected, so no escape syntax is needed for a line like /usr/bin/env. - Approval: ConsoleApprovalStrategy asks [y]es/[n]o/[a]uto before run_command, write_file, edit_file, delete and rename; reading tools never ask. The gate is Atmosphere's (ToolApprovalPolicy.custom + ApprovalStrategy, attached by AgentRunner.approval), so a denial reaches the model as its own {"status":"cancelled"} tool result and it replans. Manual is the default; --auto and the [a] answer switch the session to auto. One-shot (--prompt) has nobody to ask, so a gated call is denied unless --auto was passed. - /compact summarizes the conversation through the same model with no tools (compact-prompt.txt) and replaces the history with the summary as a user message plus a short acknowledgement. - Status line above the prompt: [manual · ctx ~3.1k/16k · 9 tools · local-model]. llama.cpp reports usage only to clients that ask (stream_options.include_usage) and Atmosphere's does not, so the count is normally an estimate, marked with ~; the window size comes from --ctx-size or the server's /props (ServerProps). - Console: answers are rendered line by line as they stream (headings, bullets, fences, inline bold/code) and tool calls print as "● tool {args}". Append-only, never redrawn, so piping stays correct. Colour is decided once (CLICOLOR_FORCE, NO_COLOR, TERM=dumb, CLICOLOR, then "is a terminal"). Tests: 61 green. ApprovalWireTest drives denial and approval through the real OpenAiCompatServer and asserts what reaches the model; ConsoleApprovalStrategyTest, SlashCommandsTest, ConsoleFormattingTest and ServerPropsTest cover the rest. Verified live with Qwen3-4B on Vulkan: /help and /status are intercepted, a denied "rm -rf build" is reported to the model, /mode auto takes effect, answers render. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01E2h8gXyyE5UeimkL9vQv9G --- CLAUDE.md | 29 +++ README.md | 10 +- llama-atmosphere-agent/README.md | 70 +++++- .../llama/atmosphere/AgentOptions.java | 18 +- .../llama/atmosphere/AgentRunner.java | 49 ++++ .../net/ladenthin/llama/atmosphere/Ansi.java | 173 ++++++++++++++ .../llama/atmosphere/ApprovalMode.java | 50 ++++ .../atmosphere/ConsoleApprovalStrategy.java | 155 ++++++++++++ .../llama/atmosphere/ConsoleSession.java | 61 ++++- .../llama/atmosphere/LocalAgent.java | 226 +++++++++++++++++- .../llama/atmosphere/MarkdownConsole.java | 127 ++++++++++ .../llama/atmosphere/ServerProps.java | 101 ++++++++ .../llama/atmosphere/SlashCommands.java | 109 +++++++++ .../llama/atmosphere/StatusLine.java | 82 +++++++ .../llama/atmosphere/compact-prompt.txt | 11 + .../atmosphere/compact-prompt.txt.license | 3 + .../net/ladenthin/llama/atmosphere/help.txt | 13 + .../llama/atmosphere/help.txt.license | 3 + .../llama/atmosphere/ApprovalWireTest.java | 166 +++++++++++++ .../ConsoleApprovalStrategyTest.java | 123 ++++++++++ .../atmosphere/ConsoleFormattingTest.java | 130 ++++++++++ .../llama/atmosphere/LocalAgentTest.java | 3 +- .../llama/atmosphere/ServerPropsTest.java | 45 ++++ .../llama/atmosphere/SlashCommandsTest.java | 57 +++++ 24 files changed, 1776 insertions(+), 38 deletions(-) create mode 100644 llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/Ansi.java create mode 100644 llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ApprovalMode.java create mode 100644 llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleApprovalStrategy.java create mode 100644 llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/MarkdownConsole.java create mode 100644 llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ServerProps.java create mode 100644 llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/SlashCommands.java create mode 100644 llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/StatusLine.java create mode 100644 llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/compact-prompt.txt create mode 100644 llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/compact-prompt.txt.license create mode 100644 llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt create mode 100644 llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt.license create mode 100644 llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ApprovalWireTest.java create mode 100644 llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleApprovalStrategyTest.java create mode 100644 llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleFormattingTest.java create mode 100644 llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ServerPropsTest.java create mode 100644 llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/SlashCommandsTest.java diff --git a/CLAUDE.md b/CLAUDE.md index a12e2f044..23a7104c9 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -2233,6 +2233,35 @@ a `jvm.config` takes no comments, so REUSE can only read its metadata from that `REUSE Compliance Check` job fails on `main` — which is how it was found, the PR run having been cancelled. Spotless (palantir) is configured in its own pom; the model-free CI job runs `spotless:check`. +**The REPL layer (commands, approval, status line, rendering).** Five small classes, no new dependency: +`SlashCommands` (a line starting with `/` whose first word names a command is handled locally — +`/help /status /tools /mode /compact /clear /exit`; **an unknown `/command` goes to the model**, which +is why no escape syntax is needed for `/usr/bin/…`), `ApprovalMode` + `ConsoleApprovalStrategy` +(`[y]es/[n]o/[a]uto` per gated call), `StatusLine`, and `Ansi` + `MarkdownConsole`. Four points that +are decisions, not details: + +1. **The approval gate is Atmosphere's, not ours.** `AgentRunner.approval(strategy, policy)` attaches + `ToolApprovalPolicy.custom(...)` (gating `run_command`, `write_file`, `edit_file`, `delete`, + `rename` — reading tools never ask) and a `ConsoleApprovalStrategy`; `ToolExecutionHelper` then + blocks the tool loop before the executor runs and turns a denial into the tool result + `{"status":"cancelled","message":"Action cancelled by user"}` for the model. Do not reimplement + that message. `ApprovalWireTest` pins both halves over the real server. +2. **One-shot (`--prompt`) denies a gated call** instead of auto-approving it — `--auto` is the + deliberate opt-in. Atmosphere itself fails closed when no strategy is wired, and this keeps that + direction: an unattended run must not be the most permissive one. +3. **The answer is rendered append-only, one completed line at a time** (`MarkdownConsole`). Redrawing + on every token is what produces the known overdraw/truncation bugs in the Ink/Bubble-Tea based + clients and breaks when the output is piped. Only headings, bullets, fences and inline + `**bold**`/`` `code` `` are handled; italics deliberately are not (`*` is more often a glob than + emphasis). Colour is decided once in `Ansi.detect()` — `CLICOLOR_FORCE`, then `NO_COLOR`, then + `TERM=dumb`/`CLICOLOR=0`, else "is a terminal" via `Console.isTerminal()` (reflective: JDK 22+; + below that `System.console() != null`). +4. **The context number in the status line is an estimate, marked `~`.** llama.cpp emits its usage + chunk only when the client sets `stream_options.include_usage`, and Atmosphere's client does not; + `ConsoleSession.usage()` takes the real count when one arrives, otherwise `LocalAgent.estimateTokens` + uses four characters per token. The window size is `--ctx-size` (in-process) or the server's + `/props` (`ServerProps`), and is omitted rather than guessed when neither answers. + **The default system prompt is general-purpose on purpose — do not narrow it back.** Every model-facing text is a resource, not a Java literal: `src/main/resources/net/ladenthin/llama/atmosphere/` holds `system-prompt.txt`, `system-prompt-shell.txt`, `system-prompt-no-shell.txt` and `run-command-tool.txt` diff --git a/README.md b/README.md index 551c09723..daa243ebd 100644 --- a/README.md +++ b/README.md @@ -1042,8 +1042,14 @@ mvn -q compile exec:java \ ``` > [!WARNING] -> With `--allow-shell` the model runs any command it decides to run, with your user's rights and -> without asking. Use a machine and a workspace you are willing to hand to the model. +> `--allow-shell` lets the model run any command with your user's rights. By default every write and +> every command is confirmed on the console (`[y]es / [n]o / [a]uto`); `--auto` turns that off. Use a +> workspace you are willing to hand to the model. + +In the REPL, `/help` lists the commands (`/status`, `/tools`, `/mode manual|auto`, `/compact`, +`/clear`, `/exit`); anything else goes to the model. A status line shows the approval mode and the +context used (`[manual · ctx ~3.1k/16k · 9 tools · local-model]`), and the answer is rendered with +headings, bullets and code spans. On Windows PowerShell quote the whole argument (`"-Dexec.args=--model models\… --allow-shell"`); for the GPU add e.g. `-Dllama.classifier=vulkan-windows-x86-64` and `--ngl 99`. The agent's diff --git a/llama-atmosphere-agent/README.md b/llama-atmosphere-agent/README.md index 293595abe..56d39e333 100644 --- a/llama-atmosphere-agent/README.md +++ b/llama-atmosphere-agent/README.md @@ -27,10 +27,10 @@ run it. Its `pom.xml` pins `llama.version` to the release these instructions des pass `-Dllama.version=…` to run against another core, e.g. a `-SNAPSHOT` before a release. > [!WARNING] -> With `--allow-shell` the model runs **any** command it decides to run, with **your** user's -> rights, without asking — deleting files, pushing to git, stopping containers included. Start it on -> a machine and account you are willing to hand to the model, and point `--workspace` at a copy of -> a project, not your only one. Without the flag it can only use the file tools inside `--workspace`. +> `--allow-shell` lets the model run **any** command with **your** user's rights. By default it asks +> first (`[y]es / [n]o / [a]uto` per call) and only writes and commands are gated — but `--auto`, and +> the `[a]` answer, turn that off for the rest of the session. Point `--workspace` at a copy of a +> project, not your only one. ## Getting started from scratch @@ -102,9 +102,9 @@ Metal with the default jar already. The root README's classifier table lists eve -Dexec.args="--base-url http://127.0.0.1:8080/v1 --workspace /path/to/project --allow-shell" ``` -A `you>` prompt appears. Type a request; the answer streams as it is generated, and every tool call -and its result are printed as `⚙ read_file {path=…}` / `↳ …` lines. `/clear` drops the history, -`/exit` quits. +A `you>` prompt appears, above it a status line. The answer streams as it is generated, and every +tool call and its result are printed as `● read_file {path=…}` / `↳ …` lines. See +[Commands, approval and the status line](#commands-approval-and-the-status-line). A single turn without the REPL: @@ -150,6 +150,7 @@ irrelevant: inference stays in the running server, the agent's JVM loads no mode | `--log-verbosity ` / `--verbose` | llama.cpp log threshold for `--model` (1 errors, 2 warnings, 3 info, 4 trace, 5 debug) / log everything | `2` / off | | `--workspace ` | directory the file tools are confined to, and where `run_command` starts | cwd | | `--allow-shell` | register `run_command`: any command line, starting in the workspace | off | +| `--auto` | run tools without asking (otherwise every write and command is confirmed) | off | | `--system ` | replace the default system prompt | built-in | | `--prompt `, `-p` | one turn, then exit | interactive | | `--temperature ` / `--max-tokens ` | sampling / per-call budget | `0.2` / `2048` | @@ -172,6 +173,56 @@ code page it saw at startup, so umlauts and emoji in the answer would turn into project's `.mvn/jvm.config` pins `-Dstdout.encoding=UTF-8 -Dstderr.encoding=UTF-8` for the `mvn` JVM so both sides agree. +### Commands, approval and the status line + +A line that starts with `/` and names a command is answered by the agent itself; anything else — an +unknown `/command` included — goes to the model: + +| Command | | +|---|---| +| `/help` (`/?`, `/commands`) | the overview below | +| `/status` | mode, context use, tools, model, workspace, history size | +| `/tools` | the tools offered, and which of them ask first | +| `/mode [manual\|auto]` (`/approve`) | show or set the approval mode | +| `/compact [focus]` | summarize the conversation and continue from the summary | +| `/clear` (`/reset`, `/new`) | drop the history | +| `/exit` (`/quit`) | leave | + +**Approval.** In the default `manual` mode every tool that writes or runs a command — +`run_command`, `write_file`, `edit_file`, `delete`, `rename` — asks before it runs: + +``` +● run_command {command=rm -rf build} +? run_command {command=rm -rf build} + allow? [y]es / [n]o / [a]uto (no more questions): +``` + +`[y]` runs it once, `[n]` cancels it *and tells the model*, so it replans instead of assuming the +command ran, `[a]` switches to `auto` for the rest of the session (`/mode manual` switches back). +Reading tools (`ls`, `read_file`, `glob`, `grep`) never ask. **In one-shot mode (`--prompt`) nobody +can answer, so a gated call is denied** — pass `--auto` to run unattended. The gate itself is +Atmosphere's (`ToolApprovalPolicy` + `ApprovalStrategy`); the agent only supplies the question and +the answer. + +**`/compact`** asks the model to summarize the conversation (goal, facts, work done, problems, state, +next step; `/compact ` adds an emphasis), then replaces the history with that summary. Use it +when the context fills up. Note the history only ever held the user texts and the final answers — +tool rounds are not replayed across turns — so nothing else is lost. + +**The status line** above the prompt reads +`[manual · ctx ~3.1k/16k · 9 tools · local-model]`: the approval mode, the context used out of the +window, the number of tools and the model id. A `~` means the number is an estimate from the text +length: llama.cpp reports token counts only to clients that ask for them +(`stream_options.include_usage`), which Atmosphere's client does not. The window size comes from +`--ctx-size` with `--model`, and from the server's `/props` with `--base-url`; when neither answers, +the line shows the count alone. + +**Colours and Markdown.** The answer is rendered line by line as it streams: headings, bullets, +fenced code blocks and inline `**bold**` / `` `code` ``. Nothing is ever redrawn, so piping the output +into a file stays correct. Colour is on only on a real terminal and obeys `NO_COLOR`, `TERM=dumb`, +`CLICOLOR=0` and `CLICOLOR_FORCE=1`. On the classic Windows `conhost.exe` escape sequences may show up +literally unless `HKCU\Console\VirtualTerminalLevel` is 1 — Windows Terminal needs nothing. + ### The system prompt Without `--system` the agent uses a built-in **general-purpose** prompt: it names the file tools and @@ -261,7 +312,10 @@ starter are the *deployment* layer on top of the same runtime — not needed for ## Limitations / next steps - Tool rounds are not kept in the cross-turn history (only `user`/`assistant` text is replayed). -- No approval prompts for destructive tools yet (`ToolDefinition.requiresApproval` exists in Atmosphere). +- Approval is per call, not per command prefix: there is no "always allow `git status`" rule yet. + Atmosphere's `ApprovalResolution` also supports approve-with-edited-arguments, which the console + does not offer. +- No auto-compaction when the context fills up; `/compact` is manual. - An engine error after the stream started ends the turn silently (see the table). - **One in-process agent per machine at a time.** The core extracts its native library to a fixed name (`jllama.dll` / `libjllama.so` in the temp directory); on Windows a second JVM cannot replace diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentOptions.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentOptions.java index cdc63af39..64634c66b 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentOptions.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentOptions.java @@ -56,6 +56,7 @@ public final class AgentOptions { private final String modelId; private final Path workspace; private final boolean allowShell; + private final boolean auto; private final double temperature; private final int maxTokens; private final int maxToolRounds; @@ -74,6 +75,7 @@ private AgentOptions(Builder b) { this.modelId = b.modelId; this.workspace = b.workspace; this.allowShell = b.allowShell; + this.auto = b.auto; this.temperature = b.temperature; this.maxTokens = b.maxTokens; this.maxToolRounds = b.maxToolRounds; @@ -97,6 +99,7 @@ public static AgentOptions parse(String[] args) { switch (a) { case "-h", "--help" -> b.help = true; case "--allow-shell" -> b.allowShell = true; + case "--auto" -> b.auto = true; case "--base-url" -> b.baseUrl = stripTrailingSlash(value(args, ++i, a)); case "--model" -> b.modelPath = value(args, ++i, a); case "--ngl", "--gpu-layers" -> b.gpuLayers = intValue(args, ++i, a); @@ -168,6 +171,7 @@ public static String usage() { "Agent:", " --workspace directory the file tools are confined to (default: cwd)", " --allow-shell add the run_command tool (runs any command line, starting in the workspace)", + " --auto run tools without asking (default: ask before writes and commands)", " --system replace the default system prompt", " --prompt , -p run one turn and exit (default: interactive; /exit to quit)", " --temperature sampling temperature (default " + DEFAULT_TEMPERATURE + ")", @@ -259,6 +263,16 @@ public Path getWorkspace() { return workspace; } + /** + * Whether tool calls run without asking. + * + * @return {@code true} when {@code --auto} was given, i.e. the session starts in + * {@link ApprovalMode#AUTO} + */ + public boolean isAuto() { + return auto; + } + /** * Whether the {@code run_command} tool is registered. * @@ -327,7 +341,8 @@ public String toString() { return "AgentOptions{baseUrl=" + baseUrl + ", modelPath=" + modelPath + ", gpuLayers=" + gpuLayers + ", ctxSize=" + ctxSize + ", logVerbosity=" + (verbose ? "verbose" : logVerbosity) + ", modelId=" + modelId + ", workspace=" + workspace - + ", allowShell=" + allowShell + ", temperature=" + temperature + ", maxTokens=" + maxTokens + + ", allowShell=" + allowShell + ", auto=" + auto + ", temperature=" + temperature + ", maxTokens=" + + maxTokens + ", maxToolRounds=" + maxToolRounds + ", prompt=" + (prompt == null ? "" : "") + "}"; } @@ -347,6 +362,7 @@ private static final class Builder { String modelId = DEFAULT_MODEL_ID; Path workspace = Paths.get("").toAbsolutePath().normalize(); boolean allowShell; + boolean auto; double temperature = DEFAULT_TEMPERATURE; int maxTokens = DEFAULT_MAX_TOKENS; int maxToolRounds = DEFAULT_MAX_TOOL_ROUNDS; diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentRunner.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentRunner.java index 3ca6c9bd0..de969f86f 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentRunner.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentRunner.java @@ -10,11 +10,14 @@ import org.atmosphere.ai.AiConfig; import org.atmosphere.ai.RetryPolicy; import org.atmosphere.ai.StreamingSession; +import org.atmosphere.ai.approval.ApprovalStrategy; +import org.atmosphere.ai.approval.ToolApprovalPolicy; import org.atmosphere.ai.llm.BuiltInAgentRuntime; import org.atmosphere.ai.llm.ChatMessage; import org.atmosphere.ai.llm.ToolLoopPolicies; import org.atmosphere.ai.llm.ToolLoopPolicy; import org.atmosphere.ai.tool.ToolDefinition; +import org.jspecify.annotations.Nullable; /** * The minimal wiring between Atmosphere's built-in OpenAI-compatible agent runtime and an @@ -36,6 +39,8 @@ public final class AgentRunner { private final String systemPrompt; private final int maxToolRounds; private RetryPolicy retryPolicy = RetryPolicy.DEFAULT; + private @Nullable ApprovalStrategy approvalStrategy; + private @Nullable ToolApprovalPolicy approvalPolicy; /** * Configure the runtime for one endpoint. @@ -83,6 +88,24 @@ public AgentRunner retryPolicy(RetryPolicy retryPolicy) { return this; } + /** + * Gate the tools {@code policy} selects behind {@code strategy}: Atmosphere's tool loop then blocks + * on the strategy before such a tool runs, and turns a denial into a {@code cancelled} tool result + * for the model on its own. + * + *

Without this, no tool is gated — Atmosphere's default policy honours a tool's own + * {@code requiresApproval()}, and none of this agent's tools set it. + * + * @param strategy what asks the user, e.g. {@link ConsoleApprovalStrategy} + * @param policy which tools it is asked about, e.g. {@link ConsoleApprovalStrategy#policy()} + * @return this runner + */ + public AgentRunner approval(ApprovalStrategy strategy, ToolApprovalPolicy policy) { + this.approvalStrategy = strategy; + this.approvalPolicy = policy; + return this; + } + /** * The model ids the endpoint advertises on {@code GET /v1/models}, falling back to the configured * id when enumeration fails. @@ -110,6 +133,29 @@ public List toolNames() { * @param session receives streamed text, tool events and the terminal complete/error */ public void run(String message, List history, StreamingSession session) { + run(message, history, session, tools, systemPrompt); + } + + /** + * Run one turn with no tools at all and a system prompt of its own — what {@code /compact} needs: + * a summary must not read files or run commands, it must only condense what is already there. + * + * @param message the user message + * @param history prior turns, replayed before the message + * @param session receives the streamed summary + * @param systemPrompt the system prompt for this one turn + */ + public void runWithoutTools( + String message, List history, StreamingSession session, String systemPrompt) { + run(message, history, session, List.of(), systemPrompt); + } + + private void run( + String message, + List history, + StreamingSession session, + List tools, + String systemPrompt) { AgentExecutionContext context = new AgentExecutionContext( message, systemPrompt, @@ -127,6 +173,9 @@ public void run(String message, List history, StreamingSession sess null, null); context = context.withRetryPolicy(retryPolicy); + if (approvalStrategy != null && approvalPolicy != null) { + context = context.withApprovalStrategy(approvalStrategy).withApprovalPolicy(approvalPolicy); + } context = ToolLoopPolicies.attach(context, ToolLoopPolicy.maxIterations(maxToolRounds)); runtime.execute(context, session); } diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/Ansi.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/Ansi.java new file mode 100644 index 000000000..aac4bed09 --- /dev/null +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/Ansi.java @@ -0,0 +1,173 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import java.util.Locale; +import java.util.function.Function; +import org.jspecify.annotations.Nullable; + +/** + * The handful of ANSI styles this agent uses, and the one decision of whether to emit them at all. + * + *

The decision is made once at startup, not per line: when colour is off every helper returns its + * argument unchanged, so no other class branches on it. The order follows the conventions the + * terminal ecosystem actually honours: + * + *

    + *
  1. {@code CLICOLOR_FORCE=1} forces colour on, even when the output is piped (for {@code less -R} + * and CI logs); + *
  2. {@code NO_COLOR} set to any non-empty value turns it off — that is the whole + * NO_COLOR specification; + *
  3. {@code TERM=dumb} or {@code CLICOLOR=0} turns it off; + *
  4. otherwise colour is on only when the output really is a terminal. + *
+ * + *

The terminal test is {@code Console.isTerminal()} where it exists (JDK 22+) and + * {@code System.console() != null} below that. The distinction matters: on JDK 22 to 24 + * {@code System.console()} also returns a console for redirected output, so the older test alone + * would colour a file. Reflection keeps the code compiling and running on JDK 21. + * + *

Windows: Windows Terminal and the VS Code terminal process escape sequences without any setup. + * The classic {@code conhost.exe} does not unless {@code HKCU\Console\VirtualTerminalLevel} is 1 — + * enabling it from the process needs native code, which this agent deliberately does not use, so + * there a user may see the raw sequences and can set {@code NO_COLOR=1}. + */ +public final class Ansi { + + private static final String RESET = "\u001b[0m"; + + /** Never emits escape sequences. */ + public static final Ansi PLAIN = new Ansi(false); + + private final boolean enabled; + + private Ansi(boolean enabled) { + this.enabled = enabled; + } + + /** + * Decide from the environment whether to use colour. + * + * @return a colouring or a plain instance + */ + public static Ansi detect() { + return detect(System::getenv, Ansi::consoleIsTerminal); + } + + /** + * The decision itself, with the environment injected so it can be tested. + * + * @param env reads an environment variable + * @param isTerminal whether standard output is a terminal + * @return a colouring or a plain instance + */ + static Ansi detect(Function env, java.util.function.BooleanSupplier isTerminal) { + if ("1".equals(trimmed(env.apply("CLICOLOR_FORCE")))) { + return new Ansi(true); + } + String noColor = env.apply("NO_COLOR"); + if (noColor != null && !noColor.isEmpty()) { + return PLAIN; + } + if ("dumb".equals(lower(env.apply("TERM"))) || "0".equals(trimmed(env.apply("CLICOLOR")))) { + return PLAIN; + } + return isTerminal.getAsBoolean() ? new Ansi(true) : PLAIN; + } + + private static boolean consoleIsTerminal() { + java.io.Console console = System.console(); + if (console == null) { + return false; + } + try { + // JDK 22+: the only reliable "is a terminal" test; JDK 21 has no such method. + return (boolean) java.io.Console.class.getMethod("isTerminal").invoke(console); + } catch (ReflectiveOperationException | RuntimeException e) { + return true; + } + } + + private static @Nullable String trimmed(@Nullable String value) { + return value == null ? null : value.trim(); + } + + private static @Nullable String lower(@Nullable String value) { + return value == null ? null : value.trim().toLowerCase(Locale.ROOT); + } + + /** + * Whether escape sequences are emitted. + * + * @return {@code true} when styling is on + */ + public boolean isEnabled() { + return enabled; + } + + private String style(String code, String text) { + return enabled ? "\u001b[" + code + "m" + text + RESET : text; + } + + /** + * Bold text. + * + * @param text the text + * @return the styled text + */ + public String bold(String text) { + return style("1", text); + } + + /** + * Dimmed text, for secondary output such as tool results. + * + * @param text the text + * @return the styled text + */ + public String dim(String text) { + return style("2", text); + } + + /** + * Cyan text, for code and tool names. + * + * @param text the text + * @return the styled text + */ + public String cyan(String text) { + return style("36", text); + } + + /** + * Green text, for the marker of a running tool. + * + * @param text the text + * @return the styled text + */ + public String green(String text) { + return style("32", text); + } + + /** + * Yellow text, for questions that need an answer. + * + * @param text the text + * @return the styled text + */ + public String yellow(String text) { + return style("33", text); + } + + /** + * Red text, for errors and denials. + * + * @param text the text + * @return the styled text + */ + public String red(String text) { + return style("31", text); + } +} diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ApprovalMode.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ApprovalMode.java new file mode 100644 index 000000000..a179f588a --- /dev/null +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ApprovalMode.java @@ -0,0 +1,50 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import java.util.Locale; + +/** + * Whether a tool call that changes something asks before it runs. + * + *

The default is {@link #MANUAL}: the model may read freely, but every write and every shell + * command is confirmed on the console. {@link #AUTO} is what the {@code --auto} flag and the + * {@code [a]} answer of a single approval prompt switch to — it stays on for the rest of the session + * until {@code /mode manual} switches back. + */ +public enum ApprovalMode { + + /** Ask before every gated tool call. */ + MANUAL, + + /** Run every tool call without asking. */ + AUTO; + + /** + * The lower-case name used on the console and in {@code /mode}. + * + * @return {@code "manual"} or {@code "auto"} + */ + public String label() { + return name().toLowerCase(Locale.ROOT); + } + + /** + * Parse a mode name as typed by the user. + * + * @param text the name, in any case, optionally surrounded by whitespace + * @return the mode + * @throws IllegalArgumentException when {@code text} names no mode + */ + public static ApprovalMode parse(String text) { + String normalized = text == null ? "" : text.trim().toLowerCase(Locale.ROOT); + for (ApprovalMode mode : values()) { + if (mode.label().equals(normalized)) { + return mode; + } + } + throw new IllegalArgumentException("Unknown mode: " + text + " (expected manual or auto)"); + } +} diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleApprovalStrategy.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleApprovalStrategy.java new file mode 100644 index 000000000..3283257c7 --- /dev/null +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleApprovalStrategy.java @@ -0,0 +1,155 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import java.io.BufferedReader; +import java.io.IOException; +import java.io.PrintStream; +import java.util.List; +import java.util.Locale; +import java.util.Set; +import java.util.concurrent.atomic.AtomicReference; +import org.atmosphere.ai.StreamingSession; +import org.atmosphere.ai.approval.ApprovalResolution; +import org.atmosphere.ai.approval.ApprovalStrategy; +import org.atmosphere.ai.approval.PendingApproval; +import org.atmosphere.ai.approval.ToolApprovalPolicy; +import org.jspecify.annotations.Nullable; + +/** + * Asks on the console before a tool that changes something runs: {@code [y]es / [n]o / [a]uto}. + * + *

Atmosphere does the gating itself — {@code ToolExecutionHelper} consults the + * {@link ToolApprovalPolicy} and, when it says the call needs approval, blocks the tool loop on + * {@link #awaitApprovalDetailed} before the executor runs. A {@link ApprovalResolution#deny() denial} + * is handed to the model as the tool result {@code {"status":"cancelled","message":"Action cancelled + * by user"}}, so it can replan instead of believing the command ran; a timeout becomes + * {@code {"status":"timeout",…}}. Nothing of that is reimplemented here. + * + *

Answering {@code a} switches the whole session to {@link ApprovalMode#AUTO} and approves — the + * remaining calls of the running turn included, because the mode is read per call. {@code /mode + * manual} switches back. + * + *

Without a console the answer is "no". In one-shot mode ({@code --prompt}) there is no one + * to ask, so a gated call is denied and the reason is printed. That is deliberate (and what Claude + * Code's non-interactive mode does): silently auto-approving would make an unattended run the most + * permissive one. Pass {@code --auto} to run unattended. + */ +public final class ConsoleApprovalStrategy implements ApprovalStrategy { + + /** + * The tools that ask before they run: the shell plus everything that writes. Reading ( + * {@code ls}, {@code read_file}, {@code glob}, {@code grep}) is never gated — it cannot change + * the machine, and gating it would make the prompt so frequent that it stops being read. + */ + public static final Set GATED_TOOLS = + Set.of(ShellTool.TOOL_NAME, "write_file", "edit_file", "delete", "rename"); + + private static final int ARGUMENT_PREVIEW_CHARS = 300; + + private final AtomicReference mode; + private final @Nullable BufferedReader input; + private final PrintStream out; + private final Ansi ansi; + + /** + * Create the strategy. + * + * @param mode the shared, mutable approval mode (also written by {@code /mode} and by an + * {@code [a]} answer) + * @param input the console the user answers on, or {@code null} when nobody can be asked + * @param out where the prompt is printed + * @param ansi the styles for the question + */ + public ConsoleApprovalStrategy( + AtomicReference mode, @Nullable BufferedReader input, PrintStream out, Ansi ansi) { + this.mode = mode; + this.input = input; + this.out = out; + this.ansi = ansi; + } + + /** + * The policy that decides which tools this strategy is asked about. + * + * @return a policy gating {@link #GATED_TOOLS} + */ + public static ToolApprovalPolicy policy() { + return ToolApprovalPolicy.custom(tool -> tool != null && GATED_TOOLS.contains(tool.name())); + } + + /** + * The gated tool names among {@code tools}, for the console. + * + * @param toolNames every offered tool name + * @return the names that will ask before running + */ + public static List gated(List toolNames) { + return toolNames.stream().filter(GATED_TOOLS::contains).toList(); + } + + @Override + public ApprovalOutcome awaitApproval(PendingApproval approval, StreamingSession session) { + return awaitApprovalDetailed(approval, session).outcome(); + } + + @Override + public ApprovalResolution awaitApprovalDetailed(PendingApproval approval, StreamingSession session) { + if (mode.get() == ApprovalMode.AUTO) { + return ApprovalResolution.approve(); + } + if (input == null) { + out.println(); + out.println(ansi.red("✗ " + approval.toolName() + " " + preview(approval) + + " — denied: no console to ask (run with --auto to allow tools unattended)")); + out.flush(); + return ApprovalResolution.deny(); + } + out.println(); + out.println(ansi.yellow("? " + approval.toolName()) + " " + ansi.dim(preview(approval))); + while (true) { + out.print(ansi.yellow(" allow? [y]es / [n]o / [a]uto (no more questions): ")); + out.flush(); + String answer = readLine(); + if (answer == null) { + // stdin closed mid-turn: the same situation as having no console at all + out.println(); + out.println(" denied (input closed)"); + out.flush(); + return ApprovalResolution.deny(); + } + switch (answer.trim().toLowerCase(Locale.ROOT)) { + case "y", "yes", "" -> { + return ApprovalResolution.approve(); + } + case "n", "no" -> { + return ApprovalResolution.deny(); + } + case "a", "auto" -> { + mode.set(ApprovalMode.AUTO); + out.println(" approval mode: auto (use /mode manual to ask again)"); + out.flush(); + return ApprovalResolution.approve(); + } + default -> out.println(" please answer y, n or a"); + } + } + } + + private @Nullable String readLine() { + try { + return input == null ? null : input.readLine(); + } catch (IOException e) { + return null; + } + } + + private static String preview(PendingApproval approval) { + String arguments = String.valueOf(approval.arguments()); + return arguments.length() <= ARGUMENT_PREVIEW_CHARS + ? arguments + : arguments.substring(0, ARGUMENT_PREVIEW_CHARS) + "… (" + arguments.length() + " chars)"; + } +} diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleSession.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleSession.java index 2da678424..c5410208a 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleSession.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleSession.java @@ -13,6 +13,7 @@ import java.util.concurrent.TimeUnit; import org.atmosphere.ai.AiEvent; import org.atmosphere.ai.StreamingSession; +import org.atmosphere.ai.TokenUsage; import org.atmosphere.ai.fs.AgentFileSystem; import org.jspecify.annotations.Nullable; @@ -30,21 +31,37 @@ public final class ConsoleSession implements StreamingSession { private static final int RESULT_PREVIEW_CHARS = 400; private final PrintStream out; + private final Ansi ansi; + private final MarkdownConsole markdown; private final Map, Object> injectables; private final StringBuilder text = new StringBuilder(); private final List chunks = new CopyOnWriteArrayList<>(); private final CountDownLatch done = new CountDownLatch(1); private volatile @Nullable Throwable failure; private volatile int toolCalls; + private volatile long inputTokens; /** - * Create a session printing to {@code out}. + * Create an unstyled session printing to {@code out}. * * @param out where streamed text and tool lines go * @param fileSystem the workspace-confined filesystem handed to the file tools */ public ConsoleSession(PrintStream out, AgentFileSystem fileSystem) { + this(out, fileSystem, Ansi.PLAIN); + } + + /** + * Create a session printing to {@code out}. + * + * @param out where streamed text and tool lines go + * @param fileSystem the workspace-confined filesystem handed to the file tools + * @param ansi the styles for the answer and the tool lines + */ + public ConsoleSession(PrintStream out, AgentFileSystem fileSystem, Ansi ansi) { this.out = out; + this.ansi = ansi; + this.markdown = new MarkdownConsole(out, ansi); this.injectables = Map.of(AgentFileSystem.class, fileSystem); } @@ -61,14 +78,33 @@ public Map, Object> injectables() { @Override public void send(String chunk) { chunks.add(chunk); + // The history keeps the raw text; only the console sees the rendered form. text.append(chunk); - out.print(chunk); - out.flush(); + markdown.append(chunk); } @Override public void sendMetadata(String key, Object value) { - // token usage, model id, tool-call argument deltas: not shown on the console + // model id, tool-call argument deltas: not shown on the console + } + + @Override + public void usage(TokenUsage usage) { + // The prompt of the last model call is what fills the context window -- the tokens generated + // in that call are part of the next call's input. Several calls happen per turn (one per tool + // round); the last one wins, which is the largest and the one the next turn continues from. + if (usage != null && usage.input() > 0) { + inputTokens = usage.input(); + } + } + + /** + * The input tokens of the last model call of this turn. + * + * @return the count, or {@code 0} when the endpoint reported no usage + */ + public long inputTokens() { + return inputTokens; } @Override @@ -78,7 +114,7 @@ public void progress(String message) { @Override public void complete() { - out.println(); + markdown.flush(); out.flush(); done.countDown(); } @@ -94,8 +130,8 @@ public void complete(String summary) { @Override public void error(Throwable t) { failure = t; - out.println(); - out.println("[error] " + t); + markdown.flush(); + out.println(ansi.red("[error] " + t)); out.flush(); done.countDown(); } @@ -110,18 +146,17 @@ public void emit(AiEvent event) { switch (event) { case AiEvent.ToolStart start -> { toolCalls++; - if (text.length() > 0 && text.charAt(text.length() - 1) != '\n') { - out.println(); - } - out.println("⚙ " + start.toolName() + " " + start.arguments()); + markdown.flush(); + out.println(ansi.green("●") + " " + ansi.bold(start.toolName()) + " " + + ansi.dim(String.valueOf(start.arguments()))); out.flush(); } case AiEvent.ToolResult result -> { - out.println("↳ " + preview(String.valueOf(result.result()))); + out.println(ansi.dim(" ↳ " + preview(String.valueOf(result.result())))); out.flush(); } case AiEvent.ToolError error -> { - out.println("↳ error: " + error.error()); + out.println(ansi.red(" ↳ error: " + error.error())); out.flush(); } default -> StreamingSession.super.emit(event); diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java index e551b26af..8e2e03477 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java @@ -14,6 +14,8 @@ import java.time.Duration; import java.util.ArrayList; import java.util.List; +import java.util.Optional; +import java.util.concurrent.atomic.AtomicReference; import net.ladenthin.llama.LlamaModel; import net.ladenthin.llama.parameters.ModelParameters; import net.ladenthin.llama.server.OpenAiCompatServer; @@ -57,6 +59,17 @@ public final class LocalAgent { /** The {@code {shell_section}} without {@code --allow-shell}. */ static final String NO_SHELL_PROMPT = "system-prompt-no-shell.txt"; + /** The {@code /help} overview. */ + static final String HELP_TEXT = "help.txt"; + + /** The instructions {@code /compact} sends; placeholder {@code {focus}}. */ + static final String COMPACT_PROMPT = "compact-prompt.txt"; + + /** The system prompt of the summarizing turn: no tools, no agent role, just condense. */ + static final String COMPACT_SYSTEM_PROMPT = + "You summarize a conversation between a user and a coding assistant. Follow the user's" + + " instructions exactly and answer with the summary only."; + private static final Duration SHELL_TIMEOUT = Duration.ofSeconds(120); private static final int SHELL_MAX_OUTPUT_CHARS = 20_000; @@ -142,31 +155,68 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream + " tools=" + runner.toolNames()); List history = new ArrayList<>(); + AtomicReference mode = + new AtomicReference<>(options.isAuto() ? ApprovalMode.AUTO : ApprovalMode.MANUAL); + boolean interactive = options.getPrompt() == null && input != null; + BufferedReader reader = input == null ? null : new BufferedReader(input); + // One-shot runs have nobody at the keyboard, so the strategy gets no console and denies + // gated calls unless --auto was passed (see ConsoleApprovalStrategy). + Ansi ansi = Ansi.detect(); + runner.approval( + new ConsoleApprovalStrategy(mode, interactive ? reader : null, out, ansi), + ConsoleApprovalStrategy.policy()); + int contextSize = options.getModelPath() != null + ? options.getCtxSize() + : ServerProps.contextSize(baseUrl, options.getApiKey()); + if (options.getPrompt() != null) { - return turn(runner, fileSystem, options.getPrompt(), history, out) ? 0 : 1; + return turn(runner, fileSystem, options.getPrompt(), history, out, ansi) + .failure() + == null + ? 0 + : 1; } - if (input == null) { + if (reader == null) { err.println("No interactive input available; pass --prompt ."); return 2; } - BufferedReader reader = new BufferedReader(input); - err.println("Interactive mode: type a request, /clear to drop the history, /exit to quit."); + err.println("Interactive mode: type a request, /help for the commands."); + long inputTokens = 0; + boolean estimated = false; while (true) { + out.println(ansi.dim(StatusLine.render( + mode.get(), inputTokens, estimated, contextSize, tools.size(), options.getModelId()))); out.print("you> "); out.flush(); String line = reader.readLine(); - if (line == null || line.trim().equals("/exit") || line.trim().equals("/quit")) { + if (line == null) { return 0; } if (line.trim().isEmpty()) { continue; } - if (line.trim().equals("/clear")) { - history.clear(); - err.println("(history cleared)"); + Optional command = SlashCommands.parse(line); + if (command.isPresent()) { + if (command.get().command() == SlashCommands.Command.EXIT) { + return 0; + } + inputTokens = handleCommand( + command.get(), + runner, + fileSystem, + history, + mode, + options, + contextSize, + inputTokens, + estimated, + out, + ansi); continue; } - turn(runner, fileSystem, line, history, out); + ConsoleSession completed = turn(runner, fileSystem, line, history, out, ansi); + estimated = completed.inputTokens() == 0; + inputTokens = estimated ? estimateTokens(systemPrompt(options), history) : completed.inputTokens(); } } finally { if (server != null) { @@ -178,17 +228,167 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream } } - private static boolean turn( - AgentRunner runner, AgentFileSystem fileSystem, String message, List history, PrintStream out) + private static ConsoleSession turn( + AgentRunner runner, + AgentFileSystem fileSystem, + String message, + List history, + PrintStream out, + Ansi ansi) throws InterruptedException { - ConsoleSession session = new ConsoleSession(out, fileSystem); + ConsoleSession session = new ConsoleSession(out, fileSystem, ansi); runner.run(message, history, session); boolean finished = session.await(TURN_TIMEOUT); history.add(ChatMessage.user(message)); if (!session.text().isEmpty()) { history.add(ChatMessage.assistant(session.text())); } - return finished && session.failure() == null; + if (!finished) { + session.error(new IllegalStateException("turn did not finish within " + TURN_TIMEOUT)); + } + return session; + } + + /** + * Run one REPL command. + * + * @param command the parsed command + * @param runner the runner (used by {@code /compact}) + * @param fileSystem the workspace filesystem (used by {@code /compact}'s session) + * @param history the conversation history, modified in place by {@code /clear} and {@code /compact} + * @param mode the approval mode, modified in place by {@code /mode} + * @param options the options, for the status output + * @param contextSize the context window in tokens, or {@link StatusLine#UNKNOWN_CONTEXT} + * @param inputTokens the input tokens of the last turn + * @param estimated whether that number is an estimate + * @param out the console + * @param ansi the console styles + * @return the input tokens to show from now on (unchanged, or the summary's after {@code /compact}) + * @throws InterruptedException if interrupted while a summary is generated + */ + private static long handleCommand( + SlashCommands command, + AgentRunner runner, + AgentFileSystem fileSystem, + List history, + AtomicReference mode, + AgentOptions options, + int contextSize, + long inputTokens, + boolean estimated, + PrintStream out, + Ansi ansi) + throws InterruptedException { + switch (command.command()) { + case HELP -> out.println(prompt(HELP_TEXT)); + case CLEAR -> { + history.clear(); + out.println("(history cleared)"); + } + case TOOLS -> { + out.println("tools: " + String.join(", ", runner.toolNames())); + out.println("asks before running (manual mode): " + + String.join(", ", ConsoleApprovalStrategy.gated(runner.toolNames()))); + } + case MODE -> { + if (command.hasArguments()) { + try { + mode.set(ApprovalMode.parse(command.arguments())); + } catch (IllegalArgumentException e) { + out.println(e.getMessage()); + return inputTokens; + } + } + out.println("approval mode: " + mode.get().label()); + } + case STATUS -> { + out.println(StatusLine.render( + mode.get(), + inputTokens, + estimated, + contextSize, + runner.toolNames().size(), + options.getModelId())); + out.println("workspace: " + options.getWorkspace()); + out.println("history: " + history.size() + " messages"); + } + case COMPACT -> { + return compact(runner, fileSystem, history, command.arguments(), out, ansi); + } + case EXIT -> { + // handled by the caller, which has to return from the loop + } + } + return inputTokens; + } + + /** + * A rough token count of what the next request will carry. + * + *

Used only for the status line, and only because llama.cpp sends its own count just to clients + * that ask for it ({@code stream_options.include_usage}), which Atmosphere's client does not. Four + * characters per token is the usual rule of thumb; the status line marks the number with a + * {@code ~} so nobody reads it as exact. + * + * @param systemPrompt the system prompt sent with every request + * @param history the conversation so far + * @return the estimated token count + */ + static long estimateTokens(String systemPrompt, List history) { + long characters = systemPrompt.length(); + for (ChatMessage message : history) { + characters += message.content() == null ? 0 : message.content().length(); + } + return characters / 4; + } + + /** + * Summarize the history and continue from the summary. + * + *

The summary is generated by the same model with no tools, then replaces the history as + * a {@code user} message plus a short assistant acknowledgement — aider's shape, and the one that + * survives a strict chat template, because a conversation may not start with two assistant turns. + * The tool rounds of a turn are not in the history to begin with (only the user text and the final + * answer are), so what is condensed here is what the next turn would have replayed anyway. + * + * @param runner the runner + * @param fileSystem the workspace filesystem for the summary session + * @param history the history, replaced in place + * @param focus optional extra instructions from {@code /compact } + * @param out the console + * @param ansi the console styles + * @return the input tokens the summarizing call reported + * @throws InterruptedException if interrupted while the summary is generated + */ + private static long compact( + AgentRunner runner, + AgentFileSystem fileSystem, + List history, + String focus, + PrintStream out, + Ansi ansi) + throws InterruptedException { + if (history.isEmpty()) { + out.println("(nothing to compact)"); + return 0; + } + String instructions = prompt(COMPACT_PROMPT) + .replace("{focus}", focus.isEmpty() ? "" : System.lineSeparator() + "Focus on: " + focus); + int before = history.size(); + out.println("(compacting " + before + " messages …)"); + ConsoleSession session = new ConsoleSession(out, fileSystem, ansi); + runner.runWithoutTools(instructions, List.copyOf(history), session, COMPACT_SYSTEM_PROMPT); + if (!session.await(TURN_TIMEOUT) || session.text().isBlank()) { + out.println("(compact failed; history kept)"); + return 0; + } + history.clear(); + history.add(ChatMessage.user("Summary of the conversation so far:" + System.lineSeparator() + + session.text().strip())); + history.add(ChatMessage.assistant("Understood, I will continue from that summary.")); + out.println("(compacted " + before + " messages into a summary of " + + session.text().strip().length() + " characters)"); + return session.inputTokens(); } /** diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/MarkdownConsole.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/MarkdownConsole.java new file mode 100644 index 000000000..ea27e904c --- /dev/null +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/MarkdownConsole.java @@ -0,0 +1,127 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import java.io.PrintStream; + +/** + * Renders the streamed answer as it arrives: just enough Markdown to make it readable. + * + *

Append-only, one line at a time. Tokens are buffered until a line is complete, then that + * line is written and never touched again. Redrawing the answer on every token — what the Ink and + * Bubble Tea based clients do — is what produces their overdraw and truncation bugs, and it needs + * cursor control that breaks as soon as the output is piped into a file. The price here is that a + * line appears only once it ends. + * + *

Handled: fenced code blocks (one bit of state), ATX headings, bullet markers, and inline + * {@code **bold**} / {@code `code`}. Italics are deliberately not handled — a lone {@code *} is more + * often a glob or a multiplication than emphasis, and getting that wrong garbles ordinary text. + * Everything else is passed through unchanged, which is what a terminal wants anyway. + * + *

Only the console sees this; {@link ConsoleSession} keeps the raw text for the history, so + * nothing that goes back to the model is affected. + */ +public final class MarkdownConsole { + + private final PrintStream out; + private final Ansi ansi; + private final StringBuilder pending = new StringBuilder(); + private boolean inFence; + + /** + * Create a renderer. + * + * @param out where the rendered text goes + * @param ansi the styles (a plain instance writes the text unchanged) + */ + public MarkdownConsole(PrintStream out, Ansi ansi) { + this.out = out; + this.ansi = ansi; + } + + /** + * Take the next streamed chunk; every complete line in it is rendered and written. + * + * @param chunk the chunk as it arrived, possibly a fragment of a line + */ + public void append(String chunk) { + pending.append(chunk); + int newline; + while ((newline = pending.indexOf("\n")) >= 0) { + String line = pending.substring(0, newline); + pending.delete(0, newline + 1); + out.println(render(line.endsWith("\r") ? line.substring(0, line.length() - 1) : line)); + } + out.flush(); + } + + /** Write what is left of an unfinished line, e.g. an answer that does not end with a newline. */ + public void flush() { + if (pending.length() > 0) { + out.println(render(pending.toString())); + pending.setLength(0); + } + inFence = false; + out.flush(); + } + + /** + * Render one complete line. + * + * @param line the line without its terminator + * @return the line with escape sequences, or unchanged when styling is off + */ + String render(String line) { + String content = line.stripLeading(); + String indent = line.substring(0, line.length() - content.length()); + if (content.startsWith("```") || content.startsWith("~~~")) { + inFence = !inFence; + return ansi.dim(line); + } + if (inFence) { + return ansi.cyan(line); + } + int heading = 0; + while (heading < content.length() && content.charAt(heading) == '#') { + heading++; + } + if (heading > 0 && heading <= 6 && content.startsWith("# ", heading - 1)) { + return indent + ansi.bold(inline(content.substring(heading + 1).strip())); + } + if (content.length() > 2 && "-*+".indexOf(content.charAt(0)) >= 0 && content.charAt(1) == ' ') { + return indent + ansi.cyan("•") + " " + inline(content.substring(2)); + } + return indent + inline(content); + } + + /** + * Style {@code **bold**} and {@code `code`} inside one line. + * + * @param text the line content + * @return the styled content + */ + String inline(String text) { + StringBuilder result = new StringBuilder(text.length()); + int index = 0; + while (index < text.length()) { + int code = text.indexOf('`', index); + int bold = text.indexOf("**", index); + boolean codeFirst = code >= 0 && (bold < 0 || code < bold); + int start = codeFirst ? code : bold; + if (start < 0) { + break; + } + String marker = codeFirst ? "`" : "**"; + int end = text.indexOf(marker, start + marker.length()); + if (end < 0) { + break; // an unclosed marker: leave the rest as typed + } + String span = text.substring(start + marker.length(), end); + result.append(text, index, start).append(codeFirst ? ansi.cyan(span) : ansi.bold(span)); + index = end + marker.length(); + } + return result.append(text.substring(index)).toString(); + } +} diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ServerProps.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ServerProps.java new file mode 100644 index 000000000..32e21bf5b --- /dev/null +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ServerProps.java @@ -0,0 +1,101 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import java.net.URI; +import java.net.http.HttpClient; +import java.net.http.HttpRequest; +import java.net.http.HttpResponse; +import java.time.Duration; +import java.util.List; +import java.util.regex.Matcher; +import java.util.regex.Pattern; + +/** + * Reads the context size of a running server from llama.cpp's {@code GET /props}, so the status line + * can show a percentage in {@code --base-url} mode too. + * + *

Both java-llama.cpp's {@code OpenAiCompatServer} and upstream {@code llama-server} answer with + * {@code {"default_generation_settings":{"n_ctx":…}}}. The endpoint sits next to the OpenAI routes + * rather than under {@code /v1}, and this project's server serves it under both, so both are tried. + * Every failure — a server without the route, a foreign OpenAI-compatible endpoint, a timeout — is + * reported as "unknown"; the status line then shows the token count without a percentage rather than + * a made-up denominator. + */ +public final class ServerProps { + + private static final Pattern N_CTX = Pattern.compile("\"n_ctx\"\\s*:\\s*(\\d+)"); + private static final Duration TIMEOUT = Duration.ofSeconds(3); + + private ServerProps() {} + + /** + * Look the context size up. + * + * @param baseUrl the OpenAI-compatible base URL, e.g. {@code http://127.0.0.1:8080/v1} + * @param apiKey the bearer token to send + * @return the context size in tokens, or {@link StatusLine#UNKNOWN_CONTEXT} when it cannot be read + */ + public static int contextSize(String baseUrl, String apiKey) { + HttpClient client = HttpClient.newBuilder().connectTimeout(TIMEOUT).build(); + for (String url : candidates(baseUrl)) { + int size = read(client, url, apiKey); + if (size != StatusLine.UNKNOWN_CONTEXT) { + return size; + } + } + return StatusLine.UNKNOWN_CONTEXT; + } + + /** + * The {@code /props} URLs tried, in order. + * + * @param baseUrl the base URL + * @return the candidate URLs + */ + static List candidates(String baseUrl) { + String trimmed = baseUrl.endsWith("/") ? baseUrl.substring(0, baseUrl.length() - 1) : baseUrl; + if (trimmed.endsWith("/v1")) { + String root = trimmed.substring(0, trimmed.length() - "/v1".length()); + return List.of(root + "/props", trimmed + "/props"); + } + return List.of(trimmed + "/props"); + } + + /** + * Extract {@code n_ctx} from a {@code /props} body. + * + * @param body the response body + * @return the context size, or {@link StatusLine#UNKNOWN_CONTEXT} when the field is absent + */ + static int parseContextSize(String body) { + Matcher matcher = N_CTX.matcher(body); + if (!matcher.find()) { + return StatusLine.UNKNOWN_CONTEXT; + } + try { + return Integer.parseInt(matcher.group(1)); + } catch (NumberFormatException e) { + return StatusLine.UNKNOWN_CONTEXT; + } + } + + private static int read(HttpClient client, String url, String apiKey) { + try { + HttpRequest request = HttpRequest.newBuilder(URI.create(url)) + .timeout(TIMEOUT) + .header("Authorization", "Bearer " + apiKey) + .GET() + .build(); + HttpResponse response = client.send(request, HttpResponse.BodyHandlers.ofString()); + return response.statusCode() == 200 ? parseContextSize(response.body()) : StatusLine.UNKNOWN_CONTEXT; + } catch (InterruptedException e) { + Thread.currentThread().interrupt(); + return StatusLine.UNKNOWN_CONTEXT; + } catch (RuntimeException | java.io.IOException e) { + return StatusLine.UNKNOWN_CONTEXT; + } + } +} diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/SlashCommands.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/SlashCommands.java new file mode 100644 index 000000000..9f8974d17 --- /dev/null +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/SlashCommands.java @@ -0,0 +1,109 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import java.util.List; +import java.util.Locale; +import java.util.Optional; +import org.jspecify.annotations.Nullable; + +/** + * The REPL's client-side commands: a line the agent answers itself instead of sending it to the model. + * + *

Dispatch is deliberately conservative. A line is a command only when it starts with {@code /} + * and its first word names a known command; everything else — including an unknown + * {@code /foo} — goes to the model verbatim. That is what makes {@code /usr/bin/env} or a line of + * Markdown work with no escape syntax, at the price of a typo being answered by the model rather than + * rejected. (aider rejects unknown commands outright and needs no escape because its REPL is + * line-oriented; this agent is asked prose far more often than it is asked commands.) + * + * @param command the command + * @param arguments the rest of the line, trimmed; empty when the line was just the command + */ +public record SlashCommands(Command command, String arguments) { + + /** The commands the REPL answers itself. */ + public enum Command { + /** Print the command overview. */ + HELP("/help", "/?", "/commands"), + /** Drop the conversation history. */ + CLEAR("/clear", "/reset", "/new"), + /** Summarize the history and continue with the summary; the argument steers the summary. */ + COMPACT("/compact"), + /** Show or set the approval mode; the argument is {@code manual} or {@code auto}. */ + MODE("/mode", "/approve"), + /** Print endpoint, model, tools, approval mode and context usage. */ + STATUS("/status"), + /** List the tools offered to the model. */ + TOOLS("/tools"), + /** Leave the REPL. */ + EXIT("/exit", "/quit"); + + private final List names; + + Command(String... names) { + this.names = List.of(names); + } + + /** + * The names that select this command, the canonical one first. + * + * @return the names, each including the leading slash + */ + public List names() { + return names; + } + + /** + * The canonical name. + * + * @return e.g. {@code "/help"} + */ + public String canonicalName() { + return names.get(0); + } + } + + /** + * Parse one REPL line. + * + * @param line the raw line as typed + * @return the command and its arguments, or empty when the line is a message for the model + */ + public static Optional parse(@Nullable String line) { + // A BOM at the start of piped input would otherwise hide the slash and send /help to the model. + String trimmed = line == null ? "" : line.replace("", "").trim(); + if (!trimmed.startsWith("/")) { + return Optional.empty(); + } + int space = indexOfWhitespace(trimmed); + String name = (space < 0 ? trimmed : trimmed.substring(0, space)).toLowerCase(Locale.ROOT); + String arguments = space < 0 ? "" : trimmed.substring(space + 1).trim(); + for (Command command : Command.values()) { + if (command.names().contains(name)) { + return Optional.of(new SlashCommands(command, arguments)); + } + } + return Optional.empty(); + } + + private static int indexOfWhitespace(String text) { + for (int i = 0; i < text.length(); i++) { + if (Character.isWhitespace(text.charAt(i))) { + return i; + } + } + return -1; + } + + /** + * Whether an argument was given. + * + * @return {@code true} when the line carried more than the command name + */ + public boolean hasArguments() { + return !arguments.isEmpty(); + } +} diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/StatusLine.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/StatusLine.java new file mode 100644 index 000000000..1668491d9 --- /dev/null +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/StatusLine.java @@ -0,0 +1,82 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import java.util.Locale; + +/** + * The one-line status shown above the {@code you>} prompt: approval mode, context usage, tool count + * and model id. + * + *

Context usage is the input side of the last completed turn — the prompt the server had to + * process, which is what fills the context window; the generated tokens of that turn are already part + * of the next request's input. Claude Code's status line computes its percentage the same way. + * The count comes from the server when it reports usage. llama.cpp only sends the trailing usage + * chunk when the client asks for it ({@code stream_options.include_usage}) and Atmosphere's client + * does not, so in practice the number is an estimate from the text length (four characters per + * token, the usual rule of thumb) and is then marked with a {@code ~}. It is meant to answer "am I + * close to the limit, should I /compact", not to be exact. + * + *

The size is known with {@code --model} (it is the + * {@code --ctx-size} the agent loaded the model with) and with {@code --base-url} it comes from the + * server's {@code /props}; when that lookup fails the line shows the token count alone instead of + * inventing a denominator. + */ +public final class StatusLine { + + /** The context size is unknown (no {@code /props}, and no in-process model). */ + public static final int UNKNOWN_CONTEXT = 0; + + private StatusLine() {} + + /** + * Render the status line. + * + * @param mode the approval mode + * @param inputTokens the input tokens of the last turn, or {@code 0} before the first one + * @param estimated whether that number is an estimate rather than the server's own count + * @param contextSize the context window in tokens, or {@link #UNKNOWN_CONTEXT} + * @param tools how many tools are offered to the model + * @param modelId the model id sent in every request + * @return one line, without a trailing newline + */ + public static String render( + ApprovalMode mode, long inputTokens, boolean estimated, int contextSize, int tools, String modelId) { + return "[" + mode.label() + " · " + context(inputTokens, estimated, contextSize) + " · " + tools + " tools · " + + modelId + "]"; + } + + /** + * The context part of the line on its own. + * + * @param inputTokens the input tokens of the last turn + * @param estimated whether that number is an estimate (rendered with a leading {@code ~}) + * @param contextSize the context window in tokens, or {@link #UNKNOWN_CONTEXT} + * @return e.g. {@code "ctx 1.2k/16k"}, or {@code "ctx 1.2k"} when the size is unknown + */ + static String context(long inputTokens, boolean estimated, int contextSize) { + // Two numbers in k, no percentage: everyone reads 12k/16k at a glance, and a percentage of a + // number that is itself an estimate suggests a precision this does not have. + String used = (estimated ? "~" : "") + abbreviate(inputTokens); + return contextSize <= UNKNOWN_CONTEXT ? "ctx " + used : "ctx " + used + "/" + abbreviate(contextSize); + } + + /** + * A short token count: {@code 812}, {@code 1.2k}, {@code 16k}. + * + * @param tokens the count + * @return the abbreviated form + */ + static String abbreviate(long tokens) { + if (tokens < 1000) { + return Long.toString(tokens); + } + double thousands = tokens / 1000.0; + // one decimal below 10k (1.2k), none above (16k) -- the decimal carries no information there + return thousands < 10 + ? String.format(Locale.ROOT, "%.1fk", thousands) + : String.format(Locale.ROOT, "%.0fk", thousands); + } +} diff --git a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/compact-prompt.txt b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/compact-prompt.txt new file mode 100644 index 000000000..2e025534a --- /dev/null +++ b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/compact-prompt.txt @@ -0,0 +1,11 @@ +Summarize the conversation so far so that it can continue in a fresh context with nothing important lost. Write the summary as plain text with these sections, and leave out a section only when it would be empty: + +1. Goal — what the user wants, in their own words where possible. +2. Facts — what was established about the project: paths, file names, commands, versions, decisions and their reasons. +3. Work done — what was changed or produced, file by file. +4. Problems — what failed, and what fixed it or is still open. +5. State — where things stand right now. +6. Next step — the single most obvious continuation, or "none" when the task is finished. + +Quote exact names, paths, commands and error messages verbatim; they are the part that cannot be reconstructed. Do not add advice, do not repeat the instructions, and do not use any tool — answer with the summary only. +{focus} diff --git a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/compact-prompt.txt.license b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/compact-prompt.txt.license new file mode 100644 index 000000000..b918686f7 --- /dev/null +++ b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/compact-prompt.txt.license @@ -0,0 +1,3 @@ +SPDX-FileCopyrightText: 2026 Bernard Ladenthin + +SPDX-License-Identifier: MIT diff --git a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt new file mode 100644 index 000000000..604e98310 --- /dev/null +++ b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt @@ -0,0 +1,13 @@ +Commands (everything else is sent to the model): + + /help this overview (/?, /commands) + /status endpoint, model, tools, mode, context use + /tools the tools offered to the model + /mode [manual|auto] show or set the approval mode (/approve) + /compact [focus] summarize the history and continue with the summary + /clear drop the history (/reset, /new) + /exit leave (/quit) + +In manual mode every tool that writes or runs a command asks first: [y]es runs it once, +[n]o tells the model the user cancelled it, [a]uto stops asking for the rest of the session. +Reading tools never ask. An unknown /command is sent to the model, not rejected. diff --git a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt.license b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt.license new file mode 100644 index 000000000..b918686f7 --- /dev/null +++ b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt.license @@ -0,0 +1,3 @@ +SPDX-FileCopyrightText: 2026 Bernard Ladenthin + +SPDX-License-Identifier: MIT diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ApprovalWireTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ApprovalWireTest.java new file mode 100644 index 000000000..cc62cc398 --- /dev/null +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ApprovalWireTest.java @@ -0,0 +1,166 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import static org.hamcrest.MatcherAssert.assertThat; +import static org.hamcrest.Matchers.contains; +import static org.hamcrest.Matchers.containsString; +import static org.hamcrest.Matchers.empty; +import static org.hamcrest.Matchers.hasSize; +import static org.hamcrest.Matchers.is; + +import com.fasterxml.jackson.databind.JsonNode; +import java.io.BufferedReader; +import java.io.ByteArrayOutputStream; +import java.io.PrintStream; +import java.io.StringReader; +import java.nio.charset.StandardCharsets; +import java.nio.file.Path; +import java.time.Duration; +import java.util.List; +import java.util.concurrent.CopyOnWriteArrayList; +import java.util.concurrent.atomic.AtomicReference; +import net.ladenthin.llama.server.OpenAiCompatServer; +import net.ladenthin.llama.server.OpenAiServerConfig; +import org.atmosphere.ai.RetryPolicy; +import org.atmosphere.ai.fs.AgentFileSystem; +import org.atmosphere.ai.fs.WorkspaceAgentFileSystem; +import org.atmosphere.ai.tool.ToolDefinition; +import org.junit.jupiter.api.Test; +import org.junit.jupiter.api.io.TempDir; + +/** + * The approval gate over the real wire: a denial must stop the tool and reach the model as a + * tool result, so it replans instead of believing the command ran. + * + *

Both halves matter and neither is ours: Atmosphere decides whether to ask + * ({@link ConsoleApprovalStrategy#policy()}), and Atmosphere turns the answer into the {@code role: + * "tool"} message. This test drives the whole path — scripted llama.cpp chunks through the real + * {@link OpenAiCompatServer}, the real tool loop, a console answer of {@code n} or {@code y} — and + * asserts what ends up on the wire. + */ +class ApprovalWireTest { + + private static final String MODEL_ID = "local-model"; + private static final String API_KEY = "k"; + private static final Duration TIMEOUT = Duration.ofSeconds(30); + + @TempDir + Path workspace; + + private final ByteArrayOutputStream console = new ByteArrayOutputStream(); + + private AgentRunner runner(OpenAiCompatServer server, List tools, String typed) { + AtomicReference mode = new AtomicReference<>(ApprovalMode.MANUAL); + PrintStream out = new PrintStream(console, true, StandardCharsets.UTF_8); + return new AgentRunner( + "http://127.0.0.1:" + server.getPort() + "/v1", + API_KEY, + MODEL_ID, + tools, + "You are a test agent.", + 0.0, + 64, + 10) + .retryPolicy(RetryPolicy.NONE) + .approval( + new ConsoleApprovalStrategy(mode, new BufferedReader(new StringReader(typed)), out, Ansi.PLAIN), + ConsoleApprovalStrategy.policy()); + } + + private ConsoleSession session() { + AgentFileSystem fs = new WorkspaceAgentFileSystem(workspace, AgentFileSystem.Limits.defaults()); + return new ConsoleSession(new PrintStream(console, true, StandardCharsets.UTF_8), fs); + } + + private static OpenAiServerConfig config() { + return OpenAiServerConfig.builder() + .host("127.0.0.1") + .port(0) + .apiKey(API_KEY) + .modelId(MODEL_ID) + .build(); + } + + private static ToolDefinition shellLike(String name, List invocations) { + return ToolDefinition.builder(name, "Test tool " + name) + .parameter("command", "The command line", "string", true) + .executor(args -> { + invocations.add(String.valueOf(args.get("command"))); + return "ran"; + }) + .build(); + } + + private static ScriptedBackend backend(String toolName) { + return new ScriptedBackend((call, request) -> call == 1 + ? ScriptedBackend.toolCallTurn("call_1", toolName, "{\"command\":\"rm -rf build\"}") + : ScriptedBackend.textTurn("Understood.")); + } + + private static String toolResult(JsonNode request) { + for (JsonNode message : request.path("messages")) { + if ("tool".equals(message.path("role").asText())) { + return message.path("content").asText(); + } + } + return ""; + } + + @Test + void aDeniedShellCallNeverRunsAndTheModelIsToldItWasCancelled() throws Exception { + List invocations = new CopyOnWriteArrayList<>(); + ScriptedBackend backend = backend(ShellTool.TOOL_NAME); + try (OpenAiCompatServer server = new OpenAiCompatServer(backend, config()).start()) { + AgentRunner runner = + runner(server, List.of(shellLike(ShellTool.TOOL_NAME, invocations)), "n" + System.lineSeparator()); + ConsoleSession session = session(); + + runner.run("Delete the build directory.", List.of(), session); + + assertThat(session.await(TIMEOUT), is(true)); + assertThat("the executor must not have run", invocations, is(empty())); + List requests = backend.requests(); + assertThat(requests, hasSize(2)); + // Atmosphere's own wording; the point is that the model is told, not what it says. + assertThat(toolResult(requests.get(1)), containsString("cancelled")); + assertThat(console.toString(StandardCharsets.UTF_8), containsString("rm -rf build")); + } + } + + @Test + void anApprovedShellCallRunsAndItsRealOutputIsSentBack() throws Exception { + List invocations = new CopyOnWriteArrayList<>(); + ScriptedBackend backend = backend(ShellTool.TOOL_NAME); + try (OpenAiCompatServer server = new OpenAiCompatServer(backend, config()).start()) { + AgentRunner runner = + runner(server, List.of(shellLike(ShellTool.TOOL_NAME, invocations)), "y" + System.lineSeparator()); + ConsoleSession session = session(); + + runner.run("Delete the build directory.", List.of(), session); + + assertThat(session.await(TIMEOUT), is(true)); + assertThat(invocations, contains("rm -rf build")); + assertThat(toolResult(backend.requests().get(1)), is("ran")); + } + } + + @Test + void aReadingToolIsNotGatedAndRunsWithoutAnyAnswer() throws Exception { + List invocations = new CopyOnWriteArrayList<>(); + ScriptedBackend backend = backend("ls"); + try (OpenAiCompatServer server = new OpenAiCompatServer(backend, config()).start()) { + // empty console input: if this tool asked, the strategy would read EOF and deny + AgentRunner runner = runner(server, List.of(shellLike("ls", invocations)), ""); + ConsoleSession session = session(); + + runner.run("List the files.", List.of(), session); + + assertThat(session.await(TIMEOUT), is(true)); + assertThat(invocations, contains("rm -rf build")); + assertThat(toolResult(backend.requests().get(1)), is("ran")); + } + } +} diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleApprovalStrategyTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleApprovalStrategyTest.java new file mode 100644 index 000000000..abe7d4629 --- /dev/null +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleApprovalStrategyTest.java @@ -0,0 +1,123 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import static org.hamcrest.MatcherAssert.assertThat; +import static org.hamcrest.Matchers.containsString; +import static org.hamcrest.Matchers.is; + +import java.io.BufferedReader; +import java.io.ByteArrayOutputStream; +import java.io.PrintStream; +import java.io.StringReader; +import java.nio.charset.StandardCharsets; +import java.time.Duration; +import java.time.Instant; +import java.util.Map; +import java.util.concurrent.atomic.AtomicReference; +import org.atmosphere.ai.approval.ApprovalStrategy.ApprovalOutcome; +import org.atmosphere.ai.approval.PendingApproval; +import org.atmosphere.ai.approval.ToolApprovalPolicy; +import org.atmosphere.ai.tool.ToolDefinition; +import org.junit.jupiter.api.Test; + +class ConsoleApprovalStrategyTest { + + private final ByteArrayOutputStream console = new ByteArrayOutputStream(); + + private PendingApproval approval() { + return new PendingApproval( + "id-1", + ShellTool.TOOL_NAME, + Map.of("command", "rm -rf build"), + null, + "console", + Instant.now().plus(Duration.ofMinutes(5))); + } + + private ApprovalOutcome ask(AtomicReference mode, String typed) { + BufferedReader reader = typed == null ? null : new BufferedReader(new StringReader(typed)); + PrintStream out = new PrintStream(console, true, StandardCharsets.UTF_8); + // The strategy never touches the session; Atmosphere passes it only so a UI can emit events. + return new ConsoleApprovalStrategy(mode, reader, out, Ansi.PLAIN).awaitApproval(approval(), null); + } + + private String consoleText() { + return console.toString(StandardCharsets.UTF_8); + } + + @Test + void yesRunsTheToolOnceAndKeepsAsking() { + AtomicReference mode = new AtomicReference<>(ApprovalMode.MANUAL); + + assertThat(ask(mode, "y\n"), is(ApprovalOutcome.APPROVED)); + assertThat(mode.get(), is(ApprovalMode.MANUAL)); + assertThat(consoleText(), containsString(ShellTool.TOOL_NAME)); + assertThat(consoleText(), containsString("rm -rf build")); + } + + @Test + void anEmptyAnswerMeansYes() { + assertThat(ask(new AtomicReference<>(ApprovalMode.MANUAL), "\n"), is(ApprovalOutcome.APPROVED)); + } + + @Test + void noDeniesAndAtmosphereTellsTheModel() { + // The cancellation text itself is Atmosphere's ("Action cancelled by user"), which is why this + // side only has to return DENIED -- see ApprovalWireTest for the message reaching the model. + assertThat(ask(new AtomicReference<>(ApprovalMode.MANUAL), "n\n"), is(ApprovalOutcome.DENIED)); + } + + @Test + void autoApprovesAndStopsAskingForTheRestOfTheSession() { + AtomicReference mode = new AtomicReference<>(ApprovalMode.MANUAL); + + assertThat(ask(mode, "a\n"), is(ApprovalOutcome.APPROVED)); + assertThat(mode.get(), is(ApprovalMode.AUTO)); + // the next call must not read anything: an empty reader would otherwise mean "input closed" + assertThat(ask(mode, ""), is(ApprovalOutcome.APPROVED)); + } + + @Test + void anUnreadableAnswerIsAskedAgain() { + assertThat(ask(new AtomicReference<>(ApprovalMode.MANUAL), "maybe\ny\n"), is(ApprovalOutcome.APPROVED)); + assertThat(consoleText(), containsString("please answer y, n or a")); + } + + @Test + void withoutAConsoleTheAnswerIsNo() { + // One-shot mode: nobody can answer, so the safe outcome is a denial with a printed reason -- + // never a silent auto-approval, which would make an unattended run the most permissive one. + assertThat(ask(new AtomicReference<>(ApprovalMode.MANUAL), null), is(ApprovalOutcome.DENIED)); + assertThat(consoleText(), containsString("--auto")); + + assertThat(ask(new AtomicReference<>(ApprovalMode.AUTO), null), is(ApprovalOutcome.APPROVED)); + } + + @Test + void closedInputDeniesRatherThanBlocking() { + assertThat(ask(new AtomicReference<>(ApprovalMode.MANUAL), ""), is(ApprovalOutcome.DENIED)); + assertThat(consoleText(), containsString("input closed")); + } + + @Test + void writingToolsAndTheShellAreGatedReadingToolsAreNot() { + ToolApprovalPolicy policy = ConsoleApprovalStrategy.policy(); + + for (String gated : new String[] {ShellTool.TOOL_NAME, "write_file", "edit_file", "delete", "rename"}) { + assertThat(gated, policy.requiresApproval(stub(gated)), is(true)); + } + for (String free : new String[] {"ls", "read_file", "glob", "grep"}) { + assertThat(free, policy.requiresApproval(stub(free)), is(false)); + } + assertThat( + ConsoleApprovalStrategy.gated(java.util.List.of("ls", "write_file", ShellTool.TOOL_NAME)), + is(java.util.List.of("write_file", ShellTool.TOOL_NAME))); + } + + private static ToolDefinition stub(String name) { + return ToolDefinition.builder(name, "test").executor(args -> "").build(); + } +} diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleFormattingTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleFormattingTest.java new file mode 100644 index 000000000..494ad180d --- /dev/null +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleFormattingTest.java @@ -0,0 +1,130 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import static org.hamcrest.MatcherAssert.assertThat; +import static org.hamcrest.Matchers.containsString; +import static org.hamcrest.Matchers.is; +import static org.hamcrest.Matchers.not; + +import java.io.ByteArrayOutputStream; +import java.io.PrintStream; +import java.nio.charset.StandardCharsets; +import java.util.Map; +import org.junit.jupiter.api.Test; + +/** The console cosmetics: colour decision, status line, and the streaming Markdown renderer. */ +class ConsoleFormattingTest { + + // ----- Ansi ----- + + private static Ansi detect(Map env, boolean terminal) { + return Ansi.detect(env::get, () -> terminal); + } + + @Test + void colourIsOnOnlyOnATerminal() { + assertThat(detect(Map.of(), true).isEnabled(), is(true)); + assertThat(detect(Map.of(), false).isEnabled(), is(false)); + } + + @Test + void theEnvironmentCanForceColourOnOrOff() { + // NO_COLOR: "when present and not an empty string ... prevents the addition of ANSI color" + assertThat(detect(Map.of("NO_COLOR", "1"), true).isEnabled(), is(false)); + assertThat(detect(Map.of("NO_COLOR", ""), true).isEnabled(), is(true)); + assertThat(detect(Map.of("TERM", "dumb"), true).isEnabled(), is(false)); + assertThat(detect(Map.of("CLICOLOR", "0"), true).isEnabled(), is(false)); + // forcing wins over "not a terminal", e.g. when piping into a pager + assertThat(detect(Map.of("CLICOLOR_FORCE", "1"), false).isEnabled(), is(true)); + // ... but NO_COLOR is checked after CLICOLOR_FORCE, which is the documented precedence + assertThat(detect(Map.of("CLICOLOR_FORCE", "1", "NO_COLOR", "1"), false).isEnabled(), is(true)); + } + + @Test + void aPlainInstanceReturnsTheTextUnchanged() { + assertThat(Ansi.PLAIN.bold("x") + Ansi.PLAIN.dim("y") + Ansi.PLAIN.red("z"), is("xyz")); + assertThat(detect(Map.of(), true).bold("x"), containsString("\u001b[")); + } + + // ----- StatusLine ----- + + @Test + void theStatusLineShowsModeContextToolsAndModel() { + String line = StatusLine.render(ApprovalMode.MANUAL, 1234, false, 16384, 9, "local-model"); + + assertThat(line, is("[manual · ctx 1.2k/16k · 9 tools · local-model]")); + } + + @Test + void contextIsShownInThousandsAndWithoutASizeWhenItIsUnknown() { + assertThat(StatusLine.context(812, false, 32768), is("ctx 812/33k")); + assertThat(StatusLine.context(16000, false, 32768), is("ctx 16k/33k")); + assertThat(StatusLine.context(0, false, StatusLine.UNKNOWN_CONTEXT), is("ctx 0")); + assertThat(StatusLine.context(2500, false, StatusLine.UNKNOWN_CONTEXT), is("ctx 2.5k")); + } + + @Test + void anEstimatedCountIsMarkedWithATilde() { + // llama.cpp reports usage only to clients that ask for it, and Atmosphere does not, so the + // number normally comes from LocalAgent.estimateTokens -- the tilde says so. + assertThat(StatusLine.context(2500, true, 16384), is("ctx ~2.5k/16k")); + assertThat( + LocalAgent.estimateTokens( + "0123456789", java.util.List.of(org.atmosphere.ai.llm.ChatMessage.user("0123456789"))), + is(5L)); + } + + // ----- MarkdownConsole ----- + + private static String render(String text, Ansi ansi) { + ByteArrayOutputStream buffer = new ByteArrayOutputStream(); + MarkdownConsole console = new MarkdownConsole(new PrintStream(buffer, true, StandardCharsets.UTF_8), ansi); + // one character at a time: the renderer must not depend on where the stream splits + for (int i = 0; i < text.length(); i++) { + console.append(text.substring(i, i + 1)); + } + console.flush(); + return buffer.toString(StandardCharsets.UTF_8); + } + + @Test + void withoutColourTheTextIsPassedThroughUnchanged() { + String markdown = "# Title\n\nSome **bold** and `code`.\n- one\n- two\n"; + + assertThat( + render(markdown, Ansi.PLAIN), + is("Title\n\nSome bold and code.\n• one\n• two\n".replace("\n", System.lineSeparator()))); + } + + @Test + void headingsBulletsAndInlineSpansAreStyled() { + Ansi ansi = detect(Map.of(), true); + String out = render("## Heading\n- item with **bold**\ntext with `code` inside\n", ansi); + + assertThat(out, containsString(ansi.bold("Heading"))); + assertThat(out, containsString(ansi.cyan("•"))); + assertThat(out, containsString(ansi.bold("bold"))); + assertThat(out, containsString(ansi.cyan("code"))); + // the markers themselves are gone, that is the point of rendering + assertThat(out, not(containsString("**"))); + assertThat(out, not(containsString("##"))); + } + + @Test + void aFencedBlockIsStyledAsAWholeAndNotParsedInside() { + Ansi ansi = detect(Map.of(), true); + String out = render("```java\nint a = b * c; // **not bold**\n```\n", ansi); + + assertThat(out, containsString(ansi.cyan("int a = b * c; // **not bold**"))); + } + + @Test + void anUnclosedMarkerIsLeftAsTyped() { + // A half-streamed "**bold" must never swallow the rest of the line. + assertThat(render("a **b\n", Ansi.PLAIN), is("a **b" + System.lineSeparator())); + assertThat(render("2 * 3 * 4\n", Ansi.PLAIN), is("2 * 3 * 4" + System.lineSeparator())); + } +} diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java index dd1566307..7e80f3d4b 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java @@ -76,7 +76,8 @@ void oneShotTurnReadsAWorkspaceFileThroughTheBuiltInFileTools() throws Exception assertThat(exit, is(0)); } String console = out.toString(StandardCharsets.UTF_8); - assertThat(console, containsString("⚙ read_file {path=hello.txt}")); + // the tool line and its result, as ConsoleSession renders them (unstyled here: not a terminal) + assertThat(console, containsString("● read_file {path=hello.txt}")); assertThat(console, containsString("↳ VALUE=42")); assertThat(console, containsString("The file says VALUE=42.")); List requests = backend.requests(); diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ServerPropsTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ServerPropsTest.java new file mode 100644 index 000000000..169df860c --- /dev/null +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ServerPropsTest.java @@ -0,0 +1,45 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import static org.hamcrest.MatcherAssert.assertThat; +import static org.hamcrest.Matchers.is; + +import java.util.List; +import org.junit.jupiter.api.Test; + +class ServerPropsTest { + + @Test + void propsIsTriedNextToTheV1RoutesFirst() { + // llama-server serves /props at the root; this project's server serves it under both. + assertThat( + ServerProps.candidates("http://127.0.0.1:8080/v1"), + is(List.of("http://127.0.0.1:8080/props", "http://127.0.0.1:8080/v1/props"))); + assertThat( + ServerProps.candidates("http://127.0.0.1:8080/v1/"), + is(List.of("http://127.0.0.1:8080/props", "http://127.0.0.1:8080/v1/props"))); + assertThat(ServerProps.candidates("http://host/api"), is(List.of("http://host/api/props"))); + } + + @Test + void theContextSizeIsReadFromTheGenerationDefaults() { + assertThat( + ServerProps.parseContextSize("{\"default_generation_settings\":{\"n_ctx\":16384,\"model\":\"m\"}}"), + is(16384)); + } + + @Test + void anythingElseIsUnknownRatherThanAGuess() { + assertThat(ServerProps.parseContextSize("{}"), is(StatusLine.UNKNOWN_CONTEXT)); + assertThat(ServerProps.parseContextSize("not json"), is(StatusLine.UNKNOWN_CONTEXT)); + assertThat(ServerProps.parseContextSize("{\"n_ctx\":\"many\"}"), is(StatusLine.UNKNOWN_CONTEXT)); + } + + @Test + void anUnreachableServerIsUnknownAndDoesNotThrow() { + assertThat(ServerProps.contextSize("http://127.0.0.1:1/v1", "k"), is(StatusLine.UNKNOWN_CONTEXT)); + } +} diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/SlashCommandsTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/SlashCommandsTest.java new file mode 100644 index 000000000..9426a00b0 --- /dev/null +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/SlashCommandsTest.java @@ -0,0 +1,57 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import static org.hamcrest.MatcherAssert.assertThat; +import static org.hamcrest.Matchers.containsString; +import static org.hamcrest.Matchers.is; + +import java.util.Optional; +import org.junit.jupiter.api.Test; + +class SlashCommandsTest { + + @Test + void aCommandIsRecognisedWithItsAliasesAndCase() { + assertThat(SlashCommands.parse("/help").orElseThrow().command(), is(SlashCommands.Command.HELP)); + assertThat(SlashCommands.parse("/?").orElseThrow().command(), is(SlashCommands.Command.HELP)); + assertThat(SlashCommands.parse(" /HELP ").orElseThrow().command(), is(SlashCommands.Command.HELP)); + assertThat(SlashCommands.parse("/quit").orElseThrow().command(), is(SlashCommands.Command.EXIT)); + assertThat(SlashCommands.parse("/reset").orElseThrow().command(), is(SlashCommands.Command.CLEAR)); + assertThat(SlashCommands.parse("/approve").orElseThrow().command(), is(SlashCommands.Command.MODE)); + } + + @Test + void theRestOfTheLineIsTheArgument() { + SlashCommands compact = + SlashCommands.parse("/compact focus on the build errors").orElseThrow(); + + assertThat(compact.command(), is(SlashCommands.Command.COMPACT)); + assertThat(compact.arguments(), is("focus on the build errors")); + assertThat(compact.hasArguments(), is(true)); + assertThat(SlashCommands.parse("/compact").orElseThrow().hasArguments(), is(false)); + assertThat(SlashCommands.parse("/mode auto").orElseThrow().arguments(), is("auto")); + } + + @Test + void anythingElseIsAMessageForTheModel() { + // An unknown command is NOT rejected: a line may legitimately start with a slash, and a user + // who types /halp would rather get an answer than an error. The trade-off is deliberate. + assertThat(SlashCommands.parse("/halp"), is(Optional.empty())); + assertThat(SlashCommands.parse("/usr/bin/env is where?"), is(Optional.empty())); + assertThat(SlashCommands.parse("what does /help do?"), is(Optional.empty())); + assertThat(SlashCommands.parse(""), is(Optional.empty())); + assertThat(SlashCommands.parse(null), is(Optional.empty())); + } + + @Test + void everyCommandIsDocumentedInTheHelpText() { + String help = LocalAgent.prompt(LocalAgent.HELP_TEXT); + + for (SlashCommands.Command command : SlashCommands.Command.values()) { + assertThat(help, containsString(command.canonicalName())); + } + } +} From 02d0d049ae5c0584a31996de1c3983bd2c06ea54 Mon Sep 17 00:00:00 2001 From: Bernard Ladenthin Date: Wed, 23 Sep 2026 00:09:03 +0200 Subject: [PATCH 02/36] llama-atmosphere-agent: add a "Try it" walkthrough to the README The new commands, the approval prompt and the rendering are documented feature by feature, but nothing showed what a first session looks like. Adds a short table: /help, /status, /tools, a shell call denied then approved, /mode auto, a Markdown answer, /compact, /exit -- in the order that exercises each piece once. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01E2h8gXyyE5UeimkL9vQv9G --- llama-atmosphere-agent/README.md | 28 ++++++++++++++++++++++++++++ 1 file changed, 28 insertions(+) diff --git a/llama-atmosphere-agent/README.md b/llama-atmosphere-agent/README.md index 56d39e333..f8728699f 100644 --- a/llama-atmosphere-agent/README.md +++ b/llama-atmosphere-agent/README.md @@ -223,6 +223,34 @@ into a file stays correct. Colour is on only on a real terminal and obeys `NO_CO `CLICOLOR=0` and `CLICOLOR_FORCE=1`. On the classic Windows `conhost.exe` escape sequences may show up literally unless `HKCU\Console\VirtualTerminalLevel` is 1 — Windows Terminal needs nothing. +### Try it + +Start the agent with shell access (add `-Dllama.classifier=…` and `--ngl 99` for a GPU; leave both +out to stay on the CPU): + +```bash +mvn -q compile exec:java \ + -Dexec.args="--model models/Qwen3-4B-Instruct-2507-Q4_K_M.gguf --ctx-size 16384 --workspace /path/to/project --allow-shell" +``` + +Then, in this order: + +| Type this | What should happen | +|---|---| +| `/help` | the command overview — the agent answers, the model never sees the line | +| `/status` | mode, context use, tools, model, workspace, history size | +| `/tools` | every tool, and which of them ask before running | +| `docker is running locally, list the images` | `? run_command {command=docker images}` and the prompt `[y]es / [n]o / [a]uto` | +| answer `n` | the command does **not** run; the model is told it was cancelled and offers an alternative | +| ask again, answer `y` | the command runs and its output goes back to the model | +| `/mode auto` | the status line flips to `auto`; nothing asks any more | +| `explain Markdown with a heading, a list, bold text and a code block` | the answer arrives rendered: heading bold, `•` bullets, code in colour | +| `/compact` | the conversation is summarized and replaces the history; `ctx` drops | +| `/exit` | leave | + +`--auto` starts in auto mode, `--verbose` brings llama.cpp's own log back, and `NO_COLOR=1` turns +the styling off. + ### The system prompt Without `--system` the agent uses a built-in **general-purpose** prompt: it names the file tools and From 1d6efbd74da09990ff95558ed8c1c3e362c406f5 Mon Sep 17 00:00:00 2001 From: Bernard Ladenthin Date: Wed, 23 Sep 2026 00:25:47 +0200 Subject: [PATCH 03/36] llama-atmosphere-agent: a real terminal via JLine (editing, history, pinned status, single key) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The REPL read lines with a BufferedReader: arrow keys produced escape sequences, there was no history, the status line scrolled away with everything else, and every approval answer needed Enter. JLine fixes all four, behind one interface so nothing else has to know: - AgentTerminal with two implementations. JLineTerminal on a real terminal: line editing, history, Tab completion of the command names, a status line (plus a rule) pinned to the bottom via JLine's Status, streamed output through LineReader.printAbove so that block stays put, single-key answers via enterRawMode, Ctrl-C drops the line instead of the session. PlainTerminal otherwise: PrintStream + BufferedReader, no cursor control, correct when the output is a file. JLineTerminal.open returns null (rather than throwing) when there is no usable terminal, and the caller falls back -- which is also why no test needs a TTY. - While a turn runs the status line shows progress: "⠙ working… (12s · 2 tool calls)". A local model can think for a while, and a silent console is indistinguishable from a hung one. - ConsoleSession, MarkdownConsole and ConsoleApprovalStrategy now write lines to the terminal instead of a PrintStream; MarkdownConsole takes a Consumer. One dependency: org.jline:jline 4.4.5, a single jar with no transitive dependencies. Verified on Windows 11 that JLine picks the windows-vtp provider, so the pinned status line, single keys, Tab completion and Ctrl-C all work there. Tests: 66 green (was 61); PlainTerminalTest covers the fallback console and that every command name is offered for completion. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01E2h8gXyyE5UeimkL9vQv9G --- CLAUDE.md | 16 +- llama-atmosphere-agent/README.md | 13 +- llama-atmosphere-agent/pom.xml | 10 ++ .../llama/atmosphere/AgentTerminal.java | 67 ++++++++ .../atmosphere/ConsoleApprovalStrategy.java | 60 +++---- .../llama/atmosphere/ConsoleSession.java | 39 ++--- .../llama/atmosphere/JLineTerminal.java | 146 ++++++++++++++++++ .../llama/atmosphere/LocalAgent.java | 124 ++++++++++----- .../llama/atmosphere/MarkdownConsole.java | 16 +- .../llama/atmosphere/PlainTerminal.java | 87 +++++++++++ .../llama/atmosphere/ApprovalWireTest.java | 15 +- .../AtmosphereToolLoopIntegrationTest.java | 5 +- .../AtmosphereWireContractTest.java | 5 +- .../ConsoleApprovalStrategyTest.java | 5 +- .../atmosphere/ConsoleFormattingTest.java | 10 +- .../llama/atmosphere/PlainTerminalTest.java | 87 +++++++++++ 16 files changed, 566 insertions(+), 139 deletions(-) create mode 100644 llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentTerminal.java create mode 100644 llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java create mode 100644 llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/PlainTerminal.java create mode 100644 llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/PlainTerminalTest.java diff --git a/CLAUDE.md b/CLAUDE.md index 23a7104c9..089474d42 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -2233,7 +2233,8 @@ a `jvm.config` takes no comments, so REUSE can only read its metadata from that `REUSE Compliance Check` job fails on `main` — which is how it was found, the PR run having been cancelled. Spotless (palantir) is configured in its own pom; the model-free CI job runs `spotless:check`. -**The REPL layer (commands, approval, status line, rendering).** Five small classes, no new dependency: +**The REPL layer (commands, approval, status line, rendering).** Eight small classes and one dependency +(`org.jline:jline`, one jar, no transitive deps): `SlashCommands` (a line starting with `/` whose first word names a command is handled locally — `/help /status /tools /mode /compact /clear /exit`; **an unknown `/command` goes to the model**, which is why no escape syntax is needed for `/usr/bin/…`), `ApprovalMode` + `ConsoleApprovalStrategy` @@ -2249,14 +2250,23 @@ are decisions, not details: 2. **One-shot (`--prompt`) denies a gated call** instead of auto-approving it — `--auto` is the deliberate opt-in. Atmosphere itself fails closed when no strategy is wired, and this keeps that direction: an unattended run must not be the most permissive one. -3. **The answer is rendered append-only, one completed line at a time** (`MarkdownConsole`). Redrawing +3. **`AgentTerminal` has exactly two implementations, chosen once at startup.** `JLineTerminal` (a real + terminal: line editing, history, Tab completion of the command names, a status line pinned to the + bottom via JLine's `Status`, single-key answers through `enterRawMode`, streamed output via + `LineReader.printAbove` so the bottom block stays put) and `PlainTerminal` (a `PrintStream` plus a + `BufferedReader`: no cursor control at all, correct when the output is a file). `JLineTerminal.open` + returns **null** instead of throwing when there is no usable terminal — piped input, a dumb + terminal, a missing native provider — and the caller falls back. Every test drives `PlainTerminal`, + which is why none of them needs a TTY. Verified on Windows: JLine picks the `windows-vtp` provider, + so ANSI works there without the registry caveat. +4. **The answer is rendered append-only, one completed line at a time** (`MarkdownConsole`). Redrawing on every token is what produces the known overdraw/truncation bugs in the Ink/Bubble-Tea based clients and breaks when the output is piped. Only headings, bullets, fences and inline `**bold**`/`` `code` `` are handled; italics deliberately are not (`*` is more often a glob than emphasis). Colour is decided once in `Ansi.detect()` — `CLICOLOR_FORCE`, then `NO_COLOR`, then `TERM=dumb`/`CLICOLOR=0`, else "is a terminal" via `Console.isTerminal()` (reflective: JDK 22+; below that `System.console() != null`). -4. **The context number in the status line is an estimate, marked `~`.** llama.cpp emits its usage +5. **The context number in the status line is an estimate, marked `~`.** llama.cpp emits its usage chunk only when the client sets `stream_options.include_usage`, and Atmosphere's client does not; `ConsoleSession.usage()` takes the real count when one arrives, otherwise `LocalAgent.estimateTokens` uses four characters per token. The window size is `--ctx-size` (in-process) or the server's diff --git a/llama-atmosphere-agent/README.md b/llama-atmosphere-agent/README.md index f8728699f..6e9e7870a 100644 --- a/llama-atmosphere-agent/README.md +++ b/llama-atmosphere-agent/README.md @@ -102,7 +102,7 @@ Metal with the default jar already. The root README's classifier table lists eve -Dexec.args="--base-url http://127.0.0.1:8080/v1 --workspace /path/to/project --allow-shell" ``` -A `you>` prompt appears, above it a status line. The answer streams as it is generated, and every +A `you>` prompt appears, with a status line pinned to the bottom of the window. The answer streams as it is generated, and every tool call and its result are printed as `● read_file {path=…}` / `↳ …` lines. See [Commands, approval and the status line](#commands-approval-and-the-status-line). @@ -188,6 +188,13 @@ unknown `/command` included — goes to the model: | `/clear` (`/reset`, `/new`) | drop the history | | `/exit` (`/quit`) | leave | +**The prompt.** On a real terminal the agent uses [JLine](https://github.com/jline/jline3): arrow keys +and the usual editing shortcuts work, ↑ recalls earlier lines, Tab completes the commands, Ctrl-C +drops the current line and Ctrl-D leaves. The status line and the rule above it stay at the bottom +while the answer scrolls past, and while a turn runs that line shows what is going on +(`⠙ working… (12s · 2 tool calls)`). With piped input, in one-shot mode and wherever JLine finds no +terminal, everything falls back to plain `println`/`readLine` — same features, no cursor tricks. + **Approval.** In the default `manual` mode every tool that writes or runs a command — `run_command`, `write_file`, `edit_file`, `delete`, `rename` — asks before it runs: @@ -198,7 +205,8 @@ unknown `/command` included — goes to the model: ``` `[y]` runs it once, `[n]` cancels it *and tells the model*, so it replans instead of assuming the -command ran, `[a]` switches to `auto` for the rest of the session (`/mode manual` switches back). +command ran, `[a]` switches to `auto` for the rest of the session (`/mode manual` switches back). On a +terminal a single key is enough — no Enter; Enter alone also means yes, and Ctrl-C means no. Reading tools (`ls`, `read_file`, `glob`, `grep`) never ask. **In one-shot mode (`--prompt`) nobody can answer, so a gated call is denied** — pass `--auto` to run unattended. The gate itself is Atmosphere's (`ToolApprovalPolicy` + `ApprovalStrategy`); the agent only supplies the question and @@ -245,6 +253,7 @@ Then, in this order: | ask again, answer `y` | the command runs and its output goes back to the model | | `/mode auto` | the status line flips to `auto`; nothing asks any more | | `explain Markdown with a heading, a list, bold text and a code block` | the answer arrives rendered: heading bold, `•` bullets, code in colour | +| press ↑ | the previous line comes back; Tab after `/` completes the commands | | `/compact` | the conversation is summarized and replaces the history; `ctx` drops | | `/exit` | leave | diff --git a/llama-atmosphere-agent/pom.xml b/llama-atmosphere-agent/pom.xml index 34c19ba1f..fa2fc4d04 100644 --- a/llama-atmosphere-agent/pom.xml +++ b/llama-atmosphere-agent/pom.xml @@ -48,6 +48,7 @@ SPDX-License-Identifier: MIT e.g. -Dllama.classifier=cuda13-linux-x86-64 or vulkan-windows-x86-64 (runtime on PATH). --> 4.0.70 + 4.4.5 2.0.19 1.0.1 6.1.3 @@ -91,6 +92,15 @@ SPDX-License-Identifier: MIT ${atmosphere.version} + + + org.jline + jline + ${jline.version} + + org.jspecify jspecify diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentTerminal.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentTerminal.java new file mode 100644 index 000000000..3f08134ba --- /dev/null +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentTerminal.java @@ -0,0 +1,67 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import org.jspecify.annotations.Nullable; + +/** + * Everything the agent needs from a console, so that a real terminal and a plain stream are + * interchangeable. + * + *

Two implementations: {@link JLineTerminal} for an interactive session (line editing, history, + * tab completion, a status line pinned to the bottom, single-key answers) and {@link PlainTerminal} + * for everything else — one-shot runs, piped input, and the tests. The agent picks one at startup and + * never branches again. + * + *

Output is line-oriented on purpose: a line is written once, complete, and is never touched + * again. That is what lets the pinned status line coexist with streamed output without redrawing + * anything the reader has already seen. + */ +public interface AgentTerminal extends AutoCloseable { + + /** + * Write one completed line, above the prompt and the status line. + * + * @param text the line, without a terminator + */ + void line(String text); + + /** + * Read one line from the user. + * + * @param prompt the prompt to show, e.g. {@code "you> "} + * @return the line, or {@code null} at end of input + */ + @Nullable + String readLine(String prompt); + + /** + * Read a single answer, without waiting for Enter where the terminal allows it. + * + * @param prompt the question to show + * @return the answer in lower case (a single key, or a whole line on a plain stream), or + * {@code null} at end of input + */ + @Nullable + String readKey(String prompt); + + /** + * Set the status line kept at the bottom of the window. + * + * @param text the line; an empty string removes it + */ + void status(String text); + + /** + * The styles to use for this console. + * + * @return a colouring instance on a terminal, a plain one otherwise + */ + Ansi ansi(); + + /** Restore the terminal. */ + @Override + void close(); +} diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleApprovalStrategy.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleApprovalStrategy.java index 3283257c7..8b0828de5 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleApprovalStrategy.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleApprovalStrategy.java @@ -4,11 +4,7 @@ package net.ladenthin.llama.atmosphere; -import java.io.BufferedReader; -import java.io.IOException; -import java.io.PrintStream; import java.util.List; -import java.util.Locale; import java.util.Set; import java.util.concurrent.atomic.AtomicReference; import org.atmosphere.ai.StreamingSession; @@ -16,7 +12,6 @@ import org.atmosphere.ai.approval.ApprovalStrategy; import org.atmosphere.ai.approval.PendingApproval; import org.atmosphere.ai.approval.ToolApprovalPolicy; -import org.jspecify.annotations.Nullable; /** * Asks on the console before a tool that changes something runs: {@code [y]es / [n]o / [a]uto}. @@ -50,8 +45,8 @@ public final class ConsoleApprovalStrategy implements ApprovalStrategy { private static final int ARGUMENT_PREVIEW_CHARS = 300; private final AtomicReference mode; - private final @Nullable BufferedReader input; - private final PrintStream out; + private final AgentTerminal terminal; + private final boolean interactive; private final Ansi ansi; /** @@ -59,16 +54,14 @@ public final class ConsoleApprovalStrategy implements ApprovalStrategy { * * @param mode the shared, mutable approval mode (also written by {@code /mode} and by an * {@code [a]} answer) - * @param input the console the user answers on, or {@code null} when nobody can be asked - * @param out where the prompt is printed - * @param ansi the styles for the question + * @param terminal where the question is asked + * @param interactive whether anybody can answer at all ({@code false} for a one-shot run) */ - public ConsoleApprovalStrategy( - AtomicReference mode, @Nullable BufferedReader input, PrintStream out, Ansi ansi) { + public ConsoleApprovalStrategy(AtomicReference mode, AgentTerminal terminal, boolean interactive) { this.mode = mode; - this.input = input; - this.out = out; - this.ansi = ansi; + this.terminal = terminal; + this.interactive = interactive; + this.ansi = terminal.ansi(); } /** @@ -100,28 +93,22 @@ public ApprovalResolution awaitApprovalDetailed(PendingApproval approval, Stream if (mode.get() == ApprovalMode.AUTO) { return ApprovalResolution.approve(); } - if (input == null) { - out.println(); - out.println(ansi.red("✗ " + approval.toolName() + " " + preview(approval) + if (!interactive) { + terminal.line(ansi.red("✗ " + approval.toolName() + " " + preview(approval) + " — denied: no console to ask (run with --auto to allow tools unattended)")); - out.flush(); return ApprovalResolution.deny(); } - out.println(); - out.println(ansi.yellow("? " + approval.toolName()) + " " + ansi.dim(preview(approval))); + terminal.line(ansi.yellow("? " + approval.toolName()) + " " + ansi.dim(preview(approval))); while (true) { - out.print(ansi.yellow(" allow? [y]es / [n]o / [a]uto (no more questions): ")); - out.flush(); - String answer = readLine(); + String answer = terminal.readKey(ansi.yellow(" allow? [y]es / [n]o / [a]uto (no more questions): ")); if (answer == null) { - // stdin closed mid-turn: the same situation as having no console at all - out.println(); - out.println(" denied (input closed)"); - out.flush(); + // input closed mid-turn: the same situation as having no console at all + terminal.line(" denied (input closed)"); return ApprovalResolution.deny(); } - switch (answer.trim().toLowerCase(Locale.ROOT)) { - case "y", "yes", "" -> { + switch (answer) { + // "\r" / "\n": Enter in raw mode, taken as yes like an empty line on a plain stream + case "y", "yes", "", "\r", "\n" -> { return ApprovalResolution.approve(); } case "n", "no" -> { @@ -129,23 +116,14 @@ public ApprovalResolution awaitApprovalDetailed(PendingApproval approval, Stream } case "a", "auto" -> { mode.set(ApprovalMode.AUTO); - out.println(" approval mode: auto (use /mode manual to ask again)"); - out.flush(); + terminal.line(" approval mode: auto (use /mode manual to ask again)"); return ApprovalResolution.approve(); } - default -> out.println(" please answer y, n or a"); + default -> terminal.line(" please answer y, n or a"); } } } - private @Nullable String readLine() { - try { - return input == null ? null : input.readLine(); - } catch (IOException e) { - return null; - } - } - private static String preview(PendingApproval approval) { String arguments = String.valueOf(approval.arguments()); return arguments.length() <= ARGUMENT_PREVIEW_CHARS diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleSession.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleSession.java index c5410208a..6d7325ee6 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleSession.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleSession.java @@ -4,7 +4,6 @@ package net.ladenthin.llama.atmosphere; -import java.io.PrintStream; import java.time.Duration; import java.util.List; import java.util.Map; @@ -30,7 +29,7 @@ public final class ConsoleSession implements StreamingSession { private static final int RESULT_PREVIEW_CHARS = 400; - private final PrintStream out; + private final AgentTerminal terminal; private final Ansi ansi; private final MarkdownConsole markdown; private final Map, Object> injectables; @@ -42,26 +41,15 @@ public final class ConsoleSession implements StreamingSession { private volatile long inputTokens; /** - * Create an unstyled session printing to {@code out}. + * Create a session writing to {@code terminal}. * - * @param out where streamed text and tool lines go + * @param terminal where streamed text and tool lines go * @param fileSystem the workspace-confined filesystem handed to the file tools */ - public ConsoleSession(PrintStream out, AgentFileSystem fileSystem) { - this(out, fileSystem, Ansi.PLAIN); - } - - /** - * Create a session printing to {@code out}. - * - * @param out where streamed text and tool lines go - * @param fileSystem the workspace-confined filesystem handed to the file tools - * @param ansi the styles for the answer and the tool lines - */ - public ConsoleSession(PrintStream out, AgentFileSystem fileSystem, Ansi ansi) { - this.out = out; - this.ansi = ansi; - this.markdown = new MarkdownConsole(out, ansi); + public ConsoleSession(AgentTerminal terminal, AgentFileSystem fileSystem) { + this.terminal = terminal; + this.ansi = terminal.ansi(); + this.markdown = new MarkdownConsole(terminal::line, ansi); this.injectables = Map.of(AgentFileSystem.class, fileSystem); } @@ -115,7 +103,6 @@ public void progress(String message) { @Override public void complete() { markdown.flush(); - out.flush(); done.countDown(); } @@ -131,8 +118,7 @@ public void complete(String summary) { public void error(Throwable t) { failure = t; markdown.flush(); - out.println(ansi.red("[error] " + t)); - out.flush(); + terminal.line(ansi.red("[error] " + t)); done.countDown(); } @@ -147,17 +133,14 @@ public void emit(AiEvent event) { case AiEvent.ToolStart start -> { toolCalls++; markdown.flush(); - out.println(ansi.green("●") + " " + ansi.bold(start.toolName()) + " " + terminal.line(ansi.green("●") + " " + ansi.bold(start.toolName()) + " " + ansi.dim(String.valueOf(start.arguments()))); - out.flush(); } case AiEvent.ToolResult result -> { - out.println(ansi.dim(" ↳ " + preview(String.valueOf(result.result())))); - out.flush(); + terminal.line(ansi.dim(" ↳ " + preview(String.valueOf(result.result())))); } case AiEvent.ToolError error -> { - out.println(ansi.red(" ↳ error: " + error.error())); - out.flush(); + terminal.line(ansi.red(" ↳ error: " + error.error())); } default -> StreamingSession.super.emit(event); } diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java new file mode 100644 index 000000000..03501356f --- /dev/null +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java @@ -0,0 +1,146 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import java.io.IOException; +import java.util.List; +import java.util.Locale; +import org.jline.reader.EndOfFileException; +import org.jline.reader.LineReader; +import org.jline.reader.LineReaderBuilder; +import org.jline.reader.UserInterruptException; +import org.jline.reader.impl.completer.StringsCompleter; +import org.jline.terminal.Attributes; +import org.jline.terminal.Terminal; +import org.jline.terminal.TerminalBuilder; +import org.jline.utils.AttributedString; +import org.jline.utils.AttributedStyle; +import org.jline.utils.Status; +import org.jspecify.annotations.Nullable; + +/** + * An {@link AgentTerminal} on a real terminal, via JLine: line editing and history at the prompt, tab + * completion of the commands, a status line pinned to the bottom of the window, and single-key + * answers. + * + *

Streamed output goes through {@link LineReader#printAbove(String)}, which scrolls it in above + * the prompt while the bottom block stays where it is — the one thing a plain {@code println} cannot + * do. Nothing is ever redrawn above that block, so the scrollback stays exactly as it was written. + * + *

{@link #open} returns {@code null} instead of throwing when there is no usable terminal (piped + * input, a "dumb" terminal, a missing native provider); the caller then uses {@link PlainTerminal}. + * Ctrl-C at the prompt clears the line and returns an empty one — it does not end the session; Ctrl-D + * ends input like end-of-file. + */ +public final class JLineTerminal implements AgentTerminal { + + private final Terminal terminal; + private final LineReader reader; + private final Status status; + private final Ansi ansi; + + private JLineTerminal(Terminal terminal, LineReader reader, Status status, Ansi ansi) { + this.terminal = terminal; + this.reader = reader; + this.status = status; + this.ansi = ansi; + } + + /** + * Open the system terminal. + * + * @param completions the words tab completes, e.g. the command names + * @return the terminal, or {@code null} when this is not an interactive terminal + */ + public static @Nullable JLineTerminal open(List completions) { + try { + Terminal terminal = TerminalBuilder.builder().system(true).build(); + if (terminal.getType().startsWith(Terminal.TYPE_DUMB)) { + terminal.close(); + return null; + } + LineReader reader = LineReaderBuilder.builder() + .terminal(terminal) + .completer(new StringsCompleter(completions)) + .build(); + return new JLineTerminal(terminal, reader, Status.getStatus(terminal), Ansi.detect()); + } catch (IOException | RuntimeException e) { + // No terminal, no native provider, a restricted environment: the plain console still works. + return null; + } + } + + @Override + public void line(String text) { + reader.printAbove(text); + } + + @Override + public @Nullable String readLine(String prompt) { + try { + return reader.readLine(prompt); + } catch (UserInterruptException e) { + return ""; // Ctrl-C: drop the line, ask again + } catch (EndOfFileException e) { + return null; // Ctrl-D + } + } + + @Override + public @Nullable String readKey(String prompt) { + terminal.writer().print(prompt); + terminal.writer().flush(); + Attributes saved = terminal.enterRawMode(); + try { + int key = terminal.reader().read(); + if (key < 0 || key == 4) { // end of input, Ctrl-D + return null; + } + if (key == 3) { // Ctrl-C: treat as "no", the safe answer + terminal.writer().print("^C" + System.lineSeparator()); + terminal.writer().flush(); + return "n"; + } + String answer = String.valueOf((char) key).toLowerCase(Locale.ROOT); + terminal.writer().print(answer + System.lineSeparator()); + terminal.writer().flush(); + return answer; + } catch (IOException e) { + return null; + } finally { + terminal.setAttributes(saved); + } + } + + @Override + public void status(String text) { + if (text.isEmpty()) { + status.update(List.of()); + return; + } + // A rule above the status line separates the live block from the scrollback, the way the + // established terminal agents frame their input. + int width = Math.max(10, terminal.getSize().getColumns()); + status.update(List.of( + new AttributedString("─".repeat(width), AttributedStyle.DEFAULT.foreground(AttributedStyle.BRIGHT)), + new AttributedString(text, AttributedStyle.DEFAULT.foreground(AttributedStyle.BRIGHT)))); + } + + @Override + public Ansi ansi() { + return ansi; + } + + @Override + public void close() { + try { + status.update(List.of()); + status.close(); + terminal.close(); + } catch (IOException | RuntimeException e) { + // closing a terminal that is already gone must not fail the session + } + } +} diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java index 8e2e03477..e7856045f 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java @@ -70,6 +70,12 @@ public final class LocalAgent { "You summarize a conversation between a user and a coding assistant. Follow the user's" + " instructions exactly and answer with the summary only."; + /** How often the activity line is refreshed while a turn runs. */ + private static final Duration ACTIVITY_INTERVAL = Duration.ofMillis(250); + + /** The spinner shown in the activity line. */ + private static final String ACTIVITY_FRAMES = "⠋⠙⠹⠸⠼⠴⠦⠧⠇⠏"; + private static final Duration SHELL_TIMEOUT = Duration.ofSeconds(120); private static final int SHELL_MAX_OUTPUT_CHARS = 20_000; @@ -116,6 +122,7 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream throws Exception { LlamaModel model = null; OpenAiCompatServer server = null; + AgentTerminal terminal = null; String baseUrl = options.getBaseUrl(); try { if (options.getModelPath() != null) { @@ -159,18 +166,19 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream new AtomicReference<>(options.isAuto() ? ApprovalMode.AUTO : ApprovalMode.MANUAL); boolean interactive = options.getPrompt() == null && input != null; BufferedReader reader = input == null ? null : new BufferedReader(input); + terminal = interactive ? JLineTerminal.open(commandNames()) : null; + if (terminal == null) { + terminal = new PlainTerminal(out, reader, Ansi.detect()); + } // One-shot runs have nobody at the keyboard, so the strategy gets no console and denies // gated calls unless --auto was passed (see ConsoleApprovalStrategy). - Ansi ansi = Ansi.detect(); - runner.approval( - new ConsoleApprovalStrategy(mode, interactive ? reader : null, out, ansi), - ConsoleApprovalStrategy.policy()); + runner.approval(new ConsoleApprovalStrategy(mode, terminal, interactive), ConsoleApprovalStrategy.policy()); int contextSize = options.getModelPath() != null ? options.getCtxSize() : ServerProps.contextSize(baseUrl, options.getApiKey()); if (options.getPrompt() != null) { - return turn(runner, fileSystem, options.getPrompt(), history, out, ansi) + return turn(runner, fileSystem, options.getPrompt(), history, terminal) .failure() == null ? 0 @@ -184,11 +192,9 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream long inputTokens = 0; boolean estimated = false; while (true) { - out.println(ansi.dim(StatusLine.render( - mode.get(), inputTokens, estimated, contextSize, tools.size(), options.getModelId()))); - out.print("you> "); - out.flush(); - String line = reader.readLine(); + terminal.status(StatusLine.render( + mode.get(), inputTokens, estimated, contextSize, tools.size(), options.getModelId())); + String line = terminal.readLine("you> "); if (line == null) { return 0; } @@ -210,15 +216,17 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream contextSize, inputTokens, estimated, - out, - ansi); + terminal); continue; } - ConsoleSession completed = turn(runner, fileSystem, line, history, out, ansi); + ConsoleSession completed = turn(runner, fileSystem, line, history, terminal); estimated = completed.inputTokens() == 0; inputTokens = estimated ? estimateTokens(systemPrompt(options), history) : completed.inputTokens(); } } finally { + if (terminal != null) { + terminal.close(); + } if (server != null) { server.close(); } @@ -233,12 +241,11 @@ private static ConsoleSession turn( AgentFileSystem fileSystem, String message, List history, - PrintStream out, - Ansi ansi) + AgentTerminal terminal) throws InterruptedException { - ConsoleSession session = new ConsoleSession(out, fileSystem, ansi); + ConsoleSession session = new ConsoleSession(terminal, fileSystem); runner.run(message, history, session); - boolean finished = session.await(TURN_TIMEOUT); + boolean finished = awaitWithActivity(session, terminal); history.add(ChatMessage.user(message)); if (!session.text().isEmpty()) { history.add(ChatMessage.assistant(session.text())); @@ -261,8 +268,7 @@ private static ConsoleSession turn( * @param contextSize the context window in tokens, or {@link StatusLine#UNKNOWN_CONTEXT} * @param inputTokens the input tokens of the last turn * @param estimated whether that number is an estimate - * @param out the console - * @param ansi the console styles + * @param terminal the console * @return the input tokens to show from now on (unchanged, or the summary's after {@code /compact}) * @throws InterruptedException if interrupted while a summary is generated */ @@ -276,18 +282,17 @@ private static long handleCommand( int contextSize, long inputTokens, boolean estimated, - PrintStream out, - Ansi ansi) + AgentTerminal terminal) throws InterruptedException { switch (command.command()) { - case HELP -> out.println(prompt(HELP_TEXT)); + case HELP -> prompt(HELP_TEXT).lines().forEach(terminal::line); case CLEAR -> { history.clear(); - out.println("(history cleared)"); + terminal.line("(history cleared)"); } case TOOLS -> { - out.println("tools: " + String.join(", ", runner.toolNames())); - out.println("asks before running (manual mode): " + terminal.line("tools: " + String.join(", ", runner.toolNames())); + terminal.line("asks before running (manual mode): " + String.join(", ", ConsoleApprovalStrategy.gated(runner.toolNames()))); } case MODE -> { @@ -295,25 +300,25 @@ private static long handleCommand( try { mode.set(ApprovalMode.parse(command.arguments())); } catch (IllegalArgumentException e) { - out.println(e.getMessage()); + terminal.line(e.getMessage()); return inputTokens; } } - out.println("approval mode: " + mode.get().label()); + terminal.line("approval mode: " + mode.get().label()); } case STATUS -> { - out.println(StatusLine.render( + terminal.line(StatusLine.render( mode.get(), inputTokens, estimated, contextSize, runner.toolNames().size(), options.getModelId())); - out.println("workspace: " + options.getWorkspace()); - out.println("history: " + history.size() + " messages"); + terminal.line("workspace: " + options.getWorkspace()); + terminal.line("history: " + history.size() + " messages"); } case COMPACT -> { - return compact(runner, fileSystem, history, command.arguments(), out, ansi); + return compact(runner, fileSystem, history, command.arguments(), terminal); } case EXIT -> { // handled by the caller, which has to return from the loop @@ -322,6 +327,34 @@ private static long handleCommand( return inputTokens; } + /** + * Wait for a turn while the status line shows that something is happening. + * + *

A local model can think for a while before the first token arrives, and a silent console is + * indistinguishable from a hung one. The line is rewritten in place (it is the pinned status area, + * not the scrollback), so nothing the user has already read moves. + * + * @param session the running turn + * @param terminal the console + * @return {@code true} when the turn finished within {@link #TURN_TIMEOUT} + * @throws InterruptedException if interrupted while waiting + */ + private static boolean awaitWithActivity(ConsoleSession session, AgentTerminal terminal) + throws InterruptedException { + long start = System.nanoTime(); + int frame = 0; + while (!session.await(ACTIVITY_INTERVAL)) { + long seconds = (System.nanoTime() - start) / 1_000_000_000L; + if (seconds > TURN_TIMEOUT.toSeconds()) { + return false; + } + terminal.status(ACTIVITY_FRAMES.charAt(frame++ % ACTIVITY_FRAMES.length()) + " working… (" + seconds + + "s · " + session.toolCalls() + " tool calls)"); + } + terminal.status(""); + return true; + } + /** * A rough token count of what the next request will carry. * @@ -355,8 +388,7 @@ static long estimateTokens(String systemPrompt, List history) { * @param fileSystem the workspace filesystem for the summary session * @param history the history, replaced in place * @param focus optional extra instructions from {@code /compact } - * @param out the console - * @param ansi the console styles + * @param terminal the console * @return the input tokens the summarizing call reported * @throws InterruptedException if interrupted while the summary is generated */ @@ -365,28 +397,27 @@ private static long compact( AgentFileSystem fileSystem, List history, String focus, - PrintStream out, - Ansi ansi) + AgentTerminal terminal) throws InterruptedException { if (history.isEmpty()) { - out.println("(nothing to compact)"); + terminal.line("(nothing to compact)"); return 0; } String instructions = prompt(COMPACT_PROMPT) .replace("{focus}", focus.isEmpty() ? "" : System.lineSeparator() + "Focus on: " + focus); int before = history.size(); - out.println("(compacting " + before + " messages …)"); - ConsoleSession session = new ConsoleSession(out, fileSystem, ansi); + terminal.line("(compacting " + before + " messages …)"); + ConsoleSession session = new ConsoleSession(terminal, fileSystem); runner.runWithoutTools(instructions, List.copyOf(history), session, COMPACT_SYSTEM_PROMPT); - if (!session.await(TURN_TIMEOUT) || session.text().isBlank()) { - out.println("(compact failed; history kept)"); + if (!awaitWithActivity(session, terminal) || session.text().isBlank()) { + terminal.line("(compact failed; history kept)"); return 0; } history.clear(); history.add(ChatMessage.user("Summary of the conversation so far:" + System.lineSeparator() + session.text().strip())); history.add(ChatMessage.assistant("Understood, I will continue from that summary.")); - out.println("(compacted " + before + " messages into a summary of " + terminal.line("(compacted " + before + " messages into a summary of " + session.text().strip().length() + " characters)"); return session.inputTokens(); } @@ -421,6 +452,17 @@ static ModelParameters modelParameters(AgentOptions options) { return parameters; } + /** + * Every command name and alias, for tab completion. + * + * @return the names, each with its leading slash + */ + static List commandNames() { + return java.util.Arrays.stream(SlashCommands.Command.values()) + .flatMap(command -> command.names().stream()) + .toList(); + } + /** * The default system prompt, or the {@code --system} override. * diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/MarkdownConsole.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/MarkdownConsole.java index ea27e904c..6e15f06fc 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/MarkdownConsole.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/MarkdownConsole.java @@ -4,7 +4,7 @@ package net.ladenthin.llama.atmosphere; -import java.io.PrintStream; +import java.util.function.Consumer; /** * Renders the streamed answer as it arrives: just enough Markdown to make it readable. @@ -25,7 +25,7 @@ */ public final class MarkdownConsole { - private final PrintStream out; + private final Consumer sink; private final Ansi ansi; private final StringBuilder pending = new StringBuilder(); private boolean inFence; @@ -33,11 +33,11 @@ public final class MarkdownConsole { /** * Create a renderer. * - * @param out where the rendered text goes + * @param sink receives each rendered line, without its terminator * @param ansi the styles (a plain instance writes the text unchanged) */ - public MarkdownConsole(PrintStream out, Ansi ansi) { - this.out = out; + public MarkdownConsole(Consumer sink, Ansi ansi) { + this.sink = sink; this.ansi = ansi; } @@ -52,19 +52,17 @@ public void append(String chunk) { while ((newline = pending.indexOf("\n")) >= 0) { String line = pending.substring(0, newline); pending.delete(0, newline + 1); - out.println(render(line.endsWith("\r") ? line.substring(0, line.length() - 1) : line)); + sink.accept(render(line.endsWith("\r") ? line.substring(0, line.length() - 1) : line)); } - out.flush(); } /** Write what is left of an unfinished line, e.g. an answer that does not end with a newline. */ public void flush() { if (pending.length() > 0) { - out.println(render(pending.toString())); + sink.accept(render(pending.toString())); pending.setLength(0); } inFence = false; - out.flush(); } /** diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/PlainTerminal.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/PlainTerminal.java new file mode 100644 index 000000000..f5d6deb06 --- /dev/null +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/PlainTerminal.java @@ -0,0 +1,87 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import java.io.BufferedReader; +import java.io.IOException; +import java.io.PrintStream; +import java.util.Locale; +import org.jspecify.annotations.Nullable; + +/** + * An {@link AgentTerminal} over a plain stream pair: what a one-shot run, a piped session and the + * tests get. + * + *

No cursor control at all — the status line is printed as an ordinary line before the prompt and + * scrolls away like everything else, and an answer needs Enter. That is the point: this + * implementation stays correct when the output is a file. + */ +public final class PlainTerminal implements AgentTerminal { + + private final PrintStream out; + private final @Nullable BufferedReader in; + private final Ansi ansi; + private String status = ""; + + /** + * Create a plain console. + * + * @param out where output goes + * @param in where input is read, or {@code null} when there is none (one-shot runs) + * @param ansi the styles + */ + public PlainTerminal(PrintStream out, @Nullable BufferedReader in, Ansi ansi) { + this.out = out; + this.in = in; + this.ansi = ansi; + } + + @Override + public void line(String text) { + out.println(text); + out.flush(); + } + + @Override + public @Nullable String readLine(String prompt) { + if (!status.isEmpty()) { + out.println(ansi.dim(status)); + } + out.print(prompt); + out.flush(); + return read(); + } + + @Override + public @Nullable String readKey(String prompt) { + out.print(prompt); + out.flush(); + String answer = read(); + return answer == null ? null : answer.trim().toLowerCase(Locale.ROOT); + } + + @Override + public void status(String text) { + this.status = text; + } + + @Override + public Ansi ansi() { + return ansi; + } + + @Override + public void close() { + out.flush(); + } + + private @Nullable String read() { + try { + return in == null ? null : in.readLine(); + } catch (IOException e) { + return null; + } + } +} diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ApprovalWireTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ApprovalWireTest.java index cc62cc398..b230fd21b 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ApprovalWireTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ApprovalWireTest.java @@ -52,9 +52,15 @@ class ApprovalWireTest { private final ByteArrayOutputStream console = new ByteArrayOutputStream(); + private AgentTerminal terminal(String typed) { + return new PlainTerminal( + new PrintStream(console, true, StandardCharsets.UTF_8), + new BufferedReader(new StringReader(typed)), + Ansi.PLAIN); + } + private AgentRunner runner(OpenAiCompatServer server, List tools, String typed) { AtomicReference mode = new AtomicReference<>(ApprovalMode.MANUAL); - PrintStream out = new PrintStream(console, true, StandardCharsets.UTF_8); return new AgentRunner( "http://127.0.0.1:" + server.getPort() + "/v1", API_KEY, @@ -65,14 +71,13 @@ private AgentRunner runner(OpenAiCompatServer server, List tools 64, 10) .retryPolicy(RetryPolicy.NONE) - .approval( - new ConsoleApprovalStrategy(mode, new BufferedReader(new StringReader(typed)), out, Ansi.PLAIN), - ConsoleApprovalStrategy.policy()); + .approval(new ConsoleApprovalStrategy(mode, terminal(typed), true), ConsoleApprovalStrategy.policy()); } private ConsoleSession session() { AgentFileSystem fs = new WorkspaceAgentFileSystem(workspace, AgentFileSystem.Limits.defaults()); - return new ConsoleSession(new PrintStream(console, true, StandardCharsets.UTF_8), fs); + return new ConsoleSession( + new PlainTerminal(new PrintStream(console, true, StandardCharsets.UTF_8), null, Ansi.PLAIN), fs); } private static OpenAiServerConfig config() { diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/AtmosphereToolLoopIntegrationTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/AtmosphereToolLoopIntegrationTest.java index 198f3bcde..b4f423a1a 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/AtmosphereToolLoopIntegrationTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/AtmosphereToolLoopIntegrationTest.java @@ -113,7 +113,10 @@ private AgentRunner runner(List tools) { private ConsoleSession session() { AgentFileSystem fs = new WorkspaceAgentFileSystem(workspace, AgentFileSystem.Limits.defaults()); - return new ConsoleSession(new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8), fs); + return new ConsoleSession( + new PlainTerminal( + new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8), null, Ansi.PLAIN), + fs); } /** diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/AtmosphereWireContractTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/AtmosphereWireContractTest.java index 02b8153cf..cfb14b82e 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/AtmosphereWireContractTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/AtmosphereWireContractTest.java @@ -74,7 +74,10 @@ private AgentRunner runner(OpenAiCompatServer server, String apiKey, List invocations, String result) { diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleApprovalStrategyTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleApprovalStrategyTest.java index abe7d4629..c223b9ee5 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleApprovalStrategyTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleApprovalStrategyTest.java @@ -39,9 +39,10 @@ private PendingApproval approval() { private ApprovalOutcome ask(AtomicReference mode, String typed) { BufferedReader reader = typed == null ? null : new BufferedReader(new StringReader(typed)); - PrintStream out = new PrintStream(console, true, StandardCharsets.UTF_8); + AgentTerminal terminal = + new PlainTerminal(new PrintStream(console, true, StandardCharsets.UTF_8), reader, Ansi.PLAIN); // The strategy never touches the session; Atmosphere passes it only so a UI can emit events. - return new ConsoleApprovalStrategy(mode, reader, out, Ansi.PLAIN).awaitApproval(approval(), null); + return new ConsoleApprovalStrategy(mode, terminal, typed != null).awaitApproval(approval(), null); } private String consoleText() { diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleFormattingTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleFormattingTest.java index 494ad180d..c406eca3a 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleFormattingTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleFormattingTest.java @@ -9,9 +9,6 @@ import static org.hamcrest.Matchers.is; import static org.hamcrest.Matchers.not; -import java.io.ByteArrayOutputStream; -import java.io.PrintStream; -import java.nio.charset.StandardCharsets; import java.util.Map; import org.junit.jupiter.api.Test; @@ -80,14 +77,15 @@ void anEstimatedCountIsMarkedWithATilde() { // ----- MarkdownConsole ----- private static String render(String text, Ansi ansi) { - ByteArrayOutputStream buffer = new ByteArrayOutputStream(); - MarkdownConsole console = new MarkdownConsole(new PrintStream(buffer, true, StandardCharsets.UTF_8), ansi); + StringBuilder buffer = new StringBuilder(); + MarkdownConsole console = + new MarkdownConsole(line -> buffer.append(line).append(System.lineSeparator()), ansi); // one character at a time: the renderer must not depend on where the stream splits for (int i = 0; i < text.length(); i++) { console.append(text.substring(i, i + 1)); } console.flush(); - return buffer.toString(StandardCharsets.UTF_8); + return buffer.toString(); } @Test diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/PlainTerminalTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/PlainTerminalTest.java new file mode 100644 index 000000000..b7643d0a1 --- /dev/null +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/PlainTerminalTest.java @@ -0,0 +1,87 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import static org.hamcrest.MatcherAssert.assertThat; +import static org.hamcrest.Matchers.containsString; +import static org.hamcrest.Matchers.hasItem; +import static org.hamcrest.Matchers.is; +import static org.hamcrest.Matchers.nullValue; + +import java.io.BufferedReader; +import java.io.ByteArrayOutputStream; +import java.io.PrintStream; +import java.io.StringReader; +import java.nio.charset.StandardCharsets; +import java.util.List; +import org.junit.jupiter.api.Test; + +/** The console used whenever there is no terminal: one-shot runs, piped input, and these tests. */ +class PlainTerminalTest { + + private final ByteArrayOutputStream out = new ByteArrayOutputStream(); + + private PlainTerminal terminal(String typed) { + return new PlainTerminal( + new PrintStream(out, true, StandardCharsets.UTF_8), + typed == null ? null : new BufferedReader(new StringReader(typed)), + Ansi.PLAIN); + } + + private String written() { + return out.toString(StandardCharsets.UTF_8); + } + + @Test + void linesAreWrittenAndInputIsReadBackLineByLine() { + PlainTerminal terminal = terminal("first" + System.lineSeparator() + "second" + System.lineSeparator()); + + terminal.line("hello"); + + assertThat(written(), containsString("hello")); + assertThat(terminal.readLine("you> "), is("first")); + assertThat(written(), containsString("you> ")); + assertThat(terminal.readKey("allow? "), is("second")); + assertThat(terminal.readLine("you> "), is(nullValue())); + } + + @Test + void theStatusLineIsPrintedBeforeThePromptBecauseNothingCanBePinned() { + PlainTerminal terminal = terminal("x" + System.lineSeparator()); + terminal.status("[manual · ctx 0/16k]"); + + terminal.readLine("you> "); + + assertThat(written(), containsString("[manual · ctx 0/16k]")); + assertThat( + written().indexOf("[manual"), + is(org.hamcrest.Matchers.lessThan(written().indexOf("you> ")))); + } + + @Test + void anAnswerIsNormalisedSoTheCallerCanCompareItToOneCase() { + assertThat(terminal("YES" + System.lineSeparator()).readKey("? "), is("yes")); + assertThat(terminal(" A " + System.lineSeparator()).readKey("? "), is("a")); + } + + @Test + void withoutInputEverythingReadsAsEndOfInput() { + PlainTerminal terminal = terminal(null); + + assertThat(terminal.readLine("you> "), is(nullValue())); + assertThat(terminal.readKey("? "), is(nullValue())); + } + + @Test + void everyCommandNameIsOfferedForCompletion() { + List names = LocalAgent.commandNames(); + + for (SlashCommands.Command command : SlashCommands.Command.values()) { + for (String name : command.names()) { + assertThat(names, hasItem(name)); + } + } + } +} From 8751cb1e60e427553d6d816919df057cda7879d5 Mon Sep 17 00:00:00 2001 From: Bernard Ladenthin Date: Wed, 23 Sep 2026 00:44:57 +0200 Subject: [PATCH 04/36] =?UTF-8?q?llama-atmosphere-agent:=20/loop=20?= =?UTF-8?q?=E2=80=94=20keep=20working=20on=20one=20task,=20state=20in=20a?= =?UTF-8?q?=20file?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit /loop [--every 5m] [--max 20] [--check ''] works on one task step by step without the user typing between steps. - Every step sends the same message: the task verbatim plus "read AGENT-LOOP.md, do one concrete step, write down what happened". The history is dropped between steps, so the file is the memory and the context cannot grow (the shape Claude Code's ralph-wiggum plugin uses). - The loop ends when a line of the answer is exactly <>. A text marker, not a "done" tool: below 7B a model emits a malformed tool call far more often than a malformed line, and mini-SWE-agent's SWE-bench results come from a plain sentinel with no tool-call API at all. A substring never counts, so the model talking about the marker cannot end the run. - --check '' only believes the marker when that command succeeds, and feeds its output back otherwise -- the cheapest defence against a 4B declaring victory. - Guards, all enforced here: step cap (20), two-hour budget, stall detection (three steps with no file change and no tool call), optional interval. A loop asks once to switch to the auto approval mode and gives up if the answer is no. The marker is checked BEFORE the stall detector: the step that only answers "done" changes no file and calls no tool, and the other order reported it as "no progress" -- caught by theLoopEndsWhenTheModelAnswersWithTheMarker while writing the tests. Also: the workspace is now part of the pinned status line, shortened to its last two segments when the path is long. Tests: 81 green (was 67). TaskLoopTest drives the loop over the real OpenAiCompatServer with scripted answers: marker, empty history per step, stall, step limit, failing check. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01E2h8gXyyE5UeimkL9vQv9G --- CLAUDE.md | 13 +- llama-atmosphere-agent/README.md | 31 +++ .../llama/atmosphere/LocalAgent.java | 77 +++++- .../llama/atmosphere/LoopOptions.java | 147 ++++++++++ .../llama/atmosphere/SlashCommands.java | 2 + .../llama/atmosphere/StatusLine.java | 39 ++- .../ladenthin/llama/atmosphere/TaskLoop.java | 250 ++++++++++++++++++ .../net/ladenthin/llama/atmosphere/help.txt | 3 + .../llama/atmosphere/loop-file-template.md | 17 ++ .../llama/atmosphere/loop-prompt.txt | 12 + .../atmosphere/ConsoleFormattingTest.java | 19 +- .../llama/atmosphere/TaskLoopTest.java | 245 +++++++++++++++++ 12 files changed, 844 insertions(+), 11 deletions(-) create mode 100644 llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LoopOptions.java create mode 100644 llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/TaskLoop.java create mode 100644 llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/loop-file-template.md create mode 100644 llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/loop-prompt.txt create mode 100644 llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/TaskLoopTest.java diff --git a/CLAUDE.md b/CLAUDE.md index 089474d42..1dcf65dc0 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -2266,7 +2266,18 @@ are decisions, not details: emphasis). Colour is decided once in `Ansi.detect()` — `CLICOLOR_FORCE`, then `NO_COLOR`, then `TERM=dumb`/`CLICOLOR=0`, else "is a terminal" via `Console.isTerminal()` (reflective: JDK 22+; below that `System.console() != null`). -5. **The context number in the status line is an estimate, marked `~`.** llama.cpp emits its usage +5. **`/loop` keeps its state in a file, not in the context, and stops on a text marker** (`TaskLoop`, + `LoopOptions`). Every step re-sends the task verbatim and drops the history, so the context cannot + grow (Claude Code's ralph-wiggum plugin does the same, and `AGENT-LOOP.md` in the workspace is the + memory). The stop signal is a line that is **exactly** `<>`, never a substring — + deliberately **not** a "done" tool: below 7B a model emits a malformed tool call far more often + than a malformed line, and mini-SWE-agent's SWE-bench results come from exactly this plain-sentinel + design. `--check ''` re-verifies the claim and feeds a failure back. Four guards, none of them + trusted to the model: step cap, wall-clock budget, stall detection (three steps with no file change + and no tool call), interval. **Order matters and a test pins it**: the marker is checked *before* + the stall detector, because the step that only answers "done" changes nothing and would otherwise + be reported as no progress. +6. **The context number in the status line is an estimate, marked `~`.** llama.cpp emits its usage chunk only when the client sets `stream_options.include_usage`, and Atmosphere's client does not; `ConsoleSession.usage()` takes the real count when one arrives, otherwise `LocalAgent.estimateTokens` uses four characters per token. The window size is `--ctx-size` (in-process) or the server's diff --git a/llama-atmosphere-agent/README.md b/llama-atmosphere-agent/README.md index 6e9e7870a..9929cc408 100644 --- a/llama-atmosphere-agent/README.md +++ b/llama-atmosphere-agent/README.md @@ -185,6 +185,7 @@ unknown `/command` included — goes to the model: | `/tools` | the tools offered, and which of them ask first | | `/mode [manual\|auto]` (`/approve`) | show or set the approval mode | | `/compact [focus]` | summarize the conversation and continue from the summary | +| `/loop [--every 5m] [--max 20] [--check ''] ` | keep working on one task until it is done | | `/clear` (`/reset`, `/new`) | drop the history | | `/exit` (`/quit`) | leave | @@ -217,6 +218,33 @@ next step; `/compact ` adds an emphasis), then replaces the history with when the context fills up. Note the history only ever held the user texts and the final answers — tool rounds are not replayed across turns — so nothing else is lost. +**`/loop`** keeps working on one task without you typing anything between steps: + +``` +/loop --check 'mvn -q test' make ShellToolTest pass on Windows +``` + +Every step sends the **same** message — the task verbatim plus "read `AGENT-LOOP.md`, do one concrete +step, write down what happened". The conversation history is **dropped between steps**: the file in +the workspace is the memory, so the context never grows and the loop can run for a long time. The +loop ends when a line of the answer is exactly + +``` +<> +``` + +A marker *mentioned* inside a sentence does not count, only a line of its own. This is a text marker +rather than a "done" tool on purpose: small local models produce a well-formed tool call far less +reliably than a line of text — mini-SWE-agent reaches its SWE-bench results with a plain sentinel and +no tool-call API at all, and Claude Code's own ralph-wiggum plugin matches an exact string too. + +With `--check ''` the marker is only believed when that command succeeds; otherwise its +output goes into the next step. That is the cheapest defence against a small model declaring victory +after one edit. Four limits stop a runaway loop, all enforced by the agent, none of them trusted to +the model: `--max` steps (20 by default), a two-hour wall-clock budget, a stall detector (three steps +in a row that write nothing and call no tool), and `--every ` for a paced run. A loop needs +the `auto` approval mode — it asks once and switches, or leaves you alone if you say no. + **The status line** above the prompt reads `[manual · ctx ~3.1k/16k · 9 tools · local-model]`: the approval mode, the context used out of the window, the number of tools and the model id. A `~` means the number is an estimate from the text @@ -255,6 +283,7 @@ Then, in this order: | `explain Markdown with a heading, a list, bold text and a code block` | the answer arrives rendered: heading bold, `•` bullets, code in colour | | press ↑ | the previous line comes back; Tab after `/` completes the commands | | `/compact` | the conversation is summarized and replaces the history; `ctx` drops | +| `/loop --max 3 add a line with the current date to notes.txt, then stop` | three steps at most, with `AGENT-LOOP.md` appearing in the workspace | | `/exit` | leave | `--auto` starts in auto mode, `--verbose` brings llama.cpp's own log back, and `NO_COLOR=1` turns @@ -353,6 +382,8 @@ starter are the *deployment* layer on top of the same runtime — not needed for Atmosphere's `ApprovalResolution` also supports approve-with-edited-arguments, which the console does not offer. - No auto-compaction when the context fills up; `/compact` is manual. +- `/loop` cannot be interrupted in the middle of a step — Ctrl-C ends the process; the loop file + survives, so restarting the same `/loop` continues where it left off. - An engine error after the stream started ends the turn silently (see the table). - **One in-process agent per machine at a time.** The core extracts its native library to a fixed name (`jllama.dll` / `libjllama.so` in the temp directory); on Windows a second JVM cannot replace diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java index e7856045f..c444e45a6 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java @@ -62,6 +62,12 @@ public final class LocalAgent { /** The {@code /help} overview. */ static final String HELP_TEXT = "help.txt"; + /** The per-step message of {@code /loop}; placeholders {@code {task}}, {@code {file}}, {@code {check_hint}}. */ + static final String LOOP_PROMPT = "loop-prompt.txt"; + + /** The skeleton written to {@link TaskLoop#LOOP_FILE}; placeholder {@code {task}}. */ + static final String LOOP_FILE_TEMPLATE = "loop-file-template.md"; + /** The instructions {@code /compact} sends; placeholder {@code {focus}}. */ static final String COMPACT_PROMPT = "compact-prompt.txt"; @@ -193,7 +199,13 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream boolean estimated = false; while (true) { terminal.status(StatusLine.render( - mode.get(), inputTokens, estimated, contextSize, tools.size(), options.getModelId())); + options.getWorkspace(), + mode.get(), + inputTokens, + estimated, + contextSize, + tools.size(), + options.getModelId())); String line = terminal.readLine("you> "); if (line == null) { return 0; @@ -236,7 +248,66 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream } } - private static ConsoleSession turn( + /** + * Run {@code /loop}: work on one task step by step until it is done. + * + *

Two things are settled before the first step. The loop needs the {@link ApprovalMode#AUTO} + * mode — a run that asks before every write is not a loop, it is a conversation — so a manual + * session is asked once and left alone if the answer is no. And the history is not the + * loop's memory: {@link TaskLoop} drops it every step and keeps the state in a file, so nothing of + * the current conversation is used or changed here. + * + * @param runner the runner + * @param fileSystem the workspace filesystem + * @param terminal the console + * @param options the agent options, for the workspace + * @param mode the approval mode, possibly switched to auto here + * @param arguments everything after {@code /loop} + * @throws InterruptedException if interrupted while a step runs + */ + private static void loop( + AgentRunner runner, + AgentFileSystem fileSystem, + AgentTerminal terminal, + AgentOptions options, + AtomicReference mode, + String arguments) + throws InterruptedException { + LoopOptions loopOptions; + try { + loopOptions = LoopOptions.parse(arguments); + } catch (IllegalArgumentException e) { + terminal.line(e.getMessage()); + return; + } + if (mode.get() != ApprovalMode.AUTO) { + String answer = + terminal.readKey("a loop cannot stop at every question — switch to auto for it? [y]es / [n]o: "); + if (answer == null || !(answer.startsWith("y") || answer.isEmpty())) { + terminal.line("loop: cancelled (use /mode auto to allow it)"); + return; + } + mode.set(ApprovalMode.AUTO); + } + if (!TaskLoop.canKeepNotes(runner.toolNames())) { + terminal.line("loop: needs the read_file and write_file tools to keep its notes"); + return; + } + TaskLoop.Outcome outcome = TaskLoop.run( + runner, + fileSystem, + terminal, + options.getWorkspace(), + loopOptions, + () -> false, + TaskLoop.DEFAULT_BUDGET); + terminal.line( + outcome.completed() + ? terminal.ansi().green("loop: " + outcome.reason()) + : terminal.ansi().yellow("loop: " + outcome.reason())); + } + + static ConsoleSession turn( AgentRunner runner, AgentFileSystem fileSystem, String message, @@ -308,6 +379,7 @@ private static long handleCommand( } case STATUS -> { terminal.line(StatusLine.render( + options.getWorkspace(), mode.get(), inputTokens, estimated, @@ -320,6 +392,7 @@ private static long handleCommand( case COMPACT -> { return compact(runner, fileSystem, history, command.arguments(), terminal); } + case LOOP -> loop(runner, fileSystem, terminal, options, mode, command.arguments()); case EXIT -> { // handled by the caller, which has to return from the loop } diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LoopOptions.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LoopOptions.java new file mode 100644 index 000000000..e9210ad09 --- /dev/null +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LoopOptions.java @@ -0,0 +1,147 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import java.time.Duration; +import java.util.Locale; +import org.jspecify.annotations.Nullable; + +/** + * What {@code /loop} was asked to do: the task, and the limits that stop it. + * + *

Syntax: {@code /loop [--every ] [--max ] [--check ] }. The flags come + * first, the rest of the line is the task verbatim — so a task may contain anything, including words + * that look like flags, as long as they are not at the front. + * + * @param task what the agent should work on, restated unchanged every step + * @param interval the pause between steps, or {@code null} to run them back to back + * @param maxSteps the hard cap on steps + * @param check a command that must succeed before {@code <>} is accepted, or + * {@code null} to trust the model + */ +public record LoopOptions( + String task, + @Nullable Duration interval, + int maxSteps, + @Nullable String check) { + + /** Steps before the loop gives up on its own. */ + public static final int DEFAULT_MAX_STEPS = 20; + + /** + * Parse the argument of {@code /loop}. + * + * @param arguments everything after the command name + * @return the parsed options + * @throws IllegalArgumentException when a flag has no value, a duration or number is malformed, or + * no task is left + */ + public static LoopOptions parse(String arguments) { + String rest = arguments == null ? "" : arguments.trim(); + Duration interval = null; + int maxSteps = DEFAULT_MAX_STEPS; + String check = null; + while (rest.startsWith("--")) { + String[] flag = split(rest); + switch (flag[0]) { + case "--every" -> { + String[] value = split(flag[1]); + interval = parseDuration(require(value[0], "--every")); + rest = value[1]; + } + case "--max" -> { + String[] value = split(flag[1]); + maxSteps = parseSteps(require(value[0], "--max")); + rest = value[1]; + } + case "--check" -> { + String[] value = splitQuoted(flag[1]); + check = require(value[0], "--check"); + rest = value[1]; + } + default -> throw new IllegalArgumentException("Unknown /loop flag: " + flag[0]); + } + } + if (rest.isEmpty()) { + throw new IllegalArgumentException("Usage: /loop [--every 5m] [--max 20] [--check ''] "); + } + return new LoopOptions(rest, interval, maxSteps, check); + } + + /** + * A duration written the way a human writes it: {@code 30s}, {@code 5m}, {@code 2h}, or a bare + * number of minutes. + * + * @param text the duration + * @return the parsed duration + * @throws IllegalArgumentException when it is not a duration or is not positive + */ + static Duration parseDuration(String text) { + String value = text.trim().toLowerCase(Locale.ROOT); + char unit = value.charAt(value.length() - 1); + String number = Character.isDigit(unit) ? value : value.substring(0, value.length() - 1); + long amount; + try { + amount = Long.parseLong(number); + } catch (NumberFormatException e) { + throw new IllegalArgumentException("Not a duration: " + text + " (try 30s, 5m or 2h)", e); + } + Duration duration = + switch (Character.isDigit(unit) ? 'm' : unit) { + case 's' -> Duration.ofSeconds(amount); + case 'm' -> Duration.ofMinutes(amount); + case 'h' -> Duration.ofHours(amount); + default -> throw new IllegalArgumentException("Unknown time unit in " + text + " (use s, m or h)"); + }; + if (duration.isZero() || duration.isNegative()) { + throw new IllegalArgumentException("The interval must be positive: " + text); + } + return duration; + } + + private static int parseSteps(String text) { + try { + int steps = Integer.parseInt(text.trim()); + if (steps <= 0) { + throw new IllegalArgumentException("--max must be positive: " + text); + } + return steps; + } catch (NumberFormatException e) { + throw new IllegalArgumentException("--max expects a number, got: " + text, e); + } + } + + private static String require(String value, String flag) { + if (value.isEmpty()) { + throw new IllegalArgumentException("Missing value for " + flag); + } + return value; + } + + /** Split off the first whitespace-separated word: {@code [word, rest]}. */ + private static String[] split(String text) { + int space = text.indexOf(' '); + return space < 0 + ? new String[] {text.trim(), ""} + : new String[] { + text.substring(0, space).trim(), text.substring(space + 1).trim() + }; + } + + /** Like {@link #split}, but a leading {@code '...'} or {@code "..."} keeps its spaces. */ + private static String[] splitQuoted(String text) { + String trimmed = text.trim(); + if (trimmed.startsWith("'") || trimmed.startsWith("\"")) { + char quote = trimmed.charAt(0); + int end = trimmed.indexOf(quote, 1); + if (end > 0) { + return new String[] { + trimmed.substring(1, end), trimmed.substring(end + 1).trim() + }; + } + } + return split(trimmed); + } +} diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/SlashCommands.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/SlashCommands.java index 9f8974d17..185ea1fa6 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/SlashCommands.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/SlashCommands.java @@ -32,6 +32,8 @@ public enum Command { CLEAR("/clear", "/reset", "/new"), /** Summarize the history and continue with the summary; the argument steers the summary. */ COMPACT("/compact"), + /** Keep working on one task until it is done; see {@link TaskLoop}. */ + LOOP("/loop"), /** Show or set the approval mode; the argument is {@code manual} or {@code auto}. */ MODE("/mode", "/approve"), /** Print endpoint, model, tools, approval mode and context usage. */ diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/StatusLine.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/StatusLine.java index 1668491d9..1dff4fd9b 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/StatusLine.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/StatusLine.java @@ -7,8 +7,8 @@ import java.util.Locale; /** - * The one-line status shown above the {@code you>} prompt: approval mode, context usage, tool count - * and model id. + * The one-line status pinned below the answer: workspace, approval mode, context usage, tool count + * and model id — the four things whose answer changes what the next request does. * *

Context usage is the input side of the last completed turn — the prompt the server had to * process, which is what fills the context window; the generated tokens of that turn are already part @@ -26,6 +26,9 @@ */ public final class StatusLine { + /** Above this many characters the workspace path is shortened to its last two segments. */ + private static final int MAX_PATH_CHARS = 40; + /** The context size is unknown (no {@code /props}, and no in-process model). */ public static final int UNKNOWN_CONTEXT = 0; @@ -34,6 +37,7 @@ private StatusLine() {} /** * Render the status line. * + * @param workspace the directory the tools work in * @param mode the approval mode * @param inputTokens the input tokens of the last turn, or {@code 0} before the first one * @param estimated whether that number is an estimate rather than the server's own count @@ -43,9 +47,34 @@ private StatusLine() {} * @return one line, without a trailing newline */ public static String render( - ApprovalMode mode, long inputTokens, boolean estimated, int contextSize, int tools, String modelId) { - return "[" + mode.label() + " · " + context(inputTokens, estimated, contextSize) + " · " + tools + " tools · " - + modelId + "]"; + java.nio.file.Path workspace, + ApprovalMode mode, + long inputTokens, + boolean estimated, + int contextSize, + int tools, + String modelId) { + return "[" + shorten(workspace) + " · " + mode.label() + " · " + context(inputTokens, estimated, contextSize) + + " · " + tools + " tools · " + modelId + "]"; + } + + /** + * The workspace path, shortened from the left when it would take over the line. + * + *

The tools work relative to this directory and the shell starts in it, so it belongs on the + * line that is always visible — but a deep path would push everything else off the screen, so only + * the last two segments survive, marked with a leading ellipsis. + * + * @param workspace the workspace directory + * @return the path, or its tail + */ + static String shorten(java.nio.file.Path workspace) { + String full = workspace.toString(); + if (full.length() <= MAX_PATH_CHARS || workspace.getNameCount() < 2) { + return full; + } + return "…" + workspace.getFileSystem().getSeparator() + + workspace.subpath(workspace.getNameCount() - 2, workspace.getNameCount()); } /** diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/TaskLoop.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/TaskLoop.java new file mode 100644 index 000000000..0686f05b4 --- /dev/null +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/TaskLoop.java @@ -0,0 +1,250 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import java.io.IOException; +import java.io.UncheckedIOException; +import java.nio.charset.StandardCharsets; +import java.nio.file.Files; +import java.nio.file.Path; +import java.time.Duration; +import java.util.List; +import java.util.regex.Pattern; +import org.jspecify.annotations.Nullable; + +/** + * {@code /loop}: keep working on one task, step by step, until it is done — the state lives in a file, + * not in the conversation. + * + *

Every step sends the same message: the task verbatim plus an instruction to read + * {@value #LOOP_FILE}, do one concrete step, and write down what happened. The conversation history is + * dropped between steps, which is the point of the file: the context never grows, so the loop + * can run for hours, and what the model knows is exactly what it wrote down. (This is the shape + * Claude Code's own ralph-wiggum plugin uses — re-inject the original prompt and let the filesystem be + * the memory.) + * + *

Why a text marker and not a "done" tool. A small local model produces a well-formed tool + * call far less reliably than a line of text — below 7B, malformed calls are the norm, and + * mini-SWE-agent reaches its SWE-bench results with a plain-text sentinel and no tool-call API at all. + * So the loop ends when a line is exactly {@value #SENTINEL}; a substring never counts, which + * is what keeps "I will answer <<TASK_COMPLETE>> when I am done" from ending the run. + * + *

The model saying "done" is not proof. With {@code --check } the sentinel is only + * accepted when that command succeeds; otherwise its output goes back into the next step. A 4B model + * declares victory early, and this is the cheapest defence against it. + * + *

Limits, all of them enforced here rather than trusted to the model: a step cap + * ({@link LoopOptions#DEFAULT_MAX_STEPS} by default), a wall-clock budget, and a stall detector — if + * neither the file nor any tool call changed anything for {@value #STALL_LIMIT} steps in a row, the + * loop stops instead of burning tokens on a model that repeats itself. + */ +public final class TaskLoop { + + /** The file the loop keeps its plan and its notes in, inside the workspace. */ + public static final String LOOP_FILE = "AGENT-LOOP.md"; + + /** The line that ends the loop. Matched as a whole line, never as a substring. */ + public static final String SENTINEL = "<>"; + + /** Steps without any change before the loop gives up. */ + public static final int STALL_LIMIT = 3; + + /** How long a loop may run before it stops on its own. */ + public static final Duration DEFAULT_BUDGET = Duration.ofHours(2); + + private static final Pattern SENTINEL_LINE = Pattern.compile("^\\s*" + Pattern.quote(SENTINEL) + "\\s*$"); + + private TaskLoop() {} + + /** + * Whether an answer ends the loop. + * + * @param answer the model's complete answer for one step + * @return {@code true} when one of its lines is exactly the sentinel + */ + public static boolean isComplete(String answer) { + return answer.lines().anyMatch(line -> SENTINEL_LINE.matcher(line).matches()); + } + + /** + * The message sent for every step. + * + * @param options the loop options + * @return the prompt, with the task, the file name and the check command filled in + */ + public static String stepPrompt(LoopOptions options) { + String checkHint = options.check() == null + ? "" + : System.lineSeparator() + "Before you declare the task done, run this and make sure it succeeds: " + + options.check(); + return LocalAgent.prompt(LocalAgent.LOOP_PROMPT) + .replace("{task}", options.task()) + .replace("{file}", LOOP_FILE) + .replace("{check_hint}", checkHint); + } + + /** + * Create the loop file when it does not exist yet. + * + * @param workspace the workspace directory + * @param task the task, written into the file + * @return the path of the loop file + */ + public static Path ensureLoopFile(Path workspace, String task) { + Path file = workspace.resolve(LOOP_FILE); + try { + if (!Files.exists(file)) { + Files.writeString( + file, + LocalAgent.prompt(LocalAgent.LOOP_FILE_TEMPLATE).replace("{task}", task) + + System.lineSeparator(), + StandardCharsets.UTF_8); + } + return file; + } catch (IOException e) { + throw new UncheckedIOException("Cannot create " + file, e); + } + } + + /** + * A fingerprint of the loop file, to tell a step that changed something from one that did not. + * + * @param file the loop file + * @return size and content hash, or {@code -1} when the file cannot be read + */ + public static long fingerprint(Path file) { + try { + return Files.exists(file) + ? Files.readString(file, StandardCharsets.UTF_8).hashCode() + : -1; + } catch (IOException e) { + return -1; + } + } + + /** + * Why a loop ended. + * + * @param reason the wording shown to the user + * @param completed whether the model declared the task finished + */ + public record Outcome(String reason, boolean completed) {} + + /** + * Run the loop until it is done, stopped, or out of budget. + * + * @param runner the runner + * @param fileSystem the workspace filesystem for the sessions + * @param terminal the console + * @param workspace the workspace directory + * @param options what to work on and for how long + * @param stopped polled between steps; {@code true} ends the loop (Ctrl-C) + * @param budget the wall-clock limit + * @return why it ended + * @throws InterruptedException if interrupted while waiting for a step or an interval + */ + public static Outcome run( + AgentRunner runner, + org.atmosphere.ai.fs.AgentFileSystem fileSystem, + AgentTerminal terminal, + Path workspace, + LoopOptions options, + java.util.function.BooleanSupplier stopped, + Duration budget) + throws InterruptedException { + Path file = ensureLoopFile(workspace, options.task()); + terminal.line("loop: " + options.task()); + terminal.line("loop: state in " + file + ", max " + options.maxSteps() + " steps, budget " + + budget.toMinutes() + " min" + + (options.interval() == null + ? "" + : ", every " + options.interval().toSeconds() + "s") + + (options.check() == null ? "" : ", check: " + options.check())); + + long deadline = System.nanoTime() + budget.toNanos(); + long lastFingerprint = fingerprint(file); + int stalled = 0; + String extra = ""; + + for (int step = 1; step <= options.maxSteps(); step++) { + if (stopped.getAsBoolean()) { + return new Outcome("stopped after " + (step - 1) + " steps", false); + } + if (System.nanoTime() > deadline) { + return new Outcome( + "budget of " + budget.toMinutes() + " min used up after " + (step - 1) + " steps", false); + } + terminal.line(terminal.ansi().dim("── loop step " + step + "/" + options.maxSteps() + " ──")); + + // A fresh history every step: the file is the memory, so the context cannot grow. + ConsoleSession session = LocalAgent.turn( + runner, fileSystem, stepPrompt(options) + extra, new java.util.ArrayList<>(), terminal); + extra = ""; + if (session.failure() != null) { + return new Outcome("step " + step + " failed: " + session.failure(), false); + } + + // The marker is checked BEFORE the stall detector: a step that only answers "done" changes + // no file and calls no tool, so the other order would report "no progress" on the very + // step that finished the task. + if (isComplete(session.text())) { + String failure = checkFailure(options, workspace); + if (failure == null) { + return new Outcome("done after " + step + " steps", true); + } + terminal.line(terminal.ansi().yellow("loop: the check failed, continuing")); + extra = System.lineSeparator() + "You answered " + SENTINEL + ", but the check (" + + options.check() + ") failed:" + System.lineSeparator() + failure + + System.lineSeparator() + "Fix that first."; + } + + long fingerprint = fingerprint(file); + boolean changed = fingerprint != lastFingerprint || session.toolCalls() > 0; + lastFingerprint = fingerprint; + stalled = changed ? 0 : stalled + 1; + if (stalled >= STALL_LIMIT) { + return new Outcome("no progress for " + STALL_LIMIT + " steps (nothing written, no tools used)", false); + } + + if (options.interval() != null && step < options.maxSteps()) { + Thread.sleep(options.interval().toMillis()); + } + } + return new Outcome("step limit of " + options.maxSteps() + " reached", false); + } + + /** + * Run the check command, if there is one. + * + * @param options the loop options + * @param workspace the directory the command runs in + * @return the command output when it failed, or {@code null} when it succeeded or there is none + */ + private static @Nullable String checkFailure(LoopOptions options, Path workspace) { + String command = options.check(); + if (command == null) { + return null; + } + try { + String result = ShellTool.run(workspace, command, Duration.ofMinutes(10), 4000); + return result.startsWith("exit code: 0") ? null : result; + } catch (IOException e) { + return "the check could not be started: " + e.getMessage(); + } catch (InterruptedException e) { + Thread.currentThread().interrupt(); + return "the check was interrupted"; + } + } + + /** + * The tool names a loop needs; used only for the console hint. + * + * @param toolNames the offered tools + * @return {@code true} when the file tools that the loop file needs are present + */ + public static boolean canKeepNotes(List toolNames) { + return toolNames.contains("read_file") && toolNames.contains("write_file"); + } +} diff --git a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt index 604e98310..0f1a4d58f 100644 --- a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt +++ b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt @@ -5,6 +5,9 @@ Commands (everything else is sent to the model): /tools the tools offered to the model /mode [manual|auto] show or set the approval mode (/approve) /compact [focus] summarize the history and continue with the summary + /loop [--every 5m] [--max 20] [--check ''] + keep working on until the model answers <>; + the state lives in AGENT-LOOP.md, not in the conversation /clear drop the history (/reset, /new) /exit leave (/quit) diff --git a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/loop-file-template.md b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/loop-file-template.md new file mode 100644 index 000000000..8642b4ffe --- /dev/null +++ b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/loop-file-template.md @@ -0,0 +1,17 @@ +# Agent loop + +## Task + +{task} + +## Plan + +- [ ] first step (replace this with the real plan) + +## Notes + +What was tried, what worked, what failed. Keep it short; this file is re-read every step. + +## Next step + +One sentence: what to do next. diff --git a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/loop-prompt.txt b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/loop-prompt.txt new file mode 100644 index 000000000..34a7b7f09 --- /dev/null +++ b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/loop-prompt.txt @@ -0,0 +1,12 @@ +Work on this task, one step at a time: + +{task} + +The file {file} is your memory between steps — the conversation itself is not kept, so anything you do not write down is lost. Read that file first. Then do exactly ONE concrete step: make the change, run the command, check the result. Afterwards update {file}: tick off what you finished, note what you learned or what failed, and write down what the next step is. + +When the whole task is finished and you have verified it, answer with exactly this line and nothing else: + +<> + +Never write that line for any other reason, never write it while planning, and never explain it. +{check_hint} diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleFormattingTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleFormattingTest.java index c406eca3a..f573d32f1 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleFormattingTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleFormattingTest.java @@ -49,10 +49,23 @@ void aPlainInstanceReturnsTheTextUnchanged() { // ----- StatusLine ----- @Test - void theStatusLineShowsModeContextToolsAndModel() { - String line = StatusLine.render(ApprovalMode.MANUAL, 1234, false, 16384, 9, "local-model"); + void theStatusLineShowsWorkspaceModeContextToolsAndModel() { + String line = StatusLine.render( + java.nio.file.Path.of("/tmp/ws"), ApprovalMode.MANUAL, 1234, false, 16384, 9, "local-model"); - assertThat(line, is("[manual · ctx 1.2k/16k · 9 tools · local-model]")); + assertThat(line, containsString("ws · manual · ctx 1.2k/16k · 9 tools · local-model]")); + } + + @Test + void aLongWorkspacePathIsShortenedToItsLastTwoSegments() { + // the path is on every line of the session, so it must not push the rest off the screen + java.nio.file.Path deep = java.nio.file.Path.of("/home/someone/projects/customer/service/backend/module"); + assertThat(StatusLine.shorten(deep), containsString("backend")); + assertThat(StatusLine.shorten(deep), containsString("module")); + assertThat(StatusLine.shorten(deep).startsWith("…"), is(true)); + assertThat( + StatusLine.shorten(java.nio.file.Path.of("/tmp/ws")), + is(java.nio.file.Path.of("/tmp/ws").toString())); } @Test diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/TaskLoopTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/TaskLoopTest.java new file mode 100644 index 000000000..3058836df --- /dev/null +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/TaskLoopTest.java @@ -0,0 +1,245 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import static org.hamcrest.MatcherAssert.assertThat; +import static org.hamcrest.Matchers.containsString; +import static org.hamcrest.Matchers.is; +import static org.hamcrest.Matchers.nullValue; + +import java.io.ByteArrayOutputStream; +import java.io.PrintStream; +import java.nio.charset.StandardCharsets; +import java.nio.file.Files; +import java.nio.file.Path; +import java.time.Duration; +import java.util.List; +import net.ladenthin.llama.server.OpenAiCompatServer; +import net.ladenthin.llama.server.OpenAiServerConfig; +import org.atmosphere.ai.RetryPolicy; +import org.atmosphere.ai.fs.AgentFileSystem; +import org.atmosphere.ai.fs.WorkspaceAgentFileSystem; +import org.junit.jupiter.api.Test; +import org.junit.jupiter.api.io.TempDir; + +/** The loop: when it stops, when it does not, and what it sends. */ +class TaskLoopTest { + + private static final String MODEL_ID = "local-model"; + private static final Duration BUDGET = Duration.ofMinutes(5); + + @TempDir + Path workspace; + + private final ByteArrayOutputStream console = new ByteArrayOutputStream(); + + private AgentTerminal terminal() { + return new PlainTerminal(new PrintStream(console, true, StandardCharsets.UTF_8), null, Ansi.PLAIN); + } + + private TaskLoop.Outcome runLoop(ScriptedBackend backend, LoopOptions options) throws Exception { + OpenAiServerConfig config = OpenAiServerConfig.builder() + .host("127.0.0.1") + .port(0) + .apiKey("k") + .modelId(MODEL_ID) + .build(); + try (OpenAiCompatServer server = new OpenAiCompatServer(backend, config).start()) { + AgentRunner runner = new AgentRunner( + "http://127.0.0.1:" + server.getPort() + "/v1", + "k", + MODEL_ID, + List.of(), + "You are a test agent.", + 0.0, + 64, + 5) + .retryPolicy(RetryPolicy.NONE); + AgentFileSystem fs = new WorkspaceAgentFileSystem(workspace, AgentFileSystem.Limits.defaults()); + return TaskLoop.run(runner, fs, terminal(), workspace, options, () -> false, BUDGET); + } + } + + // ----- the stop marker ----- + + @Test + void onlyAWholeLineEndsTheLoop() { + assertThat(TaskLoop.isComplete("all done\n" + TaskLoop.SENTINEL), is(true)); + assertThat(TaskLoop.isComplete(" " + TaskLoop.SENTINEL + " "), is(true)); + // the failure this guards against: the model talking ABOUT the marker + assertThat(TaskLoop.isComplete("I will answer " + TaskLoop.SENTINEL + " when I am done."), is(false)); + assertThat(TaskLoop.isComplete("TASK_COMPLETE"), is(false)); + assertThat(TaskLoop.isComplete("still working"), is(false)); + } + + @Test + void everyStepSendsTheTaskVerbatimAndNamesTheFile() { + String prompt = TaskLoop.stepPrompt(new LoopOptions("rename the class", null, 5, null)); + + assertThat(prompt, containsString("rename the class")); + assertThat(prompt, containsString(TaskLoop.LOOP_FILE)); + assertThat(prompt, containsString(TaskLoop.SENTINEL)); + assertThat(prompt.contains("{"), is(false)); + } + + @Test + void theCheckCommandIsPartOfTheInstructionsWhenThereIsOne() { + assertThat( + TaskLoop.stepPrompt(new LoopOptions("fix it", null, 5, "mvn -q test")), containsString("mvn -q test")); + } + + // ----- the file ----- + + @Test + void theLoopFileIsCreatedOnceAndKeepsWhatIsInIt() throws Exception { + Path file = TaskLoop.ensureLoopFile(workspace, "write a parser"); + + assertThat(Files.readString(file), containsString("write a parser")); + Files.writeString(file, "edited by the agent"); + assertThat(TaskLoop.ensureLoopFile(workspace, "write a parser"), is(file)); + assertThat(Files.readString(file), is("edited by the agent")); + } + + @Test + void theFingerprintChangesWithTheContent() throws Exception { + Path file = TaskLoop.ensureLoopFile(workspace, "task"); + long before = TaskLoop.fingerprint(file); + + Files.writeString(file, "something else"); + + assertThat(TaskLoop.fingerprint(file) != before, is(true)); + assertThat(TaskLoop.fingerprint(workspace.resolve("absent.md")), is(-1L)); + } + + // ----- the loop itself, over the real server ----- + + @Test + void theLoopEndsWhenTheModelAnswersWithTheMarker() throws Exception { + ScriptedBackend backend = new ScriptedBackend((call, request) -> call < 3 + ? ScriptedBackend.textTurn("working on step ", String.valueOf(call)) + : ScriptedBackend.textTurn(TaskLoop.SENTINEL)); + + TaskLoop.Outcome outcome = runLoop(backend, new LoopOptions("do the thing", null, 10, null)); + + assertThat(outcome.completed(), is(true)); + assertThat(outcome.reason(), containsString("3 steps")); + assertThat(backend.requests().size(), is(3)); + } + + @Test + void everyStepStartsFromAnEmptyHistorySoTheContextCannotGrow() throws Exception { + ScriptedBackend backend = new ScriptedBackend((call, request) -> + call < 2 ? ScriptedBackend.textTurn("step") : ScriptedBackend.textTurn(TaskLoop.SENTINEL)); + + runLoop(backend, new LoopOptions("do the thing", null, 10, null)); + + // system + user, every single time: the file is the memory, the conversation is not + for (var request : backend.requests()) { + assertThat(request.path("messages").size(), is(2)); + } + } + + @Test + void aModelThatRepeatsItselfWithoutChangingAnythingIsStopped() throws Exception { + // no tool calls, no file change: exactly the shape of a model looping on its own answer + ScriptedBackend backend = new ScriptedBackend((call, request) -> ScriptedBackend.textTurn("thinking…")); + + TaskLoop.Outcome outcome = runLoop(backend, new LoopOptions("do the thing", null, 50, null)); + + assertThat(outcome.completed(), is(false)); + assertThat(outcome.reason(), containsString("no progress")); + assertThat(backend.requests().size(), is(TaskLoop.STALL_LIMIT)); + } + + @Test + void theStepLimitStopsALoopThatWouldOtherwiseRunOn() throws Exception { + ScriptedBackend backend = new ScriptedBackend((call, request) -> { + // something changes every step, so the stall detector never fires + Files.writeString(workspace.resolve(TaskLoop.LOOP_FILE), "step " + call); + return ScriptedBackend.textTurn("still working"); + }); + + TaskLoop.Outcome outcome = runLoop(backend, new LoopOptions("endless", null, 2, null)); + + assertThat(outcome.completed(), is(false)); + assertThat(outcome.reason(), containsString("step limit")); + assertThat(backend.requests().size(), is(2)); + } + + @Test + void aFailingCheckRejectsTheMarkerAndFeedsTheOutputBack() throws Exception { + String failing = ShellTool.isWindows() ? "exit 7" : "exit 7"; + ScriptedBackend backend = new ScriptedBackend((call, request) -> { + Files.writeString(workspace.resolve(TaskLoop.LOOP_FILE), "step " + call); + return ScriptedBackend.textTurn(TaskLoop.SENTINEL); + }); + + TaskLoop.Outcome outcome = runLoop(backend, new LoopOptions("claim it works", null, 2, failing)); + + assertThat("the model said done, the check said otherwise", outcome.completed(), is(false)); + assertThat( + backend.requests() + .get(1) + .path("messages") + .get(1) + .path("content") + .asText(), + containsString("the check")); + } + + // ----- the argument parser ----- + + @Test + void theTaskIsEverythingAfterTheFlags() { + LoopOptions plain = LoopOptions.parse("keep the README in sync with the code"); + + assertThat(plain.task(), is("keep the README in sync with the code")); + assertThat(plain.interval(), is(nullValue())); + assertThat(plain.maxSteps(), is(LoopOptions.DEFAULT_MAX_STEPS)); + assertThat(plain.check(), is(nullValue())); + } + + @Test + void flagsAreReadFromTheFrontAndAQuotedCheckKeepsItsSpaces() { + LoopOptions options = LoopOptions.parse("--every 90s --max 3 --check 'mvn -q test' fix the build"); + + assertThat(options.interval(), is(Duration.ofSeconds(90))); + assertThat(options.maxSteps(), is(3)); + assertThat(options.check(), is("mvn -q test")); + assertThat(options.task(), is("fix the build")); + } + + @Test + void durationsAreWrittenAsPeopleWriteThem() { + assertThat(LoopOptions.parseDuration("30s"), is(Duration.ofSeconds(30))); + assertThat(LoopOptions.parseDuration("5m"), is(Duration.ofMinutes(5))); + assertThat(LoopOptions.parseDuration("2h"), is(Duration.ofHours(2))); + assertThat(LoopOptions.parseDuration("10"), is(Duration.ofMinutes(10))); + } + + @Test + void aMalformedLoopCommandSaysWhatIsWrong() { + assertThat( + org.junit.jupiter.api.Assertions.assertThrows( + IllegalArgumentException.class, () -> LoopOptions.parse("")) + .getMessage(), + containsString("Usage: /loop")); + assertThat( + org.junit.jupiter.api.Assertions.assertThrows( + IllegalArgumentException.class, () -> LoopOptions.parse("--every soon do it")) + .getMessage(), + containsString("Not a duration")); + assertThat( + org.junit.jupiter.api.Assertions.assertThrows( + IllegalArgumentException.class, () -> LoopOptions.parse("--max zero do it")) + .getMessage(), + containsString("--max expects a number")); + assertThat( + org.junit.jupiter.api.Assertions.assertThrows( + IllegalArgumentException.class, () -> LoopOptions.parse("--bogus x do it")) + .getMessage(), + containsString("Unknown /loop flag")); + } +} From 68d4cac8198052cc3e531cdca2a7f211bd82a972 Mon Sep 17 00:00:00 2001 From: Bernard Ladenthin Date: Wed, 23 Sep 2026 10:06:59 +0200 Subject: [PATCH 05/36] llama-atmosphere-agent: own read_file, edit_file and grep MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Three of Atmosphere's file tools are replaced (not added alongside — two tools that both read a file is the worst case for tool selection). They call the same AgentFileSystem, so confinement, path validation and size limits are unchanged; what changes is the parameters the model sees and the text it gets back. edit_file (TextEdits): - Normalizes line endings before matching and restores the file's own ending and BOM. The framework matches raw content, so on Windows a model's LF text never matches a CRLF file and every edit fails silently. - A miss shows the closest lines with their numbers instead of "not found"; an ambiguous match names the lines and offers replace_all. A failed edit is not a free retry: on SWE-agent trajectories the eventual success rate drops from 90.5% to 57.2% after one failure. - edits[] applies several replacements all-or-nothing. - Refuses to edit a file that was not read in this session (ReadTracker), so old_string comes from the file rather than from memory. grep (WorkspaceSearch): the framework walks alphabetically with one global 2s deadline and one global 500-hit budget, so .git, target and node_modules consume both before src is reached. This one excludes them, groups by file with line numbers, caps at 100 matches and says when it truncated. read_file: offset/limit and numbered lines; whole-file reads measure 12.7% vs 18.0% task success in the SWE-agent ablations. The numbers are display only, which the tool description and the system prompt both say. Rejected on evidence (see CLAUDE.md): unified diff/patch, fuzzy matching, an embedding index, LSP tools. Tests: 101 green (was 81). Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01E2h8gXyyE5UeimkL9vQv9G --- CLAUDE.md | 28 +- llama-atmosphere-agent/README.md | 44 +- .../llama/atmosphere/LocalAgent.java | 5 +- .../ladenthin/llama/atmosphere/TextEdits.java | 304 +++++++++++++ .../llama/atmosphere/WorkspaceSearch.java | 243 +++++++++++ .../llama/atmosphere/WorkspaceTools.java | 398 ++++++++++++++++++ .../llama/atmosphere/system-prompt.txt | 2 +- .../llama/atmosphere/LocalAgentTest.java | 7 +- .../llama/atmosphere/TextEditsTest.java | 138 ++++++ .../llama/atmosphere/WorkspaceToolsTest.java | 205 +++++++++ 10 files changed, 1364 insertions(+), 10 deletions(-) create mode 100644 llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/TextEdits.java create mode 100644 llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/WorkspaceSearch.java create mode 100644 llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/WorkspaceTools.java create mode 100644 llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/TextEditsTest.java create mode 100644 llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/WorkspaceToolsTest.java diff --git a/CLAUDE.md b/CLAUDE.md index 1dcf65dc0..f110ff36f 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -2277,7 +2277,33 @@ are decisions, not details: and no tool call), interval. **Order matters and a test pins it**: the marker is checked *before* the stall detector, because the step that only answers "done" changes nothing and would otherwise be reported as no progress. -6. **The context number in the status line is an estimate, marked `~`.** llama.cpp emits its usage +6. **Three of the file tools are this project's, not Atmosphere's** (`WorkspaceTools`, `TextEdits`, + `WorkspaceSearch`) — `read_file`, `edit_file`, `grep`. They are **replacements, not additions**: + two tools that both claim to read a file is the worst case for tool selection. They call the same + `AgentFileSystem`, so workspace confinement, path validation and size limits stay Atmosphere's. + Each replacement has a measured reason, and all three are pinned by tests: + - `edit_file`: the framework matches against the **raw** file content, so a model's LF text never + matches a CRLF file — on Windows *every* edit fails silently. `TextEdits` normalizes before + matching and restores the file's own ending and byte-order mark. It also shows the nearest lines + on a miss and the line numbers on an ambiguous match (a failed edit drops the eventual success + rate from 90.5 % to 57.2 %), offers `replace_all`, and applies a batch of edits **all-or-nothing** + — a deviation from every shipping agent, which apply sequentially and leave a half-edited file. + - `grep`: the framework walks **alphabetically** with one global 2-second deadline and one global + 500-hit budget, so `.git`, `target` and `node_modules` consume both before `src` is reached. + `WorkspaceSearch` excludes them, groups by file with line numbers, and **states** truncation. + - `read_file`: `offset`/`limit` and numbered lines (whole-file reads measure 12.7 % against 18.0 % + task success in the SWE-agent ablations). The numbers are display only, which both the tool + description and the system prompt say — leaked line numbers in `old_string` are a known failure. + - `edit_file` **refuses a file that was not read** in this session (`WorkspaceTools.ReadTracker`). + Not a staleness check: an exact unambiguous match is safe regardless; this catches the model + inventing the text. + **Rejected on evidence, do not add later without new numbers:** a unified-diff/patch tool (Meta's + ablation: search-replace 42–53 % vs 26–30 % unified diff vs 20–26 % line diff on one model; a 7B + model collapses 54 → 33 → 14 %), fuzzy matching (turns a loud miss into a silent wrong-place edit), + an embedding index (Cursor's production effect is +0.3 %), and LSP tools (the one isolation study + finds them token-negative and *worse* at multi-file rename, because renames touch comments and + strings that semantic references exclude). +7. **The context number in the status line is an estimate, marked `~`.** llama.cpp emits its usage chunk only when the client sets `stream_options.include_usage`, and Atmosphere's client does not; `ConsoleSession.usage()` takes the real count when one arrives, otherwise `LocalAgent.estimateTokens` uses four characters per token. The window size is `--ctx-size` (in-process) or the server's diff --git a/llama-atmosphere-agent/README.md b/llama-atmosphere-agent/README.md index 9929cc408..8b0a46f99 100644 --- a/llama-atmosphere-agent/README.md +++ b/llama-atmosphere-agent/README.md @@ -15,9 +15,10 @@ offline: you start yourself, or the GGUF loaded **in this process**. - **Agent:** [Atmosphere](https://github.com/Atmosphere/atmosphere)'s built-in OpenAI-compatible runtime (`org.atmosphere:atmosphere-ai`): streaming, the model→tool→model loop, and its - workspace-confined file tools (`ls`, `read_file`, `write_file`, `edit_file`, `glob`, `grep`, - `delete`, `rename`). Driven **headless** — no Spring Boot, no servlet container, no `@Agent` - scanning — through `BuiltInAgentRuntime`. + workspace-confined file tools (`ls`, `write_file`, `glob`, `delete`, `rename`). Driven **headless** + — no Spring Boot, no servlet container, no `@Agent` scanning — through `BuiltInAgentRuntime`. + `read_file`, `edit_file` and `grep` are this project's own (see [Tools](#tools)), on the same + workspace-confined filesystem. - **Shell:** an opt-in `run_command` tool (`--allow-shell`) that runs any command line through the system shell (`cmd.exe` on Windows, `sh` elsewhere). @@ -289,6 +290,43 @@ Then, in this order: `--auto` starts in auto mode, `--verbose` brings llama.cpp's own log back, and `NO_COLOR=1` turns the styling off. +### Tools + +Eight file tools plus the optional shell. Five are Atmosphere's; three are replaced here because what +they return decides how well a model can work: + +| Tool | | | +|---|---|---| +| `ls`, `write_file`, `glob`, `delete`, `rename` | Atmosphere | unchanged | +| `read_file(file_path, offset, limit)` | **ours** | a numbered window, ` 12: text`, 400 lines at a time, and it says what it left out. Reading whole files costs context and measurably lowers task success (SWE-agent: 12.7 % with whole files against 18.0 % with a 100-line window) | +| `edit_file(file_path, old_string, new_string, replace_all, edits[])` | **ours** | see below | +| `grep(pattern, dir, glob, files_only)` | **ours** | skips `.git`, `target`, `build`, `node_modules` and friends, groups matches by file with line numbers, caps at 100 matches and **says so** when it truncates | +| `run_command` | ours | opt-in via `--allow-shell` | + +**Why `edit_file` is not Atmosphere's.** Four things it does that the framework's does not, each for a +measured reason: + +1. **Line endings.** The framework compares the raw file content, so a model's LF text never matches a + CRLF file — on Windows every edit fails silently. Here the file is normalized before matching and + written back in its own ending (byte-order mark included). +2. **A miss explains itself.** Instead of "not found", it shows the closest lines in the file with + their numbers. A failed edit is not a free retry: measured on SWE-agent trajectories, an edit + attempt eventually succeeds in 90.5 % of cases, but only 57.2 % once one edit has failed. +3. **An ambiguous match names the lines** (`occurs 2 times, on lines 1, 3`) instead of asking for + "more context", and `replace_all` is offered. +4. **Several edits in one call** via `edits: [{old_string, new_string}]`, applied **all or nothing** — + every shipping agent applies them sequentially and leaves a half-edited file behind. + +`edit_file` also **refuses to edit a file that was not read** in this session, so `old_string` comes +from the file rather than from the model's memory. + +**Formats that were considered and rejected.** A unified-diff or patch tool: Meta's ablation measures +search-replace at 42–53 % against 26–30 % for unified diff and 20–26 % for line diffs on the same +model, and a 7B model drops from 54 % to 33 % to 14 % across those three. Fuzzy matching (a similarity +threshold instead of an exact match): it turns a loud "not found" into a silent edit in the wrong +place. Whole-file rewriting stays available as `write_file` — for a small model that is often the most +reliable route, and the system prompt says so. + ### The system prompt Without `--system` the agent uses a built-in **general-purpose** prompt: it names the file tools and diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java index c444e45a6..bf06ac876 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java @@ -21,7 +21,6 @@ import net.ladenthin.llama.server.OpenAiCompatServer; import net.ladenthin.llama.server.OpenAiServerConfig; import org.atmosphere.ai.fs.AgentFileSystem; -import org.atmosphere.ai.fs.FileSystemTools; import org.atmosphere.ai.fs.WorkspaceAgentFileSystem; import org.atmosphere.ai.llm.ChatMessage; import org.atmosphere.ai.tool.ToolDefinition; @@ -151,7 +150,9 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream } AgentFileSystem fileSystem = new WorkspaceAgentFileSystem(options.getWorkspace(), AgentFileSystem.Limits.defaults()); - List tools = new ArrayList<>(FileSystemTools.all()); + // Our own read_file/edit_file/grep replace the framework's (see WorkspaceTools); the + // read tracker is what lets an edit insist the file was read first. + List tools = new ArrayList<>(WorkspaceTools.all(new WorkspaceTools.ReadTracker())); if (options.isAllowShell()) { tools.add(ShellTool.definition(options.getWorkspace(), SHELL_TIMEOUT, SHELL_MAX_OUTPUT_CHARS)); } diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/TextEdits.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/TextEdits.java new file mode 100644 index 000000000..b7888630a --- /dev/null +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/TextEdits.java @@ -0,0 +1,304 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import java.util.ArrayList; +import java.util.List; +import org.jspecify.annotations.Nullable; + +/** + * The text side of editing a file: what to match, what to write back, and what to say when it does + * not match. No file access, no tools — so every rule here is testable on plain strings. + * + *

Line endings are the reason this class exists. A model answers in LF, always; a file on + * Windows is usually CRLF. Comparing the two directly never matches, which is why the naive + * implementation silently fails on every Windows file. So the file is normalized to LF before + * matching and written back in its own ending, and a byte-order mark is taken off first and + * put back afterwards — it is invisible, and it sits exactly where the first match would be. + * + *

A failed edit is not a free retry. Measured on SWE-agent trajectories: any edit attempt + * eventually succeeds in 90.5 % of cases, but only 57.2 % once a single edit has failed — a failure + * derails the rest of the run. That is why a miss does not answer "not found" but shows the closest + * lines in the file, and an ambiguous match names the line numbers instead of asking for "more + * context". aider's threshold is used for the closest-line search: similarity ≥ {@value #CANDIDATE_MIN_SIMILARITY}. + * + *

Several edits are all-or-nothing, which deviates from every shipping agent — they apply + * sequentially and leave a half-edited file behind. Here the whole batch works on one in-memory + * string and is written only if every edit matched, because at that point atomicity costs nothing and + * a half-applied batch is the state a model reasons about worst. + */ +public final class TextEdits { + + /** How similar a line must be to be worth showing after a failed match (aider's threshold). */ + public static final double CANDIDATE_MIN_SIMILARITY = 0.6; + + /** How many near misses are shown. */ + public static final int MAX_CANDIDATES = 3; + + /** The byte-order mark, as the character it decodes to. */ + private static final char BOM = ''; + + private TextEdits() {} + + /** + * One replacement. + * + * @param oldString the text to find, in LF form + * @param newString what replaces it + * @param replaceAll whether every occurrence is replaced instead of requiring exactly one + */ + public record Edit(String oldString, String newString, boolean replaceAll) {} + + /** Thrown when an edit cannot be applied; the message is what the model gets to read. */ + public static final class EditException extends RuntimeException { + private static final long serialVersionUID = 1L; + + /** + * Create the exception. + * + * @param message the explanation shown to the model + */ + public EditException(String message) { + super(message); + } + } + + /** + * The line ending a text uses. + * + * @param text the file content as read + * @return {@code "\r\n"} when the first line ending is a carriage-return pair, else {@code "\n"} + */ + public static String lineEnding(String text) { + int newline = text.indexOf('\n'); + return newline > 0 && text.charAt(newline - 1) == '\r' ? "\r\n" : "\n"; + } + + /** + * Strip a byte-order mark and convert every line ending to LF. + * + * @param text the file content as read + * @return the normalized content + */ + public static String normalize(String text) { + String withoutBom = text.isEmpty() || text.charAt(0) != BOM ? text : text.substring(1); + return withoutBom.replace("\r\n", "\n").replace("\r", "\n"); + } + + /** + * Put the file's own line ending and byte-order mark back. + * + * @param normalized the edited content in LF form + * @param original the content as it was read, for its ending and mark + * @return the content to write + */ + public static String denormalize(String normalized, String original) { + String ending = lineEnding(original); + String restored = "\n".equals(ending) ? normalized : normalized.replace("\n", ending); + return !original.isEmpty() && original.charAt(0) == BOM ? BOM + restored : restored; + } + + /** + * Apply every edit to {@code original}, or none of them. + * + * @param original the file content as read + * @param edits the edits, applied in order to the result of the previous one + * @return the content to write back, with the original line ending and mark + * @throws EditException when any edit does not match, naming what went wrong + */ + public static String apply(String original, List edits) { + if (edits.isEmpty()) { + throw new EditException("No edits were given."); + } + String content = normalize(original); + for (int i = 0; i < edits.size(); i++) { + Edit edit = edits.get(i); + String where = edits.size() == 1 ? "" : " (edit " + (i + 1) + " of " + edits.size() + ")"; + content = applyOne(content, edit, where); + } + return denormalize(content, original); + } + + private static String applyOne(String content, Edit edit, String where) { + String oldString = normalize(edit.oldString()); + if (oldString.isEmpty()) { + throw new EditException("old_string must not be empty" + where + "."); + } + if (oldString.equals(edit.newString())) { + throw new EditException("old_string and new_string are identical" + where + "."); + } + List lines = matchLines(content, oldString); + if (lines.isEmpty()) { + throw new EditException(notFoundMessage(content, oldString, where)); + } + if (lines.size() > 1 && !edit.replaceAll()) { + throw new EditException("old_string occurs " + lines.size() + " times" + where + ", on lines " + join(lines) + + ". Add surrounding lines to make it unique, or set replace_all."); + } + return edit.replaceAll() + ? content.replace(oldString, edit.newString()) + : replaceFirst(content, oldString, edit.newString()); + } + + private static String replaceFirst(String content, String oldString, String newString) { + int index = content.indexOf(oldString); + return content.substring(0, index) + newString + content.substring(index + oldString.length()); + } + + /** + * The 1-based line numbers where {@code oldString} starts. + * + * @param content the normalized content + * @param oldString the normalized text to find + * @return every match position, as line numbers + */ + static List matchLines(String content, String oldString) { + List lines = new ArrayList<>(); + int index = content.indexOf(oldString); + while (index >= 0) { + lines.add(lineOf(content, index)); + index = content.indexOf(oldString, index + 1); + } + return lines; + } + + private static int lineOf(String content, int index) { + int line = 1; + for (int i = 0; i < index; i++) { + if (content.charAt(i) == '\n') { + line++; + } + } + return line; + } + + /** + * The message for a miss: the closest lines in the file, so the next attempt can be corrected + * rather than guessed. + * + * @param content the normalized content + * @param oldString the normalized text that was not found + * @param where which edit of the batch failed + * @return the message + */ + static String notFoundMessage(String content, String oldString, String where) { + StringBuilder message = new StringBuilder("old_string was not found" + where + " — nothing was changed."); + List candidates = candidates(content, oldString); + if (candidates.isEmpty()) { + message.append(" No similar line exists; read the file again and copy the text from it" + + " (without the line numbers the read tool prints)."); + } else { + message.append(" The closest lines in the file are:"); + for (String candidate : candidates) { + message.append(System.lineSeparator()).append(" ").append(candidate); + } + } + return message.toString(); + } + + /** + * The lines most similar to the first line of {@code oldString}. + * + * @param content the normalized content + * @param oldString the text that was not found + * @return up to {@value #MAX_CANDIDATES} lines as {@code " 12: text"}, best first + */ + static List candidates(String content, String oldString) { + String needle = oldString.lines().findFirst().orElse(oldString).strip(); + if (needle.isEmpty()) { + return List.of(); + } + record Candidate(int line, String text, double score) {} + List scored = new ArrayList<>(); + String[] lines = content.split("\n", -1); + for (int i = 0; i < lines.length; i++) { + double score = similarity(needle, lines[i].strip()); + if (score >= CANDIDATE_MIN_SIMILARITY) { + scored.add(new Candidate(i + 1, lines[i], score)); + } + } + return scored.stream() + .sorted((a, b) -> Double.compare(b.score(), a.score())) + .limit(MAX_CANDIDATES) + .map(candidate -> candidate.line() + ": " + candidate.text()) + .toList(); + } + + /** + * How similar two lines are, as 1 minus the edit distance over the longer length. + * + * @param a one line + * @param b the other + * @return a value between 0 and 1 + */ + static double similarity(String a, String b) { + if (a.equals(b)) { + return 1; + } + if (a.isEmpty() || b.isEmpty()) { + return 0; + } + int distance = editDistance(a, b); + return 1.0 - (double) distance / Math.max(a.length(), b.length()); + } + + private static int editDistance(String a, String b) { + int[] previous = new int[b.length() + 1]; + int[] current = new int[b.length() + 1]; + for (int j = 0; j <= b.length(); j++) { + previous[j] = j; + } + for (int i = 1; i <= a.length(); i++) { + current[0] = i; + for (int j = 1; j <= b.length(); j++) { + int substitution = previous[j - 1] + (a.charAt(i - 1) == b.charAt(j - 1) ? 0 : 1); + current[j] = Math.min(substitution, Math.min(previous[j] + 1, current[j - 1] + 1)); + } + int[] swap = previous; + previous = current; + current = swap; + } + return previous[b.length()]; + } + + private static String join(List lines) { + StringBuilder text = new StringBuilder(); + for (int i = 0; i < lines.size(); i++) { + text.append(i == 0 ? "" : ", ").append(lines.get(i)); + } + return text.toString(); + } + + /** + * The lines around an edit, so the answer shows what the file now looks like instead of the whole + * file or nothing at all. + * + * @param content the content after the edit, in LF form + * @param oldString the text that was replaced, to locate the region + * @param newString what it was replaced with + * @param context how many lines above and below are included + * @return the snippet as {@code " 12: text"} lines, or {@code null} when the region is gone + */ + static @Nullable String snippet(String content, String oldString, String newString, int context) { + int index = newString.isEmpty() ? content.indexOf(normalize(oldString)) : content.indexOf(newString); + if (index < 0) { + return null; + } + int match = lineOf(content, index); + int start = Math.max(1, match - context); + // the replacement may be several lines; the window ends after it, not after a fixed size + int end = match + Math.max(0, (int) newString.lines().count() - 1) + context; + String[] lines = content.split("\n", -1); + StringBuilder text = new StringBuilder(); + for (int i = start; i <= Math.min(end, lines.length); i++) { + text.append(i == start ? "" : System.lineSeparator()) + .append(" ") + .append(i) + .append(": ") + .append(lines[i - 1]); + } + return text.toString(); + } +} diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/WorkspaceSearch.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/WorkspaceSearch.java new file mode 100644 index 000000000..7bca82cf8 --- /dev/null +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/WorkspaceSearch.java @@ -0,0 +1,243 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import java.io.IOException; +import java.nio.charset.StandardCharsets; +import java.nio.file.FileVisitResult; +import java.nio.file.Files; +import java.nio.file.Path; +import java.nio.file.SimpleFileVisitor; +import java.nio.file.attribute.BasicFileAttributes; +import java.util.ArrayList; +import java.util.LinkedHashMap; +import java.util.List; +import java.util.Map; +import java.util.Set; +import java.util.regex.Matcher; +import java.util.regex.Pattern; +import java.util.regex.PatternSyntaxException; + +/** + * Searching the workspace for text. + * + *

Why this exists instead of the framework's grep. Atmosphere walks the workspace + * alphabetically and spends one global 2-second deadline and one global 500-hit budget on + * that walk. In any real project {@code .git} sorts before {@code src}, so both budgets are consumed + * by the repository's own object store, the build output and {@code node_modules} before the source + * is reached — the tool then reports a truncated result that does not contain the code at all. + * Excluding those directories is the entire fix, and it is a filter in the walk. + * + *

The second reason is what the model reads. Results are grouped by file with line numbers (so a + * ranged read can follow), every line is cut at {@value #MAX_LINE_CHARS} characters (a minified file + * is otherwise a single 200 KB "line"), and truncation is always stated: a search that + * silently returns part of the matches is read as a complete answer, which is the documented way this + * class of tool misleads an agent. + */ +public final class WorkspaceSearch { + + /** Directories never searched, whatever the pattern. */ + public static final Set EXCLUDED_DIRECTORIES = Set.of( + ".git", + ".hg", + ".svn", + "node_modules", + "target", + "build", + "dist", + "out", + ".gradle", + ".idea", + ".mvn", + ".venv", + "venv", + "__pycache__", + ".cache", + ".next", + "vendor"); + + /** Files larger than this are skipped: they are generated, not written. */ + public static final long MAX_FILE_BYTES = 2_000_000; + + /** A matching line is cut here. */ + public static final int MAX_LINE_CHARS = 200; + + /** Matches shown per file before the rest is summarized. */ + public static final int MAX_MATCHES_PER_FILE = 20; + + /** The whole search stops here. */ + public static final int MAX_TOTAL_MATCHES = 100; + + /** A pattern that backtracks is cut off after this. */ + public static final long DEADLINE_MILLIS = 5_000; + + private WorkspaceSearch() {} + + /** + * One matching line. + * + * @param path the path relative to the workspace root, with {@code /} separators + * @param line the 1-based line number + * @param text the line, already cut to {@value #MAX_LINE_CHARS} characters + */ + public record Hit(String path, int line, String text) {} + + /** + * The result of a search. + * + * @param hits the matches, in walk order + * @param filesWithMatches how many files matched, including those not shown + * @param truncated whether the search stopped before the end + */ + public record Result(List hits, int filesWithMatches, boolean truncated) {} + + /** + * Search below {@code root}. + * + * @param root the workspace root + * @param pattern the regular expression + * @param subdirectory the directory to search, relative to the root, or {@code null} for all + * @param glob an optional file-name glob such as {@code *.java}, or {@code null} for all files + * @return the matches + * @throws IllegalArgumentException when the pattern is not a valid regular expression + */ + public static Result search(Path root, String pattern, String subdirectory, String glob) { + Pattern regex; + try { + regex = Pattern.compile(pattern); + } catch (PatternSyntaxException e) { + throw new IllegalArgumentException("Invalid regular expression: " + e.getMessage(), e); + } + Path start = subdirectory == null || subdirectory.isBlank() + ? root + : root.resolve(subdirectory).normalize(); + if (!start.startsWith(root) || !Files.isDirectory(start)) { + return new Result(List.of(), 0, false); + } + java.nio.file.PathMatcher nameMatcher = + glob == null || glob.isBlank() ? null : start.getFileSystem().getPathMatcher("glob:" + glob); + + List hits = new ArrayList<>(); + Set matchedFiles = new java.util.LinkedHashSet<>(); + long deadline = System.currentTimeMillis() + DEADLINE_MILLIS; + boolean[] truncated = {false}; + try { + Files.walkFileTree(start, Set.of(), Integer.MAX_VALUE, new SimpleFileVisitor<>() { + @Override + public FileVisitResult preVisitDirectory(Path directory, BasicFileAttributes attributes) { + String name = directory.getFileName() == null + ? "" + : directory.getFileName().toString(); + return EXCLUDED_DIRECTORIES.contains(name) || (!directory.equals(start) && name.startsWith(".")) + ? FileVisitResult.SKIP_SUBTREE + : FileVisitResult.CONTINUE; + } + + @Override + public FileVisitResult visitFile(Path file, BasicFileAttributes attributes) { + if (hits.size() >= MAX_TOTAL_MATCHES || System.currentTimeMillis() > deadline) { + truncated[0] = true; + return FileVisitResult.TERMINATE; + } + // Use the attributes the walk already has: a separate isRegularFile/size call per + // file goes through CreateFileW on Windows and dominates the walk. + if (!attributes.isRegularFile() || attributes.size() > MAX_FILE_BYTES) { + return FileVisitResult.CONTINUE; + } + if (nameMatcher != null && !nameMatcher.matches(file.getFileName())) { + return FileVisitResult.CONTINUE; + } + searchFile(root, file, regex, hits, matchedFiles, truncated); + return FileVisitResult.CONTINUE; + } + + @Override + public FileVisitResult visitFileFailed(Path file, IOException e) { + return FileVisitResult.CONTINUE; + } + }); + } catch (IOException e) { + throw new IllegalArgumentException("Search failed: " + e.getMessage(), e); + } + return new Result(List.copyOf(hits), matchedFiles.size(), truncated[0]); + } + + private static void searchFile( + Path root, Path file, Pattern regex, List hits, Set matchedFiles, boolean[] truncated) { + List lines; + try { + // Binary content fails to decode as UTF-8 and is skipped -- the cheap equivalent of + // ripgrep's "a file with a NUL byte is binary". + lines = Files.readAllLines(file, StandardCharsets.UTF_8); + } catch (IOException | RuntimeException e) { + return; + } + String relative = root.relativize(file).toString().replace('\\', '/'); + int inThisFile = 0; + for (int i = 0; i < lines.size(); i++) { + if (hits.size() >= MAX_TOTAL_MATCHES) { + truncated[0] = true; + return; + } + Matcher matcher = regex.matcher(lines.get(i)); + if (!matcher.find()) { + continue; + } + matchedFiles.add(relative); + inThisFile++; + if (inThisFile > MAX_MATCHES_PER_FILE) { + truncated[0] = true; + return; + } + String text = lines.get(i); + hits.add(new Hit( + relative, + i + 1, + text.length() <= MAX_LINE_CHARS ? text : text.substring(0, MAX_LINE_CHARS) + " …[cut]")); + } + } + + /** + * Render a result for the model: grouped by file, line-numbered, and honest about truncation. + * + * @param result the search result + * @param filesOnly whether to list only the file names + * @return the text the tool returns + */ + public static String format(Result result, boolean filesOnly) { + if (result.hits().isEmpty()) { + return "(no matches)"; + } + Map> byFile = new LinkedHashMap<>(); + for (Hit hit : result.hits()) { + byFile.computeIfAbsent(hit.path(), key -> new ArrayList<>()).add(hit); + } + StringBuilder text = new StringBuilder(); + if (filesOnly) { + byFile.keySet().forEach(path -> text.append(path).append(System.lineSeparator())); + } else { + for (Map.Entry> file : byFile.entrySet()) { + text.append(file.getKey()).append(System.lineSeparator()); + for (Hit hit : file.getValue()) { + text.append(" ") + .append(hit.line()) + .append(": ") + .append(hit.text()) + .append(System.lineSeparator()); + } + text.append(System.lineSeparator()); + } + } + text.append(result.hits().size()) + .append(" matches in ") + .append(result.filesWithMatches()) + .append(" files"); + if (result.truncated()) { + text.append(" (TRUNCATED — there are more; narrow the search with `glob`, `dir`" + + " or a more specific pattern)"); + } + return text.toString(); + } +} diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/WorkspaceTools.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/WorkspaceTools.java new file mode 100644 index 000000000..36970ff7d --- /dev/null +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/WorkspaceTools.java @@ -0,0 +1,398 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import java.nio.file.Path; +import java.util.ArrayList; +import java.util.LinkedHashSet; +import java.util.List; +import java.util.Map; +import java.util.Set; +import org.atmosphere.ai.fs.AgentFileSystem; +import org.atmosphere.ai.fs.FileSystemTools; +import org.atmosphere.ai.fs.WorkspaceAgentFileSystem; +import org.atmosphere.ai.tool.ToolDefinition; +import org.atmosphere.ai.tool.ToolExecutor; +import org.atmosphere.ai.tool.ToolKind; +import org.atmosphere.ai.tool.ToolParameter; +import org.jspecify.annotations.Nullable; + +/** + * The three tools this agent provides itself, replacing the framework's versions of the same names: + * {@code read_file}, {@code edit_file} and {@code grep}. Everything else — {@code ls}, + * {@code write_file}, {@code glob}, {@code delete}, {@code rename} — stays Atmosphere's. + * + *

They are replacements, not additions, on purpose: two tools that both claim to "read a file" is + * the worst outcome for tool selection. Underneath they use the same {@link AgentFileSystem} the + * framework binds to the session, so path validation, the workspace confinement and the size limits + * are unchanged; only the parameters the model sees and the text it gets back are ours. + * + *

Why each one differs from the framework's: + * + *

    + *
  • read_file takes {@code offset}/{@code limit} and prints line numbers. Reading a whole + * file costs context for nothing and measurably lowers task success (SWE-agent: 12.7 % with + * whole files against 18.0 % with a 100-line window). + *
  • edit_file normalizes line endings (the framework compares raw content, so no edit ever + * matches in a CRLF file), explains a miss with the nearest lines, names the line numbers of an + * ambiguous match, and can apply several edits at once — see {@link TextEdits}. + *
  • grep skips {@code .git}, build output and dependency directories — the framework's walk + * is alphabetical with global budgets, so those directories eat the result before the source is + * reached — and states when it truncated; see {@link WorkspaceSearch}. + *
+ */ +public final class WorkspaceTools { + + /** Lines returned by {@code read_file} when the model gives no limit. */ + public static final int DEFAULT_READ_LIMIT = 400; + + /** Lines of context shown around a completed edit. */ + private static final int EDIT_SNIPPET_CONTEXT = 3; + + private WorkspaceTools() {} + + /** + * Remembers which files were read, so an edit can insist on it. + * + *

Not a staleness check: an {@code old_string} that matches the current content exactly and + * unambiguously is safe whether or not the file changed meanwhile. What this prevents is the + * other case — a model inventing the text it wants to replace. One instance per session. + */ + public static final class ReadTracker { + private final Set read = new LinkedHashSet<>(); + + /** + * Note that a file was read. + * + * @param path the path as the model wrote it + */ + public void markRead(String path) { + read.add(normalizePath(path)); + } + + /** + * Whether a file was read in this session. + * + * @param path the path as the model wrote it + * @return {@code true} when it was read before + */ + public boolean wasRead(String path) { + return read.contains(normalizePath(path)); + } + + private static String normalizePath(String path) { + return path.replace('\\', '/').replaceAll("^\\./", ""); + } + } + + /** + * The tool set: Atmosphere's tools where they are fine, ours where they are not. + * + * @param tracker the read tracker shared by {@code read_file} and {@code edit_file} + * @return the tools to offer the model + */ + public static List all(ReadTracker tracker) { + return List.of( + FileSystemTools.ls(), + readFile(tracker), + FileSystemTools.writeFile(), + editFile(tracker), + FileSystemTools.glob(), + grep(), + FileSystemTools.delete(), + FileSystemTools.rename()); + } + + /** + * {@code read_file}: a window of a file, with line numbers. + * + * @param tracker records the read so an edit is allowed afterwards + * @return the tool + */ + public static ToolDefinition readFile(ReadTracker tracker) { + return ToolDefinition.builder( + FileSystemTools.READ_FILE, + "Read a file from the workspace. Returns the lines numbered as ` 12: text`." + + " Reads at most " + DEFAULT_READ_LIMIT + " lines at a time; use offset and limit to" + + " page through a longer file. The line numbers are display only — never include them" + + " in old_string when you edit.") + .parameter("file_path", "File path relative to the workspace root", "string", true) + .parameter("offset", "First line to show, 1-based (default 1)", "integer", false) + .parameter("limit", "How many lines to show (default " + DEFAULT_READ_LIMIT + ")", "integer", false) + .returnType("string") + .executor(withFileSystem((arguments, injectables) -> { + AgentFileSystem fs = fileSystem(injectables); + if (fs == null) { + return "File tools unavailable: no agent filesystem is bound to this session."; + } + String path = string(arguments, "file_path"); + if (path == null) { + return "Error: file_path is required"; + } + try { + List lines = fs.read(path).lines().toList(); + int offset = Math.max(1, integer(arguments, "offset", 1)); + int limit = Math.max(1, integer(arguments, "limit", DEFAULT_READ_LIMIT)); + tracker.markRead(path); + return window(lines, offset, limit); + } catch (IllegalArgumentException e) { + return "Error: " + e.getMessage(); + } + })) + .kind(ToolKind.READ) + .build(); + } + + /** + * Render the requested window with line numbers and say what was left out. + * + * @param lines the file's lines + * @param offset the first line, 1-based + * @param limit how many lines + * @return the text for the model + */ + static String window(List lines, int offset, int limit) { + if (lines.isEmpty()) { + return "(empty file)"; + } + if (offset > lines.size()) { + return "(offset " + offset + " is past the end; the file has " + lines.size() + " lines)"; + } + int last = Math.min(lines.size(), offset + limit - 1); + StringBuilder text = new StringBuilder(); + for (int i = offset; i <= last; i++) { + text.append(" ").append(i).append(": ").append(lines.get(i - 1)).append(System.lineSeparator()); + } + if (offset > 1 || last < lines.size()) { + text.append("(showing lines ") + .append(offset) + .append("–") + .append(last) + .append(" of ") + .append(lines.size()) + .append("; use offset/limit for the rest)"); + } + return text.toString(); + } + + /** + * {@code edit_file}: exact replacement with line-ending handling, corrective errors, and several + * edits in one call. + * + * @param tracker enforces that the file was read first + * @return the tool + */ + public static ToolDefinition editFile(ReadTracker tracker) { + ToolParameter edits = ToolParameter.ofArray( + "edits", + "Several replacements applied to the same file, in order. Use instead of" + + " old_string/new_string. All of them must match, or the file is left untouched.", + false, + new ToolParameter( + "edit", + "One replacement", + "object", + false, + List.of(), + null, + List.of( + new ToolParameter("old_string", "The exact text to replace", "string", true), + new ToolParameter("new_string", "The replacement", "string", true), + new ToolParameter("replace_all", "Replace every occurrence", "boolean", false)))); + return ToolDefinition.builder( + FileSystemTools.EDIT_FILE, + "Edit a file by replacing exact text. Read the file first and copy old_string from it" + + " (without the line numbers the read tool prints). old_string must match exactly" + + " once unless replace_all is set. Line endings are handled for you.") + .parameter("file_path", "File path relative to the workspace root", "string", true) + .parameter("old_string", "The exact text to replace", "string", false) + .parameter("new_string", "The replacement text", "string", false) + .parameter("replace_all", "Replace every occurrence instead of requiring exactly one", "boolean", false) + .parameter(edits) + .returnType("string") + .executor(withFileSystem((arguments, injectables) -> { + AgentFileSystem fs = fileSystem(injectables); + if (fs == null) { + return "File tools unavailable: no agent filesystem is bound to this session."; + } + String path = string(arguments, "file_path"); + if (path == null) { + return "Error: file_path is required"; + } + List list; + try { + list = parseEdits(arguments); + } catch (IllegalArgumentException e) { + return "Error: " + e.getMessage(); + } + if (!tracker.wasRead(path)) { + return "Error: read " + path + " before editing it, so old_string comes from the file" + + " rather than from memory."; + } + try { + String original = fs.read(path); + String edited = TextEdits.apply(original, list); + fs.write(path, edited); + return "Edited " + path + " (" + list.size() + (list.size() == 1 ? " edit)" : " edits)") + + describe(edited, list); + } catch (TextEdits.EditException e) { + return "Error: " + e.getMessage(); + } catch (IllegalArgumentException e) { + return "Error: " + e.getMessage(); + } + })) + .kind(ToolKind.EDIT) + .build(); + } + + private static String describe(String edited, List edits) { + TextEdits.Edit last = edits.get(edits.size() - 1); + String snippet = TextEdits.snippet( + TextEdits.normalize(edited), last.oldString(), last.newString(), EDIT_SNIPPET_CONTEXT); + return snippet == null ? "" : System.lineSeparator() + snippet; + } + + /** + * Read the edits from the call, in either shape. + * + * @param arguments the tool arguments + * @return the edits, in order + * @throws IllegalArgumentException when neither shape is present or an entry is incomplete + */ + static List parseEdits(Map arguments) { + Object many = arguments == null ? null : arguments.get("edits"); + if (many instanceof Iterable entries) { + List edits = new ArrayList<>(); + for (Object entry : entries) { + if (!(entry instanceof Map map)) { + throw new IllegalArgumentException( + "every entry of edits must be an object with old_string" + " and new_string"); + } + Object oldString = map.get("old_string"); + Object newString = map.get("new_string"); + if (oldString == null) { + throw new IllegalArgumentException("every entry of edits needs old_string"); + } + edits.add(new TextEdits.Edit( + oldString.toString(), + newString == null ? "" : newString.toString(), + Boolean.TRUE.equals(map.get("replace_all")) + || "true".equalsIgnoreCase(String.valueOf(map.get("replace_all"))))); + } + if (edits.isEmpty()) { + throw new IllegalArgumentException("edits was empty"); + } + return edits; + } + String oldString = string(arguments, "old_string"); + if (oldString == null) { + throw new IllegalArgumentException("old_string is required (or pass edits)"); + } + String newString = string(arguments, "new_string"); + return List.of( + new TextEdits.Edit(oldString, newString == null ? "" : newString, bool(arguments, "replace_all"))); + } + + /** + * {@code grep}: search the workspace, skipping what is not source. + * + * @return the tool + */ + public static ToolDefinition grep() { + return ToolDefinition.builder( + FileSystemTools.GREP, + "Search the workspace with a regular expression. Returns matching lines grouped by file" + + " as ` 12: text`, capped at " + WorkspaceSearch.MAX_TOTAL_MATCHES + " matches;" + + " .git, build output and dependency directories are skipped. Says so when the" + + " result is truncated.") + .parameter("pattern", "The regular expression to search for", "string", true) + .parameter("dir", "Directory to search under, relative to the workspace root", "string", false) + .parameter("glob", "Only search files matching this name pattern, e.g. *.java", "string", false) + .parameter("files_only", "Return just the file names instead of the matching lines", "boolean", false) + .returnType("string") + .executor(withFileSystem((arguments, injectables) -> { + AgentFileSystem fs = fileSystem(injectables); + if (!(fs instanceof WorkspaceAgentFileSystem workspace)) { + return "Search unavailable: this session has no workspace directory."; + } + String pattern = string(arguments, "pattern"); + if (pattern == null) { + return "Error: pattern is required"; + } + try { + Path root = workspace.root(); + WorkspaceSearch.Result result = WorkspaceSearch.search( + root, pattern, string(arguments, "dir"), string(arguments, "glob")); + return WorkspaceSearch.format(result, bool(arguments, "files_only")); + } catch (IllegalArgumentException e) { + return "Error: " + e.getMessage(); + } + })) + .kind(ToolKind.READ) + .build(); + } + + /** + * Adapt a two-argument function to {@link ToolExecutor}, whose single abstract method takes only + * the arguments — the injectables arrive through its default overload, and that is where the + * session's filesystem lives. + * + * @param body what the tool does + * @return the executor + */ + private static ToolExecutor withFileSystem(ToolBody body) { + return new ToolExecutor() { + @Override + public Object execute(Map arguments) { + return execute(arguments, Map.of()); + } + + @Override + public Object execute(Map arguments, Map, Object> injectables) { + return body.run(arguments, injectables); + } + }; + } + + /** The body of a tool: arguments plus the session's injectables. */ + @FunctionalInterface + private interface ToolBody { + /** + * Run the tool. + * + * @param arguments the call's arguments + * @param injectables the session scope, carrying the filesystem + * @return what the model sees + */ + Object run(Map arguments, Map, Object> injectables); + } + + private static @Nullable AgentFileSystem fileSystem(@Nullable Map, Object> injectables) { + return FileSystemTools.resolveFileSystem(injectables == null ? Map.of() : injectables) + .orElse(null); + } + + private static @Nullable String string(@Nullable Map arguments, String name) { + Object value = arguments == null ? null : arguments.get(name); + return value == null || value.toString().isEmpty() ? null : value.toString(); + } + + private static boolean bool(@Nullable Map arguments, String name) { + Object value = arguments == null ? null : arguments.get(name); + return Boolean.TRUE.equals(value) || "true".equalsIgnoreCase(String.valueOf(value)); + } + + private static int integer(@Nullable Map arguments, String name, int fallback) { + Object value = arguments == null ? null : arguments.get(name); + if (value instanceof Number number) { + return number.intValue(); + } + try { + return value == null ? fallback : Integer.parseInt(value.toString().trim()); + } catch (NumberFormatException e) { + return fallback; + } + } +} diff --git a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/system-prompt.txt b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/system-prompt.txt index 1cc6d722c..43d1afa39 100644 --- a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/system-prompt.txt +++ b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/system-prompt.txt @@ -4,4 +4,4 @@ The file tools ls, read_file, write_file, edit_file, glob, grep, delete and rena {shell_section} -Work step by step: read a file before you edit it, check the result after a change, and finish with a short summary. Answer in the user's language. +Work step by step. read_file shows a numbered window of a file — those numbers are display only, never copy them into old_string. Edit a file only after reading it (edit_file refuses otherwise), copy old_string from what you read, and check the result after a change. Finish with a short summary. Answer in the user's language. diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java index 7e80f3d4b..982fe8782 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java @@ -53,7 +53,7 @@ private static OpenAiCompatServer server(ScriptedBackend backend) throws Excepti void oneShotTurnReadsAWorkspaceFileThroughTheBuiltInFileTools() throws Exception { Files.writeString(workspace.resolve("hello.txt"), "VALUE=42\n"); ScriptedBackend backend = new ScriptedBackend((call, request) -> call == 1 - ? ScriptedBackend.toolCallTurn("call_1", "read_file", "{\"path\":\"hello.txt\"}") + ? ScriptedBackend.toolCallTurn("call_1", "read_file", "{\"file_path\":\"hello.txt\"}") : ScriptedBackend.textTurn("The file says VALUE=42.")); ByteArrayOutputStream out = new ByteArrayOutputStream(); ByteArrayOutputStream err = new ByteArrayOutputStream(); @@ -77,8 +77,9 @@ void oneShotTurnReadsAWorkspaceFileThroughTheBuiltInFileTools() throws Exception } String console = out.toString(StandardCharsets.UTF_8); // the tool line and its result, as ConsoleSession renders them (unstyled here: not a terminal) - assertThat(console, containsString("● read_file {path=hello.txt}")); - assertThat(console, containsString("↳ VALUE=42")); + assertThat(console, containsString("● read_file {file_path=hello.txt}")); + // our read_file numbers the lines, so the result is " 1: VALUE=42" + assertThat(console, containsString("↳ 1: VALUE=42")); assertThat(console, containsString("The file says VALUE=42.")); List requests = backend.requests(); assertThat(requests, hasSize(2)); diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/TextEditsTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/TextEditsTest.java new file mode 100644 index 000000000..2ab35252f --- /dev/null +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/TextEditsTest.java @@ -0,0 +1,138 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import static org.hamcrest.MatcherAssert.assertThat; +import static org.hamcrest.Matchers.containsString; +import static org.hamcrest.Matchers.is; +import static org.hamcrest.Matchers.not; +import static org.junit.jupiter.api.Assertions.assertThrows; + +import java.util.List; +import org.junit.jupiter.api.Test; + +/** The matching and rewriting rules that decide whether an edit lands. */ +class TextEditsTest { + + private static TextEdits.Edit edit(String oldString, String newString) { + return new TextEdits.Edit(oldString, newString, false); + } + + @Test + void aCrlfFileIsEditedByAModelThatOnlyWritesLineFeeds() { + // THE bug this class exists for: the framework compares raw content, so on Windows the + // model's "a\nb" never matches the file's "a\r\nb" and every edit silently fails. + String file = "int a = 1;\r\nint b = 2;\r\n"; + + String edited = TextEdits.apply(file, List.of(edit("int a = 1;\nint b = 2;", "int a = 3;\nint b = 4;"))); + + assertThat(edited, is("int a = 3;\r\nint b = 4;\r\n")); + } + + @Test + void theFilesOwnLineEndingAndByteOrderMarkSurvive() { + assertThat(TextEdits.apply("alpha\r\nbeta\r\n", List.of(edit("beta", "gamma"))), is("alpha\r\ngamma\r\n")); + assertThat(TextEdits.apply("alpha\nbeta\n", List.of(edit("beta", "gamma"))), is("alpha\ngamma\n")); + assertThat(TextEdits.lineEnding("a\r\nb"), is("\r\n")); + assertThat(TextEdits.lineEnding("a\nb"), is("\n")); + assertThat(TextEdits.lineEnding("no newline at all"), is("\n")); + } + + @Test + void anAmbiguousMatchNamesTheLineNumbers() { + String file = "x();\ny();\nx();\n"; + + TextEdits.EditException error = + assertThrows(TextEdits.EditException.class, () -> TextEdits.apply(file, List.of(edit("x();", "z();")))); + + assertThat(error.getMessage(), containsString("occurs 2 times")); + assertThat( + "naming the lines is what makes the next attempt possible", + error.getMessage(), + containsString("lines 1, 3")); + assertThat(error.getMessage(), containsString("replace_all")); + } + + @Test + void replaceAllTakesEveryOccurrence() { + assertThat( + TextEdits.apply("x();\ny();\nx();\n", List.of(new TextEdits.Edit("x();", "z();", true))), + is("z();\ny();\nz();\n")); + } + + @Test + void aMissShowsTheNearestLinesInsteadOfJustSayingNo() { + String file = "public void handle(String name) {\n log(name);\n}\n"; + + TextEdits.EditException error = assertThrows( + TextEdits.EditException.class, + () -> TextEdits.apply(file, List.of(edit("public void handle(String value) {", "x")))); + + assertThat(error.getMessage(), containsString("not found")); + assertThat(error.getMessage(), containsString("closest lines")); + assertThat(error.getMessage(), containsString("1: public void handle(String name) {")); + } + + @Test + void aMissWithNothingSimilarSaysToReadTheFileAgain() { + TextEdits.EditException error = assertThrows( + TextEdits.EditException.class, + () -> TextEdits.apply("alpha\nbeta\n", List.of(edit("zzzzzzzzzzzzzzzz", "x")))); + + assertThat(error.getMessage(), containsString("read the file again")); + // and it warns about the trap that causes many of these misses + assertThat(error.getMessage(), containsString("line numbers")); + } + + @Test + void severalEditsAreAllAppliedOrNoneAreAtAll() { + String file = "one\ntwo\nthree\n"; + + String edited = TextEdits.apply(file, List.of(edit("one", "1"), edit("three", "3"))); + assertThat(edited, is("1\ntwo\n3\n")); + + TextEdits.EditException error = assertThrows( + TextEdits.EditException.class, + () -> TextEdits.apply(file, List.of(edit("one", "1"), edit("absent", "x")))); + assertThat(error.getMessage(), containsString("edit 2 of 2")); + } + + @Test + void anEmptyOrUnchangedEditIsRejectedRatherThanSilentlyDoingNothing() { + assertThat( + assertThrows(TextEdits.EditException.class, () -> TextEdits.apply("a\n", List.of(edit("", "x")))) + .getMessage(), + containsString("must not be empty")); + assertThat( + assertThrows(TextEdits.EditException.class, () -> TextEdits.apply("a\n", List.of(edit("a", "a")))) + .getMessage(), + containsString("identical")); + assertThat( + assertThrows(TextEdits.EditException.class, () -> TextEdits.apply("a\n", List.of())) + .getMessage(), + containsString("No edits")); + } + + @Test + void theAnswerShowsTheEditedRegionWithLineNumbers() { + String edited = TextEdits.apply("a\nb\nc\nd\ne\n", List.of(edit("c", "CHANGED"))); + + String snippet = TextEdits.snippet(TextEdits.normalize(edited), "c", "CHANGED", 1); + + assertThat(snippet, containsString("3: CHANGED")); + assertThat(snippet, containsString("2: b")); + assertThat("the whole file is never echoed back", snippet, not(containsString("5: e"))); + } + + @Test + void similarityIsBetweenZeroAndOneAndOrdersTheCandidates() { + assertThat(TextEdits.similarity("abc", "abc"), is(1.0)); + assertThat(TextEdits.similarity("", "abc"), is(0.0)); + assertThat(TextEdits.similarity("int a = 1;", "int a = 2;") > TextEdits.CANDIDATE_MIN_SIMILARITY, is(true)); + assertThat( + TextEdits.similarity("int a = 1;", "completely different") < TextEdits.CANDIDATE_MIN_SIMILARITY, + is(true)); + } +} diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/WorkspaceToolsTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/WorkspaceToolsTest.java new file mode 100644 index 000000000..2bf03d4dd --- /dev/null +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/WorkspaceToolsTest.java @@ -0,0 +1,205 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import static org.hamcrest.MatcherAssert.assertThat; +import static org.hamcrest.Matchers.contains; +import static org.hamcrest.Matchers.containsString; +import static org.hamcrest.Matchers.is; +import static org.hamcrest.Matchers.not; + +import java.nio.charset.StandardCharsets; +import java.nio.file.Files; +import java.nio.file.Path; +import java.util.List; +import java.util.Map; +import org.atmosphere.ai.fs.AgentFileSystem; +import org.atmosphere.ai.fs.WorkspaceAgentFileSystem; +import org.atmosphere.ai.tool.ToolDefinition; +import org.junit.jupiter.api.BeforeEach; +import org.junit.jupiter.api.Test; +import org.junit.jupiter.api.io.TempDir; + +/** The three replaced tools, driven the way Atmosphere drives them. */ +class WorkspaceToolsTest { + + @TempDir + Path workspace; + + private Map, Object> scope; + private WorkspaceTools.ReadTracker tracker; + + @BeforeEach + void bindFileSystem() { + AgentFileSystem fs = new WorkspaceAgentFileSystem(workspace, AgentFileSystem.Limits.defaults()); + scope = Map.of(AgentFileSystem.class, fs); + tracker = new WorkspaceTools.ReadTracker(); + } + + private String call(ToolDefinition tool, Map arguments) throws Exception { + return String.valueOf(tool.executor().execute(arguments, scope)); + } + + private void write(String name, String content) throws Exception { + Files.writeString(workspace.resolve(name), content, StandardCharsets.UTF_8); + } + + // ----- read_file ----- + + @Test + void readingReturnsNumberedLinesAndSaysWhatItLeftOut() throws Exception { + write("big.txt", "l1\nl2\nl3\nl4\nl5\n"); + + String all = call(WorkspaceTools.readFile(tracker), Map.of("file_path", "big.txt")); + assertThat(all, containsString(" 1: l1")); + assertThat("a file that fits is shown without a footer", all, not(containsString("showing lines"))); + + String window = call(WorkspaceTools.readFile(tracker), Map.of("file_path", "big.txt", "offset", 2, "limit", 2)); + assertThat(window, containsString(" 2: l2")); + assertThat(window, containsString(" 3: l3")); + assertThat(window, not(containsString("l4"))); + assertThat(window, containsString("showing lines 2–3 of 5")); + } + + @Test + void readingPastTheEndAndReadingNothingAreExplained() throws Exception { + write("small.txt", "only\n"); + write("empty.txt", ""); + + assertThat( + call(WorkspaceTools.readFile(tracker), Map.of("file_path", "small.txt", "offset", 99)), + containsString("past the end")); + assertThat(call(WorkspaceTools.readFile(tracker), Map.of("file_path", "empty.txt")), containsString("empty")); + assertThat(call(WorkspaceTools.readFile(tracker), Map.of()), containsString("file_path is required")); + } + + // ----- edit_file ----- + + @Test + void anEditIsRefusedUntilTheFileWasRead() throws Exception { + write("code.txt", "alpha\n"); + + String refused = call( + WorkspaceTools.editFile(tracker), + Map.of("file_path", "code.txt", "old_string", "alpha", "new_string", "beta")); + assertThat(refused, containsString("read code.txt before editing")); + assertThat("nothing was written", Files.readString(workspace.resolve("code.txt")), is("alpha\n")); + + call(WorkspaceTools.readFile(tracker), Map.of("file_path", "code.txt")); + String edited = call( + WorkspaceTools.editFile(tracker), + Map.of("file_path", "code.txt", "old_string", "alpha", "new_string", "beta")); + assertThat(edited, containsString("Edited code.txt")); + assertThat(Files.readString(workspace.resolve("code.txt")), is("beta\n")); + } + + @Test + void theAnswerOfASuccessfulEditShowsTheChangedRegion() throws Exception { + write("code.txt", "a\nb\nc\nd\n"); + call(WorkspaceTools.readFile(tracker), Map.of("file_path", "code.txt")); + + String answer = call( + WorkspaceTools.editFile(tracker), + Map.of("file_path", "code.txt", "old_string", "c", "new_string", "CHANGED")); + + assertThat(answer, containsString("3: CHANGED")); + } + + @Test + void severalEditsArriveAsAListAndAreAllOrNothing() throws Exception { + write("code.txt", "one\ntwo\nthree\n"); + call(WorkspaceTools.readFile(tracker), Map.of("file_path", "code.txt")); + + String failed = call( + WorkspaceTools.editFile(tracker), + Map.of( + "file_path", + "code.txt", + "edits", + List.of( + Map.of("old_string", "one", "new_string", "1"), + Map.of("old_string", "absent", "new_string", "x")))); + assertThat(failed, containsString("edit 2 of 2")); + assertThat( + "a failed batch leaves the file untouched", + Files.readString(workspace.resolve("code.txt")), + is("one\ntwo\nthree\n")); + + String applied = call( + WorkspaceTools.editFile(tracker), + Map.of( + "file_path", + "code.txt", + "edits", + List.of( + Map.of("old_string", "one", "new_string", "1"), + Map.of("old_string", "three", "new_string", "3")))); + assertThat(applied, containsString("2 edits")); + assertThat(Files.readString(workspace.resolve("code.txt")), is("1\ntwo\n3\n")); + } + + @Test + void aFailedEditExplainsItselfInsteadOfSayingNo() throws Exception { + write("code.txt", "public void handle(String name) {\n}\n"); + call(WorkspaceTools.readFile(tracker), Map.of("file_path", "code.txt")); + + String answer = call( + WorkspaceTools.editFile(tracker), + Map.of("file_path", "code.txt", "old_string", "public void handle(String value) {", "new_string", "x")); + + assertThat(answer, containsString("closest lines")); + assertThat(answer, containsString("1: public void handle(String name) {")); + } + + // ----- grep ----- + + @Test + void theSearchSkipsBuildOutputAndRepositoryInternals() throws Exception { + Files.createDirectories(workspace.resolve("src")); + Files.createDirectories(workspace.resolve("target/classes")); + Files.createDirectories(workspace.resolve(".git")); + Files.writeString(workspace.resolve("src/Main.java"), "class Main { void needle() {} }\n"); + Files.writeString(workspace.resolve("target/classes/Main.txt"), "needle in build output\n"); + Files.writeString(workspace.resolve(".git/config"), "needle in the repository\n"); + + String answer = call(WorkspaceTools.grep(), Map.of("pattern", "needle")); + + assertThat(answer, containsString("src/Main.java")); + assertThat("build output is not source", answer, not(containsString("target/"))); + assertThat("the repository's own files are not source", answer, not(containsString(".git"))); + assertThat(answer, containsString("1 matches in 1 files")); + } + + @Test + void theSearchCanBeNarrowedAndCanListFilesOnly() throws Exception { + Files.writeString(workspace.resolve("A.java"), "needle\n"); + Files.writeString(workspace.resolve("B.txt"), "needle\n"); + + assertThat( + call(WorkspaceTools.grep(), Map.of("pattern", "needle", "glob", "*.java")), + not(containsString("B.txt"))); + String filesOnly = call(WorkspaceTools.grep(), Map.of("pattern", "needle", "files_only", true)); + assertThat(filesOnly, containsString("A.java")); + assertThat("files_only means no line content", filesOnly, not(containsString(" 1: needle"))); + } + + @Test + void anInvalidPatternIsAnErrorMessageNotAnException() throws Exception { + assertThat(call(WorkspaceTools.grep(), Map.of("pattern", "[unclosed")), containsString("Invalid regular")); + assertThat(call(WorkspaceTools.grep(), Map.of()), containsString("pattern is required")); + } + + @Test + void theToolSetReplacesTheFrameworksReadEditAndGrepWithoutDuplicates() { + List names = + WorkspaceTools.all(tracker).stream().map(ToolDefinition::name).toList(); + + assertThat(names, contains("ls", "read_file", "write_file", "edit_file", "glob", "grep", "delete", "rename")); + assertThat( + "no tool name may appear twice", + names.size(), + is(names.stream().distinct().count() == 8L ? 8 : -1)); + } +} From 356fef8afd53d12643d57035f3f91367e43edf81 Mon Sep 17 00:00:00 2001 From: Bernard Ladenthin Date: Wed, 23 Sep 2026 10:46:43 +0200 Subject: [PATCH 06/36] llama-atmosphere-agent: keep tool calls in the history, add /calls A real session showed the failure this fixes: after three turns a 4B model stopped calling tools and started describing the work instead -- it reported JUnit tests, a Maven build and a .bat script it had run, with exit codes, while the workspace contained only the two files from turn one. Nothing had been executed, so nothing asked for approval either. The cause is on our side: the history held the user's messages and the model's prose and nothing else, so from turn three the model saw only its own paragraphs and no evidence it had ever used a tool -- and continued that pattern. - LocalAgent.withToolNotes prefixes each turn's answer in the history with "(tools I actually ran this turn: -> )". A text note rather than real tool_calls messages because Atmosphere's AbstractAgentRuntime.assembleMessages rebuilds every history entry as new ChatMessage(role, content) -- the tool-call array and id are dropped, so protocol-faithful replay is impossible through the framework's history. - ToolCallLog + /calls (/log): every call of the session with its result. An empty list is the proof that nothing ran, whatever the prose says. - System prompt: never claim to have run a command, created a file or seen a result without having called the tool. Tests: 102 green. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01E2h8gXyyE5UeimkL9vQv9G --- CLAUDE.md | 14 ++- llama-atmosphere-agent/README.md | 16 +++ .../llama/atmosphere/ConsoleSession.java | 33 ++++++ .../llama/atmosphere/LocalAgent.java | 80 ++++++++++++-- .../llama/atmosphere/SlashCommands.java | 2 + .../ladenthin/llama/atmosphere/TaskLoop.java | 12 ++- .../llama/atmosphere/ToolCallLog.java | 100 ++++++++++++++++++ .../net/ladenthin/llama/atmosphere/help.txt | 1 + .../llama/atmosphere/system-prompt.txt | 2 +- .../llama/atmosphere/LocalAgentTest.java | 45 ++++++++ .../llama/atmosphere/TaskLoopTest.java | 2 +- 11 files changed, 292 insertions(+), 15 deletions(-) create mode 100644 llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ToolCallLog.java diff --git a/CLAUDE.md b/CLAUDE.md index f110ff36f..4f53ef8e8 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -2303,7 +2303,19 @@ are decisions, not details: an embedding index (Cursor's production effect is +0.3 %), and LSP tools (the one isolation study finds them token-negative and *worse* at multi-file rename, because renames touch comments and strings that semantic references exclude). -7. **The context number in the status line is an estimate, marked `~`.** llama.cpp emits its usage +7. **Tool calls are carried into the conversation as a text note, and logged for `/calls`.** + `LocalAgent.withToolNotes` prefixes each turn's answer in the history with + `(tools I actually ran this turn: -> )`, and `ToolCallLog` + keeps the same data for the `/calls` command. **Why it is a note and not real `tool_calls` + messages:** `AbstractAgentRuntime.assembleMessages` rebuilds every history entry as + `new ChatMessage(h.role(), h.content())` — the tool-call array and the tool-call id never leave the + framework, so protocol-faithful replay through `context.history()` is impossible; content is what + survives. **Why it exists at all:** with only user text and assistant prose in the history, a 4B + model stopped calling tools after the third turn of a real session and *described* the work instead + — inventing JUnit tests, a Maven build and a `.bat` script, complete with exit codes, while the + workspace stayed empty. The system prompt also forbids claiming an action without the call. + Pinned by `LocalAgentTest.aToolCallStaysInTheHistorySoTheNextTurnSeesItHappened`. +8. **The context number in the status line is an estimate, marked `~`.** llama.cpp emits its usage chunk only when the client sets `stream_options.include_usage`, and Atmosphere's client does not; `ConsoleSession.usage()` takes the real count when one arrives, otherwise `LocalAgent.estimateTokens` uses four characters per token. The window size is `--ctx-size` (in-process) or the server's diff --git a/llama-atmosphere-agent/README.md b/llama-atmosphere-agent/README.md index 8b0a46f99..03a5628cc 100644 --- a/llama-atmosphere-agent/README.md +++ b/llama-atmosphere-agent/README.md @@ -184,6 +184,7 @@ unknown `/command` included — goes to the model: | `/help` (`/?`, `/commands`) | the overview below | | `/status` | mode, context use, tools, model, workspace, history size | | `/tools` | the tools offered, and which of them ask first | +| `/calls` (`/log`) | every tool call of this session with its result — the receipt | | `/mode [manual\|auto]` (`/approve`) | show or set the approval mode | | `/compact [focus]` | summarize the conversation and continue from the summary | | `/loop [--every 5m] [--max 20] [--check ''] ` | keep working on one task until it is done | @@ -214,6 +215,19 @@ can answer, so a gated call is denied** — pass `--auto` to run unattended. The Atmosphere's (`ToolApprovalPolicy` + `ApprovalStrategy`); the agent only supplies the question and the answer. +**`/calls` is the receipt.** It lists every tool call of the session, one line each, with the +arguments and a short result. Use it when an answer sounds too good: a model that has drifted starts +*describing* work — "the tests passed, the jar was created" — while calling nothing at all. The +scrollback reads the same either way; this list only grows when something really ran. + +To make that drift less likely, each turn's calls are also carried into the conversation as a short +note above the answer (`(tools I actually ran this turn: …)`). Without it the history holds only the +user's messages and the model's own prose, and a small model then continues the pattern it sees — +prose. This is not theoretical: in a real session a 4B model invented JUnit tests, a Maven build and a +`.bat` script it had never written, three turns in a row. The note is plain text rather than proper +`tool_calls` messages because Atmosphere's `assembleMessages` rebuilds every history entry as +`new ChatMessage(role, content)` and drops the rest; the content is what survives. + **`/compact`** asks the model to summarize the conversation (goal, facts, work done, problems, state, next step; `/compact ` adds an emphasis), then replaces the history with that summary. Use it when the context fills up. Note the history only ever held the user texts and the final answers — @@ -420,6 +434,8 @@ starter are the *deployment* layer on top of the same runtime — not needed for Atmosphere's `ApprovalResolution` also supports approve-with-edited-arguments, which the console does not offer. - No auto-compaction when the context fills up; `/compact` is manual. +- A small model still drifts into describing instead of doing, especially after several turns; + `/calls` makes it visible, the note in the history makes it rarer, a bigger model makes it go away. - `/loop` cannot be interrupted in the middle of a step — Ctrl-C ends the process; the loop file survives, so restarting the same `/loop` continues where it left off. - An engine error after the stream started ends the turn silently (see the table). diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleSession.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleSession.java index 6d7325ee6..a16b7ba05 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleSession.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleSession.java @@ -39,6 +39,7 @@ public final class ConsoleSession implements StreamingSession { private volatile @Nullable Throwable failure; private volatile int toolCalls; private volatile long inputTokens; + private final List rounds = new CopyOnWriteArrayList<>(); /** * Create a session writing to {@code terminal}. @@ -53,6 +54,24 @@ public ConsoleSession(AgentTerminal terminal, AgentFileSystem fileSystem) { this.injectables = Map.of(AgentFileSystem.class, fileSystem); } + /** + * One tool call of this turn, kept so the next turn can see that it happened. + * + * @param name the tool + * @param argumentsJson the arguments as JSON + * @param result what the tool returned, already shortened + */ + public record ToolRound(String name, String argumentsJson, String result) {} + + /** + * The tool calls of this turn, in order. + * + * @return the rounds, empty when the model only wrote text + */ + public List rounds() { + return List.copyOf(rounds); + } + @Override public String sessionId() { return "console"; @@ -132,20 +151,34 @@ public void emit(AiEvent event) { switch (event) { case AiEvent.ToolStart start -> { toolCalls++; + rounds.add(new ToolRound(start.toolName(), String.valueOf(start.arguments()), "")); markdown.flush(); terminal.line(ansi.green("●") + " " + ansi.bold(start.toolName()) + " " + ansi.dim(String.valueOf(start.arguments()))); } case AiEvent.ToolResult result -> { + recordResult(String.valueOf(result.result())); terminal.line(ansi.dim(" ↳ " + preview(String.valueOf(result.result())))); } case AiEvent.ToolError error -> { + recordResult("error: " + error.error()); terminal.line(ansi.red(" ↳ error: " + error.error())); } default -> StreamingSession.super.emit(event); } } + /** Attach a result to the round that is still waiting for one. */ + private void recordResult(String result) { + for (int i = rounds.size() - 1; i >= 0; i--) { + ToolRound round = rounds.get(i); + if (round.result().isEmpty()) { + rounds.set(i, new ToolRound(round.name(), round.argumentsJson(), result)); + return; + } + } + } + private static String preview(String value) { String oneLine = value.replace("\r\n", "\n").replace('\n', ' '); return oneLine.length() <= RESULT_PREVIEW_CHARS diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java index bf06ac876..500a67bc0 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java @@ -75,6 +75,9 @@ public final class LocalAgent { "You summarize a conversation between a user and a coding assistant. Follow the user's" + " instructions exactly and answer with the summary only."; + /** How much of a tool result is kept in the history of later turns. */ + private static final int HISTORY_RESULT_CHARS = 400; + /** How often the activity line is refreshed while a turn runs. */ private static final Duration ACTIVITY_INTERVAL = Duration.ofMillis(250); @@ -169,6 +172,7 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream + " tools=" + runner.toolNames()); List history = new ArrayList<>(); + ToolCallLog callLog = new ToolCallLog(); AtomicReference mode = new AtomicReference<>(options.isAuto() ? ApprovalMode.AUTO : ApprovalMode.MANUAL); boolean interactive = options.getPrompt() == null && input != null; @@ -185,7 +189,7 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream : ServerProps.contextSize(baseUrl, options.getApiKey()); if (options.getPrompt() != null) { - return turn(runner, fileSystem, options.getPrompt(), history, terminal) + return turn(runner, fileSystem, options.getPrompt(), history, terminal, callLog, 1) .failure() == null ? 0 @@ -198,6 +202,7 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream err.println("Interactive mode: type a request, /help for the commands."); long inputTokens = 0; boolean estimated = false; + int turnNumber = 0; while (true) { terminal.status(StatusLine.render( options.getWorkspace(), @@ -229,10 +234,12 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream contextSize, inputTokens, estimated, - terminal); + terminal, + callLog); continue; } - ConsoleSession completed = turn(runner, fileSystem, line, history, terminal); + turnNumber++; + ConsoleSession completed = turn(runner, fileSystem, line, history, terminal, callLog, turnNumber); estimated = completed.inputTokens() == 0; inputTokens = estimated ? estimateTokens(systemPrompt(options), history) : completed.inputTokens(); } @@ -264,6 +271,7 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream * @param options the agent options, for the workspace * @param mode the approval mode, possibly switched to auto here * @param arguments everything after {@code /loop} + * @param callLog records the steps' tool calls for {@code /calls} * @throws InterruptedException if interrupted while a step runs */ private static void loop( @@ -272,7 +280,8 @@ private static void loop( AgentTerminal terminal, AgentOptions options, AtomicReference mode, - String arguments) + String arguments, + ToolCallLog callLog) throws InterruptedException { LoopOptions loopOptions; try { @@ -301,7 +310,8 @@ private static void loop( options.getWorkspace(), loopOptions, () -> false, - TaskLoop.DEFAULT_BUDGET); + TaskLoop.DEFAULT_BUDGET, + callLog); terminal.line( outcome.completed() ? terminal.ansi().green("loop: " + outcome.reason()) @@ -313,14 +323,18 @@ static ConsoleSession turn( AgentFileSystem fileSystem, String message, List history, - AgentTerminal terminal) + AgentTerminal terminal, + ToolCallLog callLog, + int turnNumber) throws InterruptedException { ConsoleSession session = new ConsoleSession(terminal, fileSystem); runner.run(message, history, session); boolean finished = awaitWithActivity(session, terminal); history.add(ChatMessage.user(message)); - if (!session.text().isEmpty()) { - history.add(ChatMessage.assistant(session.text())); + callLog.add(turnNumber, session.rounds()); + String answer = withToolNotes(session.rounds(), session.text()); + if (!answer.isEmpty()) { + history.add(ChatMessage.assistant(answer)); } if (!finished) { session.error(new IllegalStateException("turn did not finish within " + TURN_TIMEOUT)); @@ -328,6 +342,49 @@ static ConsoleSession turn( return session; } + /** + * Put this turn's tool calls in front of its answer, so the next turn can see they happened. + * + *

Why this matters more than it looks. Without it the history holds the user's messages + * and the model's prose, and nothing else — so from the third or fourth turn on, a small model + * sees only its own paragraphs and no evidence that it ever used a tool. It then continues that + * pattern: it describes creating a file and running a build, reports an exit code, and + * writes nothing at all. That is not hypothetical; it happened on a real session, with the model + * inventing test results and a jar that never existed. + * + *

Why a note rather than real {@code tool_calls} messages. Atmosphere's + * {@code AbstractAgentRuntime.assembleMessages} rebuilds every history entry as + * {@code new ChatMessage(role, content)} — the tool-call array and the tool-call id are dropped on + * the way out. Protocol-faithful replay is therefore impossible through the framework's history; + * what survives is the content, so the evidence goes there. It is also cheaper: one line per call + * instead of a message pair, with the result cut to {@value #HISTORY_RESULT_CHARS} characters. + * + * @param rounds the tool calls of the finished turn + * @param text the model's answer + * @return the answer with the calls noted above it + */ + static String withToolNotes(List rounds, String text) { + if (rounds.isEmpty()) { + return text; + } + StringBuilder note = new StringBuilder("(tools I actually ran this turn:"); + for (ConsoleSession.ToolRound round : rounds) { + String result = round.result().replace("\r\n", " ").replace('\n', ' '); + note.append(System.lineSeparator()) + .append("- ") + .append(round.name()) + .append(" ") + .append(round.argumentsJson()) + .append(" -> ") + .append( + result.length() <= HISTORY_RESULT_CHARS + ? result + : result.substring(0, HISTORY_RESULT_CHARS) + " …[cut]"); + } + note.append(")"); + return text.isEmpty() ? note.toString() : note + System.lineSeparator() + System.lineSeparator() + text; + } + /** * Run one REPL command. * @@ -341,6 +398,7 @@ static ConsoleSession turn( * @param inputTokens the input tokens of the last turn * @param estimated whether that number is an estimate * @param terminal the console + * @param callLog every tool call of the session, for {@code /calls} * @return the input tokens to show from now on (unchanged, or the summary's after {@code /compact}) * @throws InterruptedException if interrupted while a summary is generated */ @@ -354,7 +412,8 @@ private static long handleCommand( int contextSize, long inputTokens, boolean estimated, - AgentTerminal terminal) + AgentTerminal terminal, + ToolCallLog callLog) throws InterruptedException { switch (command.command()) { case HELP -> prompt(HELP_TEXT).lines().forEach(terminal::line); @@ -362,6 +421,7 @@ private static long handleCommand( history.clear(); terminal.line("(history cleared)"); } + case CALLS -> callLog.render().lines().forEach(terminal::line); case TOOLS -> { terminal.line("tools: " + String.join(", ", runner.toolNames())); terminal.line("asks before running (manual mode): " @@ -393,7 +453,7 @@ private static long handleCommand( case COMPACT -> { return compact(runner, fileSystem, history, command.arguments(), terminal); } - case LOOP -> loop(runner, fileSystem, terminal, options, mode, command.arguments()); + case LOOP -> loop(runner, fileSystem, terminal, options, mode, command.arguments(), callLog); case EXIT -> { // handled by the caller, which has to return from the loop } diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/SlashCommands.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/SlashCommands.java index 185ea1fa6..a58504638 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/SlashCommands.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/SlashCommands.java @@ -40,6 +40,8 @@ public enum Command { STATUS("/status"), /** List the tools offered to the model. */ TOOLS("/tools"), + /** Show every tool call of this session — the receipt for what really happened. */ + CALLS("/calls", "/log"), /** Leave the REPL. */ EXIT("/exit", "/quit"); diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/TaskLoop.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/TaskLoop.java index 0686f05b4..579e53b17 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/TaskLoop.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/TaskLoop.java @@ -142,6 +142,7 @@ public record Outcome(String reason, boolean completed) {} * @param options what to work on and for how long * @param stopped polled between steps; {@code true} ends the loop (Ctrl-C) * @param budget the wall-clock limit + * @param callLog records every tool call of every step, so /calls shows what the loop did * @return why it ended * @throws InterruptedException if interrupted while waiting for a step or an interval */ @@ -152,7 +153,8 @@ public static Outcome run( Path workspace, LoopOptions options, java.util.function.BooleanSupplier stopped, - Duration budget) + Duration budget, + ToolCallLog callLog) throws InterruptedException { Path file = ensureLoopFile(workspace, options.task()); terminal.line("loop: " + options.task()); @@ -180,7 +182,13 @@ public static Outcome run( // A fresh history every step: the file is the memory, so the context cannot grow. ConsoleSession session = LocalAgent.turn( - runner, fileSystem, stepPrompt(options) + extra, new java.util.ArrayList<>(), terminal); + runner, + fileSystem, + stepPrompt(options) + extra, + new java.util.ArrayList<>(), + terminal, + callLog, + step); extra = ""; if (session.failure() != null) { return new Outcome("step " + step + " failed: " + session.failure(), false); diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ToolCallLog.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ToolCallLog.java new file mode 100644 index 000000000..9530719a1 --- /dev/null +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ToolCallLog.java @@ -0,0 +1,100 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import java.time.LocalTime; +import java.time.format.DateTimeFormatter; +import java.util.ArrayList; +import java.util.List; + +/** + * Every tool call of the session, in order, for {@code /calls}. + * + *

It answers one question the transcript cannot: did that actually happen? A model that + * runs out of context, or simply drifts, starts describing work instead of doing it — reporting an + * exit code, a created file, a passing test, all invented. The scrollback looks convincing, because + * the prose is the same either way. This log only ever grows when a tool really ran, so an empty or + * short list is the proof. + * + *

Kept small on purpose: one line per call, arguments and result cut hard. It is a receipt, not a + * second transcript. + */ +public final class ToolCallLog { + + /** Characters kept of the arguments and of the result. */ + private static final int PREVIEW_CHARS = 120; + + private static final DateTimeFormatter TIME = DateTimeFormatter.ofPattern("HH:mm:ss"); + + private final List entries = new ArrayList<>(); + + /** + * One recorded call. + * + * @param time when it ran + * @param turn the user turn it belonged to, counting from 1 + * @param name the tool + * @param arguments the arguments, shortened + * @param result what came back, shortened + */ + public record Entry(LocalTime time, int turn, String name, String arguments, String result) {} + + /** + * Record the calls of one finished turn. + * + * @param turn the turn number + * @param rounds the calls, in order + */ + public void add(int turn, List rounds) { + for (ConsoleSession.ToolRound round : rounds) { + entries.add(new Entry( + LocalTime.now(), turn, round.name(), cut(round.argumentsJson()), cut(oneLine(round.result())))); + } + } + + /** + * How many calls were made in this session. + * + * @return the count + */ + public int size() { + return entries.size(); + } + + /** + * The log as the console shows it. + * + * @return one line per call, or a sentence saying there were none + */ + public String render() { + if (entries.isEmpty()) { + return "No tool has been called in this session — everything so far was text only."; + } + StringBuilder text = new StringBuilder(); + for (Entry entry : entries) { + text.append(entry.time().format(TIME)) + .append(" turn ") + .append(entry.turn()) + .append(" ") + .append(entry.name()) + .append(" ") + .append(entry.arguments()) + .append(System.lineSeparator()) + .append(" ↳ ") + .append(entry.result()) + .append(System.lineSeparator()); + } + text.append(entries.size()).append(" calls"); + return text.toString(); + } + + private static String oneLine(String text) { + return text.replace("\r\n", " ").replace('\n', ' ').strip(); + } + + private static String cut(String text) { + return text.length() <= PREVIEW_CHARS ? text : text.substring(0, PREVIEW_CHARS) + " …"; + } +} diff --git a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt index 0f1a4d58f..3a281efc9 100644 --- a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt +++ b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt @@ -3,6 +3,7 @@ Commands (everything else is sent to the model): /help this overview (/?, /commands) /status endpoint, model, tools, mode, context use /tools the tools offered to the model + /calls every tool call of this session, with its result (/log) /mode [manual|auto] show or set the approval mode (/approve) /compact [focus] summarize the history and continue with the summary /loop [--every 5m] [--max 20] [--check ''] diff --git a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/system-prompt.txt b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/system-prompt.txt index 43d1afa39..1e2f59396 100644 --- a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/system-prompt.txt +++ b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/system-prompt.txt @@ -4,4 +4,4 @@ The file tools ls, read_file, write_file, edit_file, glob, grep, delete and rena {shell_section} -Work step by step. read_file shows a numbered window of a file — those numbers are display only, never copy them into old_string. Edit a file only after reading it (edit_file refuses otherwise), copy old_string from what you read, and check the result after a change. Finish with a short summary. Answer in the user's language. +Work step by step. read_file shows a numbered window of a file — those numbers are display only, never copy them into old_string. Edit a file only after reading it (edit_file refuses otherwise), copy old_string from what you read, and check the result after a change. Finish with a short summary. Never claim that you ran a command, created a file or saw a result unless you actually called the tool in this turn and read what it returned; if you did not, say so plainly instead of describing what would happen. Answer in the user's language. diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java index 982fe8782..bd49c01f7 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java @@ -5,6 +5,7 @@ package net.ladenthin.llama.atmosphere; import static org.hamcrest.MatcherAssert.assertThat; +import static org.hamcrest.Matchers.contains; import static org.hamcrest.Matchers.containsString; import static org.hamcrest.Matchers.hasItem; import static org.hamcrest.Matchers.hasSize; @@ -19,6 +20,7 @@ import java.nio.charset.StandardCharsets; import java.nio.file.Files; import java.nio.file.Path; +import java.util.ArrayList; import java.util.List; import net.ladenthin.llama.server.OpenAiCompatServer; import net.ladenthin.llama.server.OpenAiServerConfig; @@ -187,4 +189,47 @@ void mavenJvmConfigPinsAUtf8ConsoleForExecJava() throws Exception { assertThat(content, containsString("-Dstdout.encoding=UTF-8")); assertThat(content, containsString("-Dstderr.encoding=UTF-8")); } + + @Test + void aToolCallStaysInTheHistorySoTheNextTurnSeesItHappened() throws Exception { + // The failure this pins: with only user text and the model's prose in the history, a small + // model stops calling tools after a few turns and starts DESCRIBING the work instead -- + // reporting exit codes and files that never existed. The evidence has to stay in the context. + Files.writeString(workspace.resolve("hello.txt"), "VALUE=42\n"); + ByteArrayOutputStream out = new ByteArrayOutputStream(); + ByteArrayOutputStream err = new ByteArrayOutputStream(); + ScriptedBackend backend = new ScriptedBackend((call, request) -> switch (call) { + case 1 -> ScriptedBackend.toolCallTurn("call_1", "read_file", "{\"file_path\":\"hello.txt\"}"); + case 2 -> ScriptedBackend.textTurn("The file says VALUE=42."); + default -> ScriptedBackend.textTurn("Understood."); + }); + try (OpenAiCompatServer server = server(backend)) { + AgentOptions options = AgentOptions.parse(new String[] { + "--base-url", "http://127.0.0.1:" + server.getPort() + "/v1", "--workspace", workspace.toString() + }); + + LocalAgent.run( + options, + new StringReader( + "what does hello.txt say?" + System.lineSeparator() + "and now?" + System.lineSeparator()), + new PrintStream(out, true, StandardCharsets.UTF_8), + new PrintStream(err, true, StandardCharsets.UTF_8)); + } + + // third request = second turn: the call and its result must be in what the model gets to see. + // Not as real tool_calls messages -- Atmosphere's assembleMessages rebuilds history entries as + // new ChatMessage(role, content) and drops everything else -- so they ride in the content. + List requests = backend.requests(); + assertThat(requests.size(), is(3)); + List roles = new ArrayList<>(); + for (JsonNode message : requests.get(2).path("messages")) { + roles.add(message.path("role").asText()); + } + assertThat(roles, contains("system", "user", "assistant", "user")); + String replayedAnswer = + requests.get(2).path("messages").get(2).path("content").asText(); + assertThat(replayedAnswer, containsString("read_file")); + assertThat(replayedAnswer, containsString("VALUE=42")); + assertThat("and the answer itself is still there", replayedAnswer, containsString("The file says VALUE=42.")); + } } diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/TaskLoopTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/TaskLoopTest.java index 3058836df..49facd04e 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/TaskLoopTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/TaskLoopTest.java @@ -58,7 +58,7 @@ private TaskLoop.Outcome runLoop(ScriptedBackend backend, LoopOptions options) t 5) .retryPolicy(RetryPolicy.NONE); AgentFileSystem fs = new WorkspaceAgentFileSystem(workspace, AgentFileSystem.Limits.defaults()); - return TaskLoop.run(runner, fs, terminal(), workspace, options, () -> false, BUDGET); + return TaskLoop.run(runner, fs, terminal(), workspace, options, () -> false, BUDGET, new ToolCallLog()); } } From 0785f4cf54283dde284af60a594f61fc5533f6d2 Mon Sep 17 00:00:00 2001 From: Bernard Ladenthin Date: Wed, 23 Sep 2026 12:57:51 +0200 Subject: [PATCH 07/36] llama-atmosphere-agent: live progress, and stop the model parroting the tool record MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Two problems from one session. The tool record added in the previous commit rode in front of the assistant's answer -- and the model copied it into its next reply, so the user read "(tools I actually ran this turn:" as the first line of an answer. And a long command showed nothing at all until it finished. - The record now rides in front of the NEXT USER message. An assistant message is text a model imitates; a mid-history system message would be cleaner but Mistral's template requires strict user/assistant alternation and Gemma has no system role, so this is the placement every template accepts. It also says "do not repeat it". - The pinned line names what is happening: "⠙ thinking… (5s)" while the model generates, "⠙ run_command… (47s of 61s · 2 tool calls)" while a tool runs. A two-minute build was indistinguishable from a hang. - run_command prints its output line by line while it runs instead of dumping it at the end. Reading the pipe incrementally also keeps the child from blocking once the pipe buffer is full -- about 4 KB on Windows. Tests: 105 green (was 102). Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01E2h8gXyyE5UeimkL9vQv9G --- CLAUDE.md | 10 +++ llama-atmosphere-agent/README.md | 26 +++++-- .../llama/atmosphere/ConsoleSession.java | 24 ++++++ .../llama/atmosphere/LocalAgent.java | 77 +++++++++++++++---- .../ladenthin/llama/atmosphere/ShellTool.java | 58 ++++++++++++-- .../llama/atmosphere/LocalAgentTest.java | 49 +++++++++--- .../llama/atmosphere/ShellToolTest.java | 19 +++++ 7 files changed, 227 insertions(+), 36 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 4f53ef8e8..c72686959 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -2314,7 +2314,17 @@ are decisions, not details: model stopped calling tools after the third turn of a real session and *described* the work instead — inventing JUnit tests, a Maven build and a `.bat` script, complete with exit codes, while the workspace stayed empty. The system prompt also forbids claiming an action without the call. + **Placement was found by failing twice, so do not "simplify" it:** in front of the assistant's + answer made the model copy the record into its own replies (the user saw `(tools I actually ran + this turn: …)` as the first line of an answer); real `tool_calls` messages are impossible (see + above); a mid-history system message is cleanest but Mistral's template requires strict + user/assistant alternation and Gemma has no system role. It therefore rides in front of the **next + user message**, which every template accepts. Pinned by `LocalAgentTest.aToolCallStaysInTheHistorySoTheNextTurnSeesItHappened`. + **Live feedback while a turn runs** (`LocalAgent.activityLine`, `ShellTool`'s line-by-line output): + the pinned line names the running tool and its own elapsed time, and shell output is printed as it + arrives. Reading the pipe incrementally is not only cosmetic — an unread pipe blocks the child once + it is full, which on Windows is roughly 4 KB. 8. **The context number in the status line is an estimate, marked `~`.** llama.cpp emits its usage chunk only when the client sets `stream_options.include_usage`, and Atmosphere's client does not; `ConsoleSession.usage()` takes the real count when one arrives, otherwise `LocalAgent.estimateTokens` diff --git a/llama-atmosphere-agent/README.md b/llama-atmosphere-agent/README.md index 03a5628cc..e42a61891 100644 --- a/llama-atmosphere-agent/README.md +++ b/llama-atmosphere-agent/README.md @@ -220,13 +220,18 @@ arguments and a short result. Use it when an answer sounds too good: a model tha *describing* work — "the tests passed, the jar was created" — while calling nothing at all. The scrollback reads the same either way; this list only grows when something really ran. -To make that drift less likely, each turn's calls are also carried into the conversation as a short -note above the answer (`(tools I actually ran this turn: …)`). Without it the history holds only the -user's messages and the model's own prose, and a small model then continues the pattern it sees — -prose. This is not theoretical: in a real session a 4B model invented JUnit tests, a Maven build and a -`.bat` script it had never written, three turns in a row. The note is plain text rather than proper -`tool_calls` messages because Atmosphere's `assembleMessages` rebuilds every history entry as -`new ChatMessage(role, content)` and drops the rest; the content is what survives. +To make that drift less likely, each turn's calls ride along with the **next** message as a short +record. Without it the history holds only the user's messages and the model's own prose, and a small +model then continues the pattern it sees — prose. This is not theoretical: in a real session a 4B +model invented JUnit tests, a Maven build and a `.bat` script it had never written, three turns in a +row. + +Two placements were tried and discarded, both visible failures: in front of the assistant's answer +(the model copied it into its next reply, so the record appeared as the first line of an answer), and +as real `tool_calls` messages (impossible — Atmosphere's `assembleMessages` rebuilds every history +entry as `new ChatMessage(role, content)` and drops the rest). A system message mid-history would be +cleaner, but Mistral's template requires strict user/assistant alternation and Gemma has no system +role at all. **`/compact`** asks the model to summarize the conversation (goal, facts, work done, problems, state, next step; `/compact ` adds an emphasis), then replaces the history with that summary. Use it @@ -260,6 +265,13 @@ the model: `--max` steps (20 by default), a two-hour wall-clock budget, a stall in a row that write nothing and call no tool), and `--every ` for a paced run. A loop needs the `auto` approval mode — it asks once and switches, or leaves you alone if you say no. +**While a turn runs, the pinned line says what is happening**: `⠙ thinking… (5s)` while the model +generates, and `⠙ run_command… (47s of 61s · 2 tool calls)` while a tool is executing. A build that +takes two minutes is otherwise indistinguishable from a hang. `run_command` additionally prints its +output **line by line while it runs** (dimmed, `│ `-prefixed) instead of dumping it at the end — which +also keeps the pipe drained; a process whose output nobody reads blocks once the buffer is full, and +on Windows that buffer is about 4 KB. + **The status line** above the prompt reads `[manual · ctx ~3.1k/16k · 9 tools · local-model]`: the approval mode, the context used out of the window, the number of tools and the model id. A `~` means the number is an estimate from the text diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleSession.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleSession.java index a16b7ba05..53d2f38c9 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleSession.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleSession.java @@ -40,6 +40,8 @@ public final class ConsoleSession implements StreamingSession { private volatile int toolCalls; private volatile long inputTokens; private final List rounds = new CopyOnWriteArrayList<>(); + private volatile @Nullable String runningTool; + private volatile long runningSince; /** * Create a session writing to {@code terminal}. @@ -151,16 +153,20 @@ public void emit(AiEvent event) { switch (event) { case AiEvent.ToolStart start -> { toolCalls++; + runningTool = start.toolName(); + runningSince = System.nanoTime(); rounds.add(new ToolRound(start.toolName(), String.valueOf(start.arguments()), "")); markdown.flush(); terminal.line(ansi.green("●") + " " + ansi.bold(start.toolName()) + " " + ansi.dim(String.valueOf(start.arguments()))); } case AiEvent.ToolResult result -> { + runningTool = null; recordResult(String.valueOf(result.result())); terminal.line(ansi.dim(" ↳ " + preview(String.valueOf(result.result())))); } case AiEvent.ToolError error -> { + runningTool = null; recordResult("error: " + error.error()); terminal.line(ansi.red(" ↳ error: " + error.error())); } @@ -168,6 +174,24 @@ public void emit(AiEvent event) { } } + /** + * The tool that is executing right now, if any. + * + * @return the tool name, or {@code null} when the model is generating rather than running something + */ + public @Nullable String runningTool() { + return runningTool; + } + + /** + * How long the running tool has been running. + * + * @return the seconds since it started, or {@code 0} when nothing runs + */ + public long runningSeconds() { + return runningTool == null ? 0 : (System.nanoTime() - runningSince) / 1_000_000_000L; + } + /** Attach a result to the round that is still waiting for one. */ private void recordResult(String result) { for (int i = rounds.size() - 1; i >= 0; i--) { diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java index 500a67bc0..f45503f22 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java @@ -157,7 +157,13 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream // read tracker is what lets an edit insist the file was read first. List tools = new ArrayList<>(WorkspaceTools.all(new WorkspaceTools.ReadTracker())); if (options.isAllowShell()) { - tools.add(ShellTool.definition(options.getWorkspace(), SHELL_TIMEOUT, SHELL_MAX_OUTPUT_CHARS)); + // Live output: a two-minute build has to show that it is doing something. + AgentTerminal console = terminal; + tools.add(ShellTool.definition( + options.getWorkspace(), + SHELL_TIMEOUT, + SHELL_MAX_OUTPUT_CHARS, + line -> console.line(console.ansi().dim(" │ " + line)))); } AgentRunner runner = new AgentRunner( baseUrl, @@ -203,6 +209,7 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream long inputTokens = 0; boolean estimated = false; int turnNumber = 0; + String pendingNote = ""; while (true) { terminal.status(StatusLine.render( options.getWorkspace(), @@ -239,7 +246,17 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream continue; } turnNumber++; - ConsoleSession completed = turn(runner, fileSystem, line, history, terminal, callLog, turnNumber); + ConsoleSession completed = turn( + runner, + fileSystem, + pendingNote.isEmpty() + ? line + : pendingNote + System.lineSeparator() + System.lineSeparator() + line, + history, + terminal, + callLog, + turnNumber); + pendingNote = toolNote(completed.rounds()); estimated = completed.inputTokens() == 0; inputTokens = estimated ? estimateTokens(systemPrompt(options), history) : completed.inputTokens(); } @@ -332,9 +349,8 @@ static ConsoleSession turn( boolean finished = awaitWithActivity(session, terminal); history.add(ChatMessage.user(message)); callLog.add(turnNumber, session.rounds()); - String answer = withToolNotes(session.rounds(), session.text()); - if (!answer.isEmpty()) { - history.add(ChatMessage.assistant(answer)); + if (!session.text().isEmpty()) { + history.add(ChatMessage.assistant(session.text())); } if (!finished) { session.error(new IllegalStateException("turn did not finish within " + TURN_TIMEOUT)); @@ -352,6 +368,14 @@ static ConsoleSession turn( * writes nothing at all. That is not hypothetical; it happened on a real session, with the model * inventing test results and a jar that never existed. * + *

Where it goes, and why not somewhere more obvious. In front of the next user + * message. The first attempt put it in front of the assistant's own answer, and the model + * promptly copied it into its next reply — the user read "(tools I actually ran this turn: …)" as + * the first line of an answer, because text attributed to the assistant is text a model imitates. + * A system message mid-history would be cleaner still, but not every chat template accepts one: + * Mistral's requires strict user/assistant alternation and Gemma has no system role at all. + * Riding along with the next user message keeps the sequence template-safe everywhere. + * *

Why a note rather than real {@code tool_calls} messages. Atmosphere's * {@code AbstractAgentRuntime.assembleMessages} rebuilds every history entry as * {@code new ChatMessage(role, content)} — the tool-call array and the tool-call id are dropped on @@ -360,14 +384,14 @@ static ConsoleSession turn( * instead of a message pair, with the result cut to {@value #HISTORY_RESULT_CHARS} characters. * * @param rounds the tool calls of the finished turn - * @param text the model's answer - * @return the answer with the calls noted above it + * @return the note, or an empty string when no tool ran */ - static String withToolNotes(List rounds, String text) { + static String toolNote(List rounds) { if (rounds.isEmpty()) { - return text; + return ""; } - StringBuilder note = new StringBuilder("(tools I actually ran this turn:"); + StringBuilder note = new StringBuilder("Record of the tools that actually ran in the previous turn." + + " This is a log for your reference; do not repeat it and do not mention it."); for (ConsoleSession.ToolRound round : rounds) { String result = round.result().replace("\r\n", " ").replace('\n', ' '); note.append(System.lineSeparator()) @@ -381,8 +405,7 @@ static String withToolNotes(List rounds, String text) ? result : result.substring(0, HISTORY_RESULT_CHARS) + " …[cut]"); } - note.append(")"); - return text.isEmpty() ? note.toString() : note + System.lineSeparator() + System.lineSeparator() + text; + return note.toString(); } /** @@ -482,13 +505,39 @@ private static boolean awaitWithActivity(ConsoleSession session, AgentTerminal t if (seconds > TURN_TIMEOUT.toSeconds()) { return false; } - terminal.status(ACTIVITY_FRAMES.charAt(frame++ % ACTIVITY_FRAMES.length()) + " working… (" + seconds - + "s · " + session.toolCalls() + " tool calls)"); + terminal.status(activityLine( + ACTIVITY_FRAMES.charAt(frame++ % ACTIVITY_FRAMES.length()), + seconds, + session.runningTool(), + session.runningSeconds(), + session.toolCalls())); } terminal.status(""); return true; } + /** + * What the pinned line says while a turn is running. + * + *

Naming the running tool is the point: a build or a test run can take minutes, and + * "working…" during a two-minute {@code mvn test} is indistinguishable from a hang. When no tool + * runs, the model is generating, which is its own kind of waiting. + * + * @param frame the spinner character + * @param seconds how long the whole turn has been running + * @param runningTool the tool executing right now, or {@code null} + * @param toolSeconds how long that tool has been running + * @param toolCalls how many tools ran in this turn so far + * @return the line + */ + static String activityLine( + char frame, long seconds, @Nullable String runningTool, long toolSeconds, int toolCalls) { + String what = runningTool == null + ? "thinking… (" + seconds + "s" + : runningTool + "… (" + toolSeconds + "s of " + seconds + "s"; + return frame + " " + what + (toolCalls == 0 ? "" : " · " + toolCalls + " tool calls") + ")"; + } + /** * A rough token count of what the next request will carry. * diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ShellTool.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ShellTool.java index 7e2604d27..a3493ac4d 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ShellTool.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ShellTool.java @@ -4,7 +4,9 @@ package net.ladenthin.llama.atmosphere; +import java.io.BufferedReader; import java.io.IOException; +import java.io.InputStreamReader; import java.nio.charset.StandardCharsets; import java.nio.file.Path; import java.time.Duration; @@ -12,6 +14,7 @@ import java.util.concurrent.ExecutionException; import java.util.concurrent.TimeUnit; import java.util.concurrent.TimeoutException; +import java.util.function.Consumer; import org.atmosphere.ai.tool.ToolDefinition; /** @@ -41,6 +44,20 @@ private ShellTool() {} * @return the definition */ public static ToolDefinition definition(Path workspace, Duration defaultTimeout, int maxOutputChars) { + return definition(workspace, defaultTimeout, maxOutputChars, line -> {}); + } + + /** + * Build the tool definition with live output. + * + * @param workspace the working directory of every command + * @param defaultTimeout the timeout applied when the model does not pass {@code timeout_seconds} + * @param maxOutputChars output is truncated to this many characters (tail kept, head marked) + * @param liveOutput receives each output line while the command is still running + * @return the definition + */ + public static ToolDefinition definition( + Path workspace, Duration defaultTimeout, int maxOutputChars, Consumer liveOutput) { return ToolDefinition.builder( TOOL_NAME, LocalAgent.prompt(DESCRIPTION_RESOURCE).replace("{shell}", shellName())) .parameter(PARAM_COMMAND, "The command line to run through " + shellName(), "string", true) @@ -51,7 +68,7 @@ public static ToolDefinition definition(Path workspace, Duration defaultTimeout, return "Error: '" + PARAM_COMMAND + "' is required"; } Duration timeout = timeoutOf(args.get(PARAM_TIMEOUT), defaultTimeout); - return run(workspace, command.toString(), timeout, maxOutputChars); + return run(workspace, command.toString(), timeout, maxOutputChars, liveOutput); }) .build(); } @@ -106,18 +123,47 @@ private static Duration timeoutOf(Object raw, Duration fallback) { */ static String run(Path workspace, String command, Duration timeout, int maxOutputChars) throws IOException, InterruptedException { + return run(workspace, command, timeout, maxOutputChars, line -> {}); + } + + /** + * Run one command through the platform shell, reporting its output as it arrives. + * + *

The output is read line by line rather than in one go at the end. That is what lets the + * console show a long build while it runs — a silent minute is indistinguishable from a hang — and + * it is also what keeps the pipe drained: a process whose output nobody reads blocks once the + * pipe buffer is full, which on Windows is roughly 4 KB. + * + * @param workspace the working directory + * @param command the command line + * @param timeout kill the process after this long + * @param maxOutputChars truncate the captured output to this many characters + * @param liveOutput receives each line as it is read + * @return a text block starting with {@code exit code: N}, followed by the output + * @throws IOException if the process cannot be started + * @throws InterruptedException if interrupted while waiting + */ + static String run(Path workspace, String command, Duration timeout, int maxOutputChars, Consumer liveOutput) + throws IOException, InterruptedException { ProcessBuilder builder = isWindows() ? new ProcessBuilder("cmd.exe", "/c", command) : new ProcessBuilder("sh", "-c", command); builder.directory(workspace.toFile()); builder.redirectErrorStream(true); Process process = builder.start(); process.getOutputStream().close(); - CompletableFuture output = CompletableFuture.supplyAsync(() -> { - try { - return process.getInputStream().readAllBytes(); + CompletableFuture output = CompletableFuture.supplyAsync(() -> { + StringBuilder collected = new StringBuilder(); + try (BufferedReader reader = + new BufferedReader(new InputStreamReader(process.getInputStream(), StandardCharsets.UTF_8))) { + String line; + while ((line = reader.readLine()) != null) { + collected.append(line).append(System.lineSeparator()); + liveOutput.accept(line); + } } catch (IOException e) { - return ("[output unreadable: " + e.getMessage() + "]").getBytes(StandardCharsets.UTF_8); + collected.append("[output unreadable: ").append(e.getMessage()).append("]"); } + return collected.toString(); }); boolean finished = process.waitFor(timeout.toMillis(), TimeUnit.MILLISECONDS); if (!finished) { @@ -129,7 +175,7 @@ static String run(Path workspace, String command, Duration timeout, int maxOutpu } String text; try { - text = new String(output.get(5, TimeUnit.SECONDS), StandardCharsets.UTF_8); + text = output.get(5, TimeUnit.SECONDS); } catch (ExecutionException | TimeoutException e) { text = "[output unavailable: " + e.getMessage() + "]"; } diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java index bd49c01f7..17d6a90f8 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java @@ -216,20 +216,51 @@ void aToolCallStaysInTheHistorySoTheNextTurnSeesItHappened() throws Exception { new PrintStream(err, true, StandardCharsets.UTF_8)); } - // third request = second turn: the call and its result must be in what the model gets to see. - // Not as real tool_calls messages -- Atmosphere's assembleMessages rebuilds history entries as - // new ChatMessage(role, content) and drops everything else -- so they ride in the content. + // third request = second turn: the calls of turn one ride along with the new user message, so + // the model sees that they happened. Not as tool_calls messages (Atmosphere's assembleMessages + // rebuilds history as new ChatMessage(role, content) and drops the rest), and not as assistant + // text either (the model copied that into its own answers). List requests = backend.requests(); assertThat(requests.size(), is(3)); List roles = new ArrayList<>(); for (JsonNode message : requests.get(2).path("messages")) { roles.add(message.path("role").asText()); } - assertThat(roles, contains("system", "user", "assistant", "user")); - String replayedAnswer = - requests.get(2).path("messages").get(2).path("content").asText(); - assertThat(replayedAnswer, containsString("read_file")); - assertThat(replayedAnswer, containsString("VALUE=42")); - assertThat("and the answer itself is still there", replayedAnswer, containsString("The file says VALUE=42.")); + assertThat( + "strict alternation keeps every chat template happy", + roles, + contains("system", "user", "assistant", "user")); + String secondUserMessage = + requests.get(2).path("messages").get(3).path("content").asText(); + assertThat(secondUserMessage, containsString("read_file")); + assertThat(secondUserMessage, containsString("VALUE=42")); + assertThat(secondUserMessage, containsString("do not repeat it")); + assertThat("and the user's own words are still there", secondUserMessage, containsString("and now?")); + assertThat( + "the answer itself stays the model's own", + requests.get(2).path("messages").get(2).path("content").asText(), + is("The file says VALUE=42.")); + } + + @Test + void theActivityLineNamesTheRunningToolSoALongBuildLooksAlive() { + // "working… (90s)" during a two-minute mvn test is indistinguishable from a hang. + assertThat( + LocalAgent.activityLine('x', 12, "run_command", 9, 2), is("x run_command… (9s of 12s · 2 tool calls)")); + assertThat(LocalAgent.activityLine('x', 5, null, 0, 0), is("x thinking… (5s)")); + assertThat(LocalAgent.activityLine('x', 30, null, 0, 3), is("x thinking… (30s · 3 tool calls)")); + } + + @Test + void theToolNoteIsAddressedToTheModelAndNotWrittenAsItsOwnWords() { + // It first rode in front of the assistant's answer -- and the model copied it into its next + // reply, so the user read "(tools I actually ran this turn: …)" as the first line of an answer. + String note = LocalAgent.toolNote( + List.of(new ConsoleSession.ToolRound("grep", "{pattern=Test}", "3 matches in 2 files"))); + + assertThat(note, containsString("do not repeat it")); + assertThat(note, containsString("grep")); + assertThat(note, containsString("3 matches in 2 files")); + assertThat(LocalAgent.toolNote(List.of()), is("")); } } diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ShellToolTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ShellToolTest.java index 0d92693e1..2b1e8fe1d 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ShellToolTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ShellToolTest.java @@ -12,6 +12,7 @@ import java.nio.file.Files; import java.nio.file.Path; import java.time.Duration; +import java.util.List; import java.util.Map; import org.atmosphere.ai.tool.ToolDefinition; import org.junit.jupiter.api.Test; @@ -90,4 +91,22 @@ void timeoutKillsTheProcess() throws Exception { assertThat(result, startsWith("exit code: (killed after 0 s)")); } + + @Test + void outputIsReportedLineByLineWhileTheCommandStillRuns() throws Exception { + List live = new java.util.concurrent.CopyOnWriteArrayList<>(); + + String result = ShellTool.run( + workspace, + shell("echo one; echo two", "echo one& echo two"), + Duration.ofSeconds(30), + 10_000, + live::add); + + // the same lines reach the console while the process runs and the model afterwards + assertThat(live, is(java.util.List.of("one", "two"))); + assertThat(result, containsString("one")); + assertThat(result, containsString("two")); + assertThat(result, startsWith("exit code: 0")); + } } From 2a1bca816e928343d2e31852f5bf2bca378b310c Mon Sep 17 00:00:00 2001 From: Bernard Ladenthin Date: Wed, 23 Sep 2026 14:29:23 +0200 Subject: [PATCH 08/36] llama-atmosphere-agent: our own spinner words MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The activity line now draws a word from spinner-words.txt once per turn: "⠙ Fettling… (run_command 47s of 61s · 2 tool calls)". Deliberately our own two dozen words, not Claude Code's: that list is extracted from a proprietary binary, and the public copies are either unlicensed (all rights reserved) or CC BY-NC-SA, which is incompatible with MIT and with REUSE. A test asserts that none of their words appear here, that the list has no duplicates and that every entry is a plain capitalised word. The tool name and the elapsed times stay next to the word: with a local model "which tool, for how long" carries more information than the joke does. Tests: 106 green. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01E2h8gXyyE5UeimkL9vQv9G --- llama-atmosphere-agent/README.md | 10 ++++-- .../llama/atmosphere/LocalAgent.java | 35 ++++++++++++++++--- .../llama/atmosphere/spinner-words.txt | 26 ++++++++++++++ .../atmosphere/spinner-words.txt.license | 3 ++ .../llama/atmosphere/LocalAgentTest.java | 32 ++++++++++++++--- 5 files changed, 95 insertions(+), 11 deletions(-) create mode 100644 llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/spinner-words.txt create mode 100644 llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/spinner-words.txt.license diff --git a/llama-atmosphere-agent/README.md b/llama-atmosphere-agent/README.md index e42a61891..dca0093f9 100644 --- a/llama-atmosphere-agent/README.md +++ b/llama-atmosphere-agent/README.md @@ -265,8 +265,14 @@ the model: `--max` steps (20 by default), a two-hour wall-clock budget, a stall in a row that write nothing and call no tool), and `--every ` for a paced run. A loop needs the `auto` approval mode — it asks once and switches, or leaves you alone if you say no. -**While a turn runs, the pinned line says what is happening**: `⠙ thinking… (5s)` while the model -generates, and `⠙ run_command… (47s of 61s · 2 tool calls)` while a tool is executing. A build that +**While a turn runs, the pinned line says what is happening**: `⠙ Fettling… (5s)` while the model +generates, and `⠙ Fettling… (run_command 47s of 61s · 2 tool calls)` while a tool is executing. The +word is drawn once per turn from +[`spinner-words.txt`](src/main/resources/net/ladenthin/llama/atmosphere/spinner-words.txt) — our own +two dozen, because Claude Code's list is extracted from a proprietary binary and the public copies of +it are either unlicensed or CC BY-NC-SA, neither of which fits an MIT project. Edit the file to +change them. The numbers stay next to the word on purpose: with a local model, "which tool, for how +long" is worth more than the joke. A build that takes two minutes is otherwise indistinguishable from a hang. `run_command` additionally prints its output **line by line while it runs** (dimmed, `│ `-prefixed) instead of dumping it at the end — which also keeps the pipe drained; a process whose output nobody reads blocks once the buffer is full, and diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java index f45503f22..331295378 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java @@ -84,6 +84,9 @@ public final class LocalAgent { /** The spinner shown in the activity line. */ private static final String ACTIVITY_FRAMES = "⠋⠙⠹⠸⠼⠴⠦⠧⠇⠏"; + /** The whimsical words the activity line picks from, one per turn. */ + static final String SPINNER_WORDS = "spinner-words.txt"; + private static final Duration SHELL_TIMEOUT = Duration.ofSeconds(120); private static final int SHELL_MAX_OUTPUT_CHARS = 20_000; @@ -500,6 +503,8 @@ private static boolean awaitWithActivity(ConsoleSession session, AgentTerminal t throws InterruptedException { long start = System.nanoTime(); int frame = 0; + // one word per turn, not per frame: a word that changes ten times a second is noise + String word = spinnerWord(); while (!session.await(ACTIVITY_INTERVAL)) { long seconds = (System.nanoTime() - start) / 1_000_000_000L; if (seconds > TURN_TIMEOUT.toSeconds()) { @@ -507,6 +512,7 @@ private static boolean awaitWithActivity(ConsoleSession session, AgentTerminal t } terminal.status(activityLine( ACTIVITY_FRAMES.charAt(frame++ % ACTIVITY_FRAMES.length()), + word, seconds, session.runningTool(), session.runningSeconds(), @@ -531,11 +537,30 @@ private static boolean awaitWithActivity(ConsoleSession session, AgentTerminal t * @return the line */ static String activityLine( - char frame, long seconds, @Nullable String runningTool, long toolSeconds, int toolCalls) { - String what = runningTool == null - ? "thinking… (" + seconds + "s" - : runningTool + "… (" + toolSeconds + "s of " + seconds + "s"; - return frame + " " + what + (toolCalls == 0 ? "" : " · " + toolCalls + " tool calls") + ")"; + char frame, String word, long seconds, @Nullable String runningTool, long toolSeconds, int toolCalls) { + String inside = runningTool == null ? seconds + "s" : runningTool + " " + toolSeconds + "s of " + seconds + "s"; + return frame + " " + word + "… (" + inside + (toolCalls == 0 ? "" : " · " + toolCalls + " tool calls") + ")"; + } + + /** + * A word for the activity line, drawn once per turn. + * + *

Our own list ({@value #SPINNER_WORDS}), not the one Claude Code ships: that one is extracted + * from a proprietary binary, and the public collections of it are either unlicensed or + * CC BY-NC-SA — neither is compatible with this project's MIT licence or with REUSE. Edit the + * resource to change them; no Java involved. + * + * @return one word, or {@code "Thinking"} when the list cannot be read + */ + static String spinnerWord() { + List words = prompt(SPINNER_WORDS) + .lines() + .map(String::strip) + .filter(word -> !word.isEmpty()) + .toList(); + return words.isEmpty() + ? "Thinking" + : words.get(java.util.concurrent.ThreadLocalRandom.current().nextInt(words.size())); } /** diff --git a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/spinner-words.txt b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/spinner-words.txt new file mode 100644 index 000000000..e16bfe437 --- /dev/null +++ b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/spinner-words.txt @@ -0,0 +1,26 @@ +Fettling +Whittling +Rummaging +Squirrelling +Beavering +Trundling +Ambling +Pootling +Cobbling +Kerfuffling +Whirligigging +Bustling +Beetling +Scuttling +Lumbering +Chugging +Puffing +Clattering +Burbling +Purring +Cranking +Rejigging +Tootling +Fossicking +Wombling +Head-scratching diff --git a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/spinner-words.txt.license b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/spinner-words.txt.license new file mode 100644 index 000000000..b918686f7 --- /dev/null +++ b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/spinner-words.txt.license @@ -0,0 +1,3 @@ +SPDX-FileCopyrightText: 2026 Bernard Ladenthin + +SPDX-License-Identifier: MIT diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java index 17d6a90f8..d53f4c08f 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java @@ -244,11 +244,35 @@ void aToolCallStaysInTheHistorySoTheNextTurnSeesItHappened() throws Exception { @Test void theActivityLineNamesTheRunningToolSoALongBuildLooksAlive() { - // "working… (90s)" during a two-minute mvn test is indistinguishable from a hang. + // "working… (90s)" during a two-minute mvn test is indistinguishable from a hang, so the line + // names the tool and how long IT has been running, next to the turn's total. assertThat( - LocalAgent.activityLine('x', 12, "run_command", 9, 2), is("x run_command… (9s of 12s · 2 tool calls)")); - assertThat(LocalAgent.activityLine('x', 5, null, 0, 0), is("x thinking… (5s)")); - assertThat(LocalAgent.activityLine('x', 30, null, 0, 3), is("x thinking… (30s · 3 tool calls)")); + LocalAgent.activityLine('x', "Fettling", 12, "run_command", 9, 2), + is("x Fettling… (run_command 9s of 12s · 2 tool calls)")); + assertThat(LocalAgent.activityLine('x', "Fettling", 5, null, 0, 0), is("x Fettling… (5s)")); + assertThat(LocalAgent.activityLine('x', "Fettling", 30, null, 0, 3), is("x Fettling… (30s · 3 tool calls)")); + } + + @Test + void theSpinnerWordsAreOursAndHarmless() { + List words = LocalAgent.prompt(LocalAgent.SPINNER_WORDS) + .lines() + .map(String::strip) + .filter(word -> !word.isEmpty()) + .toList(); + + assertThat("enough variety to not repeat every other turn", words.size() > 15, is(true)); + assertThat("no duplicates", words.size(), is((int) + words.stream().distinct().count())); + for (String word : words) { + assertThat(word, word.matches("[A-Z][a-z-]+"), is(true)); + } + // Claude Code's own list is extracted from a proprietary binary and the public copies of it are + // unlicensed or CC BY-NC-SA; none of its words may appear here. + for (String theirs : List.of("Razzmatazzing", "Clauding", "Flibbertigibbeting", "Simmering", "Vibing")) { + assertThat(words.contains(theirs), is(false)); + } + assertThat(words.contains(LocalAgent.spinnerWord()), is(true)); } @Test From b4fd0552afe56c831ffba246d45c6e7d4126daa6 Mon Sep 17 00:00:00 2001 From: Bernard Ladenthin Date: Wed, 23 Sep 2026 14:29:48 +0200 Subject: [PATCH 09/36] llama-atmosphere-agent: add the missing REUSE sidecars for the /loop resources loop-prompt.txt and loop-file-template.md shipped without a .license file, which fails the REUSE Compliance Check on main -- the same mistake as the agent's .mvn/jvm.config, caught this time by running `reuse lint` before pushing. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01E2h8gXyyE5UeimkL9vQv9G --- .../ladenthin/llama/atmosphere/loop-file-template.md.license | 3 +++ .../net/ladenthin/llama/atmosphere/loop-prompt.txt.license | 3 +++ 2 files changed, 6 insertions(+) create mode 100644 llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/loop-file-template.md.license create mode 100644 llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/loop-prompt.txt.license diff --git a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/loop-file-template.md.license b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/loop-file-template.md.license new file mode 100644 index 000000000..b918686f7 --- /dev/null +++ b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/loop-file-template.md.license @@ -0,0 +1,3 @@ +SPDX-FileCopyrightText: 2026 Bernard Ladenthin + +SPDX-License-Identifier: MIT diff --git a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/loop-prompt.txt.license b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/loop-prompt.txt.license new file mode 100644 index 000000000..b918686f7 --- /dev/null +++ b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/loop-prompt.txt.license @@ -0,0 +1,3 @@ +SPDX-FileCopyrightText: 2026 Bernard Ladenthin + +SPDX-License-Identifier: MIT From 6172e12eeb6365cbca67ff72e3746a9ec3f97fde Mon Sep 17 00:00:00 2001 From: Bernard Ladenthin Date: Wed, 23 Sep 2026 14:38:07 +0200 Subject: [PATCH 10/36] llama-atmosphere-agent: print the status above the prompt instead of pinning it The status line was invisible in practice. JLine's Status pins to the bottom of the WINDOW, so in a tall terminal with a few lines of output the prompt sits near the top and the status sits at the bottom edge, visually disconnected -- and it never appears when the scrollback is copied, because it is not in the scrollback. It is now printed as an ordinary dimmed line right above each prompt. The pinned area keeps its one good use: the activity line while a turn runs, where the streamed output fills the screen down to it. PlainTerminal.status() is now a no-op: the caller prints the line itself, and a spinner refreshed four times a second would otherwise write four lines a second into a piped log. Tests: 106 green. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01E2h8gXyyE5UeimkL9vQv9G --- llama-atmosphere-agent/README.md | 6 ++++++ .../llama/atmosphere/LocalAgent.java | 21 ++++++++++++------- .../llama/atmosphere/PlainTerminal.java | 7 ++----- .../llama/atmosphere/PlainTerminalTest.java | 14 +++++++------ 4 files changed, 29 insertions(+), 19 deletions(-) diff --git a/llama-atmosphere-agent/README.md b/llama-atmosphere-agent/README.md index dca0093f9..d283007cc 100644 --- a/llama-atmosphere-agent/README.md +++ b/llama-atmosphere-agent/README.md @@ -278,6 +278,12 @@ output **line by line while it runs** (dimmed, `│ `-prefixed) instead of dumpi also keeps the pipe drained; a process whose output nobody reads blocks once the buffer is full, and on Windows that buffer is about 4 KB. +**Where the two lines live.** The status is printed **above the prompt**, as an ordinary line, every +time you are asked for input — JLine can pin a line to the bottom of the window, but in a tall +terminal with little output that bottom edge is nowhere near the cursor and the line goes unnoticed. +The pinned area is used only while a turn runs, where it works well because the streamed output fills +the screen right down to it. + **The status line** above the prompt reads `[manual · ctx ~3.1k/16k · 9 tools · local-model]`: the approval mode, the context used out of the window, the number of tools and the model id. A `~` means the number is an estimate from the text diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java index 331295378..4233c77d3 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java @@ -214,14 +214,19 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream int turnNumber = 0; String pendingNote = ""; while (true) { - terminal.status(StatusLine.render( - options.getWorkspace(), - mode.get(), - inputTokens, - estimated, - contextSize, - tools.size(), - options.getModelId())); + // Printed right above the prompt, not pinned: JLine pins to the bottom of the WINDOW, + // which in a tall terminal with little output sits far below the cursor and is easy to + // miss entirely. The pinned area is worth more while a turn runs (see awaitWithActivity), + // because output then fills the screen down to it. + terminal.line(terminal.ansi() + .dim(StatusLine.render( + options.getWorkspace(), + mode.get(), + inputTokens, + estimated, + contextSize, + tools.size(), + options.getModelId()))); String line = terminal.readLine("you> "); if (line == null) { return 0; diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/PlainTerminal.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/PlainTerminal.java index f5d6deb06..9d0008b6b 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/PlainTerminal.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/PlainTerminal.java @@ -23,7 +23,6 @@ public final class PlainTerminal implements AgentTerminal { private final PrintStream out; private final @Nullable BufferedReader in; private final Ansi ansi; - private String status = ""; /** * Create a plain console. @@ -46,9 +45,6 @@ public void line(String text) { @Override public @Nullable String readLine(String prompt) { - if (!status.isEmpty()) { - out.println(ansi.dim(status)); - } out.print(prompt); out.flush(); return read(); @@ -64,7 +60,8 @@ public void line(String text) { @Override public void status(String text) { - this.status = text; + // Nothing can be pinned on a plain stream, and the caller already prints the status line above + // the prompt. Dropping it here is what keeps a piped session free of half-drawn spinner lines. } @Override diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/PlainTerminalTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/PlainTerminalTest.java index b7643d0a1..e9244b45b 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/PlainTerminalTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/PlainTerminalTest.java @@ -8,6 +8,7 @@ import static org.hamcrest.Matchers.containsString; import static org.hamcrest.Matchers.hasItem; import static org.hamcrest.Matchers.is; +import static org.hamcrest.Matchers.not; import static org.hamcrest.Matchers.nullValue; import java.io.BufferedReader; @@ -48,16 +49,17 @@ void linesAreWrittenAndInputIsReadBackLineByLine() { } @Test - void theStatusLineIsPrintedBeforeThePromptBecauseNothingCanBePinned() { + void aPinnedStatusIsDroppedRatherThanPrintedRepeatedly() { + // Nothing can be pinned on a plain stream: a spinner refreshed four times a second would + // otherwise produce four lines a second in a piped log. The caller prints the status itself, + // once, above the prompt. PlainTerminal terminal = terminal("x" + System.lineSeparator()); - terminal.status("[manual · ctx 0/16k]"); + terminal.status("⠙ Fettling… (5s)"); terminal.readLine("you> "); - assertThat(written(), containsString("[manual · ctx 0/16k]")); - assertThat( - written().indexOf("[manual"), - is(org.hamcrest.Matchers.lessThan(written().indexOf("you> ")))); + assertThat(written(), not(containsString("Fettling"))); + assertThat(written(), containsString("you> ")); } @Test From 311dbec7272f83f992f678523b13ef8a5388b343 Mon Sep 17 00:00:00 2001 From: Bernard Ladenthin Date: Wed, 23 Sep 2026 14:47:24 +0200 Subject: [PATCH 11/36] llama-atmosphere-agent: pin the status again on a real terminal The previous commit moved it above the prompt, which removed the fixed bottom block the terminal is supposed to have. It is pinned again wherever pinning works, and printed above the prompt only where it does not (plain stream: piped input, one-shot runs, tests) -- decided by the new AgentTerminal.pinsStatus(). The gap between the prompt and the pinned block in a tall, mostly empty window is inherent to a line-oriented REPL: JLine pins to the bottom of the window, not below the cursor. Documented rather than worked around. Tests: 106 green. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01E2h8gXyyE5UeimkL9vQv9G --- llama-atmosphere-agent/README.md | 11 +++---- .../llama/atmosphere/AgentTerminal.java | 9 ++++++ .../llama/atmosphere/JLineTerminal.java | 5 ++++ .../llama/atmosphere/LocalAgent.java | 29 ++++++++++--------- .../llama/atmosphere/PlainTerminal.java | 5 ++++ 5 files changed, 41 insertions(+), 18 deletions(-) diff --git a/llama-atmosphere-agent/README.md b/llama-atmosphere-agent/README.md index d283007cc..c92765705 100644 --- a/llama-atmosphere-agent/README.md +++ b/llama-atmosphere-agent/README.md @@ -278,11 +278,12 @@ output **line by line while it runs** (dimmed, `│ `-prefixed) instead of dumpi also keeps the pipe drained; a process whose output nobody reads blocks once the buffer is full, and on Windows that buffer is about 4 KB. -**Where the two lines live.** The status is printed **above the prompt**, as an ordinary line, every -time you are asked for input — JLine can pin a line to the bottom of the window, but in a tall -terminal with little output that bottom edge is nowhere near the cursor and the line goes unnoticed. -The pinned area is used only while a turn runs, where it works well because the streamed output fills -the screen right down to it. +**Where the two lines live.** On a real terminal both are pinned to the bottom of the window, below a +rule: the status while you type, the activity while a turn runs. Note where "the bottom" is — the +bottom of the *window*, not the line under the cursor; in a tall terminal with only a few lines of +output there is a gap between your prompt and the pinned block. On a plain stream (piped input, a +one-shot run) nothing can be pinned, so the status is printed above the prompt instead and the +activity line is dropped rather than repeated into the log. **The status line** above the prompt reads `[manual · ctx ~3.1k/16k · 9 tools · local-model]`: the approval mode, the context used out of the diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentTerminal.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentTerminal.java index 3f08134ba..ca93d2bdf 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentTerminal.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentTerminal.java @@ -54,6 +54,15 @@ public interface AgentTerminal extends AutoCloseable { */ void status(String text); + /** + * Whether {@link #status} really pins the line to the bottom of the window. + * + *

{@code false} on a plain stream, where the caller has to print the status itself. + * + * @return {@code true} on a real terminal + */ + boolean pinsStatus(); + /** * The styles to use for this console. * diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java index 03501356f..fc170aa9a 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java @@ -128,6 +128,11 @@ public void status(String text) { new AttributedString(text, AttributedStyle.DEFAULT.foreground(AttributedStyle.BRIGHT)))); } + @Override + public boolean pinsStatus() { + return true; + } + @Override public Ansi ansi() { return ansi; diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java index 4233c77d3..427d15f50 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java @@ -214,19 +214,22 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream int turnNumber = 0; String pendingNote = ""; while (true) { - // Printed right above the prompt, not pinned: JLine pins to the bottom of the WINDOW, - // which in a tall terminal with little output sits far below the cursor and is easy to - // miss entirely. The pinned area is worth more while a turn runs (see awaitWithActivity), - // because output then fills the screen down to it. - terminal.line(terminal.ansi() - .dim(StatusLine.render( - options.getWorkspace(), - mode.get(), - inputTokens, - estimated, - contextSize, - tools.size(), - options.getModelId()))); + // Pinned to the bottom of the window on a real terminal; printed above the prompt on a + // plain stream, where there is nothing to pin and a repeatedly refreshed line would + // just fill a piped log. + String status = StatusLine.render( + options.getWorkspace(), + mode.get(), + inputTokens, + estimated, + contextSize, + tools.size(), + options.getModelId()); + if (terminal.pinsStatus()) { + terminal.status(status); + } else { + terminal.line(terminal.ansi().dim(status)); + } String line = terminal.readLine("you> "); if (line == null) { return 0; diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/PlainTerminal.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/PlainTerminal.java index 9d0008b6b..23d8526e5 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/PlainTerminal.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/PlainTerminal.java @@ -64,6 +64,11 @@ public void status(String text) { // the prompt. Dropping it here is what keeps a piped session free of half-drawn spinner lines. } + @Override + public boolean pinsStatus() { + return false; + } + @Override public Ansi ansi() { return ansi; From daaf47e63dcb150d665037af077aadf014bb3ae1 Mon Sep 17 00:00:00 2001 From: Bernard Ladenthin Date: Wed, 23 Sep 2026 14:52:37 +0200 Subject: [PATCH 12/36] llama-atmosphere-agent: two-row status block, so the spinner word is always visible MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The block showed either the state line or the activity line, never both -- so the spinner words only existed during a turn, in the place the eye was not, and the block changed height whenever it switched. It now has two fixed rows under the rule: what the agent is doing ("… waiting for input …" or "⠙ Fettling… (run_command 47s of 61s · 2 tool calls)") and the session state (workspace · mode · ctx · tools · model). Fixed height means the output above it does not jump on a refresh. AgentTerminal.status takes a list of lines instead of one string; JLineTerminal renders rule + rows, PlainTerminal still ignores it. Tests: 106 green. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01E2h8gXyyE5UeimkL9vQv9G --- llama-atmosphere-agent/README.md | 23 ++++++++--- .../llama/atmosphere/AgentTerminal.java | 10 +++-- .../llama/atmosphere/JLineTerminal.java | 17 +++++---- .../llama/atmosphere/LocalAgent.java | 38 +++++++++++-------- .../llama/atmosphere/PlainTerminal.java | 2 +- .../ladenthin/llama/atmosphere/TaskLoop.java | 3 +- .../llama/atmosphere/PlainTerminalTest.java | 2 +- 7 files changed, 61 insertions(+), 34 deletions(-) diff --git a/llama-atmosphere-agent/README.md b/llama-atmosphere-agent/README.md index c92765705..0a2d206ac 100644 --- a/llama-atmosphere-agent/README.md +++ b/llama-atmosphere-agent/README.md @@ -278,12 +278,23 @@ output **line by line while it runs** (dimmed, `│ `-prefixed) instead of dumpi also keeps the pipe drained; a process whose output nobody reads blocks once the buffer is full, and on Windows that buffer is about 4 KB. -**Where the two lines live.** On a real terminal both are pinned to the bottom of the window, below a -rule: the status while you type, the activity while a turn runs. Note where "the bottom" is — the -bottom of the *window*, not the line under the cursor; in a tall terminal with only a few lines of -output there is a gap between your prompt and the pinned block. On a plain stream (piped input, a -one-shot run) nothing can be pinned, so the status is printed above the prompt instead and the -activity line is dropped rather than repeated into the log. +**The block at the bottom has two rows**, below a rule, and both are always present: + +``` +──────────────────────────────────────────────────────────────────── +⠙ Fettling… (run_command 47s of 61s · 2 tool calls) +[/path/to/project · manual · ctx ~3.1k/16k · 9 tools · local-model] +``` + +The first row is what the agent is doing — `… waiting for input …` when it is your turn, the spinner +with the running tool and its elapsed time while it works. The second row is the session's state. Two +fixed rows rather than one changing one: a block that changes height makes the output above it jump on +every refresh. + +Note where "the bottom" is: the bottom of the *window*, not the line under the cursor. In a tall +terminal with only a few lines of output there is a gap between your prompt and the block. On a plain +stream (piped input, a one-shot run) nothing can be pinned, so the state line is printed above the +prompt instead and the activity row is dropped rather than repeated into the log. **The status line** above the prompt reads `[manual · ctx ~3.1k/16k · 9 tools · local-model]`: the approval mode, the context used out of the diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentTerminal.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentTerminal.java index ca93d2bdf..8da5e07e8 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentTerminal.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentTerminal.java @@ -48,11 +48,15 @@ public interface AgentTerminal extends AutoCloseable { String readKey(String prompt); /** - * Set the status line kept at the bottom of the window. + * Set the block kept at the bottom of the window. * - * @param text the line; an empty string removes it + *

Two lines in practice: what the agent is doing right now, and the session's state. Keeping + * both there at all times is what stops the block from changing height, which would make the + * output above it jump on every update. + * + * @param lines the lines, top to bottom; an empty list removes the block */ - void status(String text); + void status(java.util.List lines); /** * Whether {@link #status} really pins the line to the bottom of the window. diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java index fc170aa9a..0abe0e2b7 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java @@ -115,17 +115,20 @@ public void line(String text) { } @Override - public void status(String text) { - if (text.isEmpty()) { + public void status(List lines) { + if (lines.isEmpty()) { status.update(List.of()); return; } - // A rule above the status line separates the live block from the scrollback, the way the - // established terminal agents frame their input. + // A rule above the block separates it from the scrollback, the way the established terminal + // agents frame their input. int width = Math.max(10, terminal.getSize().getColumns()); - status.update(List.of( - new AttributedString("─".repeat(width), AttributedStyle.DEFAULT.foreground(AttributedStyle.BRIGHT)), - new AttributedString(text, AttributedStyle.DEFAULT.foreground(AttributedStyle.BRIGHT)))); + List block = new java.util.ArrayList<>(); + block.add(new AttributedString("─".repeat(width), AttributedStyle.DEFAULT.foreground(AttributedStyle.BRIGHT))); + for (String line : lines) { + block.add(new AttributedString(line, AttributedStyle.DEFAULT.foreground(AttributedStyle.BRIGHT))); + } + status.update(block); } @Override diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java index 427d15f50..85d1d1135 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java @@ -84,6 +84,9 @@ public final class LocalAgent { /** The spinner shown in the activity line. */ private static final String ACTIVITY_FRAMES = "⠋⠙⠹⠸⠼⠴⠦⠧⠇⠏"; + /** The first line of the block while nothing is running. */ + static final String IDLE_LINE = "… waiting for input …"; + /** The whimsical words the activity line picks from, one per turn. */ static final String SPINNER_WORDS = "spinner-words.txt"; @@ -198,7 +201,7 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream : ServerProps.contextSize(baseUrl, options.getApiKey()); if (options.getPrompt() != null) { - return turn(runner, fileSystem, options.getPrompt(), history, terminal, callLog, 1) + return turn(runner, fileSystem, options.getPrompt(), history, terminal, callLog, 1, "") .failure() == null ? 0 @@ -226,7 +229,7 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream tools.size(), options.getModelId()); if (terminal.pinsStatus()) { - terminal.status(status); + terminal.status(List.of(IDLE_LINE, status)); } else { terminal.line(terminal.ansi().dim(status)); } @@ -266,7 +269,8 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream history, terminal, callLog, - turnNumber); + turnNumber, + status); pendingNote = toolNote(completed.rounds()); estimated = completed.inputTokens() == 0; inputTokens = estimated ? estimateTokens(systemPrompt(options), history) : completed.inputTokens(); @@ -353,11 +357,12 @@ static ConsoleSession turn( List history, AgentTerminal terminal, ToolCallLog callLog, - int turnNumber) + int turnNumber, + String stateLine) throws InterruptedException { ConsoleSession session = new ConsoleSession(terminal, fileSystem); runner.run(message, history, session); - boolean finished = awaitWithActivity(session, terminal); + boolean finished = awaitWithActivity(session, terminal, stateLine); history.add(ChatMessage.user(message)); callLog.add(turnNumber, session.rounds()); if (!session.text().isEmpty()) { @@ -504,10 +509,11 @@ private static long handleCommand( * * @param session the running turn * @param terminal the console + * @param stateLine the second line of the block, kept in place so it does not change height * @return {@code true} when the turn finished within {@link #TURN_TIMEOUT} * @throws InterruptedException if interrupted while waiting */ - private static boolean awaitWithActivity(ConsoleSession session, AgentTerminal terminal) + private static boolean awaitWithActivity(ConsoleSession session, AgentTerminal terminal, String stateLine) throws InterruptedException { long start = System.nanoTime(); int frame = 0; @@ -518,15 +524,17 @@ private static boolean awaitWithActivity(ConsoleSession session, AgentTerminal t if (seconds > TURN_TIMEOUT.toSeconds()) { return false; } - terminal.status(activityLine( - ACTIVITY_FRAMES.charAt(frame++ % ACTIVITY_FRAMES.length()), - word, - seconds, - session.runningTool(), - session.runningSeconds(), - session.toolCalls())); + terminal.status(List.of( + activityLine( + ACTIVITY_FRAMES.charAt(frame++ % ACTIVITY_FRAMES.length()), + word, + seconds, + session.runningTool(), + session.runningSeconds(), + session.toolCalls()), + stateLine)); } - terminal.status(""); + terminal.status(List.of(IDLE_LINE, stateLine)); return true; } @@ -625,7 +633,7 @@ private static long compact( terminal.line("(compacting " + before + " messages …)"); ConsoleSession session = new ConsoleSession(terminal, fileSystem); runner.runWithoutTools(instructions, List.copyOf(history), session, COMPACT_SYSTEM_PROMPT); - if (!awaitWithActivity(session, terminal) || session.text().isBlank()) { + if (!awaitWithActivity(session, terminal, "/compact") || session.text().isBlank()) { terminal.line("(compact failed; history kept)"); return 0; } diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/PlainTerminal.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/PlainTerminal.java index 23d8526e5..1312f39ef 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/PlainTerminal.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/PlainTerminal.java @@ -59,7 +59,7 @@ public void line(String text) { } @Override - public void status(String text) { + public void status(java.util.List lines) { // Nothing can be pinned on a plain stream, and the caller already prints the status line above // the prompt. Dropping it here is what keeps a piped session free of half-drawn spinner lines. } diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/TaskLoop.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/TaskLoop.java index 579e53b17..1c8168af1 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/TaskLoop.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/TaskLoop.java @@ -188,7 +188,8 @@ public static Outcome run( new java.util.ArrayList<>(), terminal, callLog, - step); + step, + "loop step " + step + "/" + options.maxSteps() + " · " + options.task()); extra = ""; if (session.failure() != null) { return new Outcome("step " + step + " failed: " + session.failure(), false); diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/PlainTerminalTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/PlainTerminalTest.java index e9244b45b..8f64e9987 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/PlainTerminalTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/PlainTerminalTest.java @@ -54,7 +54,7 @@ void aPinnedStatusIsDroppedRatherThanPrintedRepeatedly() { // otherwise produce four lines a second in a piped log. The caller prints the status itself, // once, above the prompt. PlainTerminal terminal = terminal("x" + System.lineSeparator()); - terminal.status("⠙ Fettling… (5s)"); + terminal.status(java.util.List.of("⠙ Fettling… (5s)", "[state]")); terminal.readLine("you> "); From ae899c3d9f2d419c7a2cc6db43dba3c248b1b3a1 Mon Sep 17 00:00:00 2001 From: Bernard Ladenthin Date: Wed, 23 Sep 2026 14:59:44 +0200 Subject: [PATCH 13/36] llama-atmosphere-agent: run the turn on its own thread, so the activity line moves MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The block sat on "… waiting for input …" for the whole turn and the spinner words were never seen. The cause was not the rendering: Atmosphere's execute() is synchronous and returns only once the turn including every tool round has finished, so the activity loop started when there was nothing left to show. - turn() and compact() now run the runtime on a daemon thread and drive the status block from the console thread while it works. - TurnActivity pauses the redraw while the approval prompt is open: that prompt reads a single key in raw mode on the worker thread, and a status redraw arriving mid-read writes escape sequences across the question. Tests: 107 green (was 106), including that the prompt hands the terminal back. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01E2h8gXyyE5UeimkL9vQv9G --- CLAUDE.md | 9 +++- .../atmosphere/ConsoleApprovalStrategy.java | 15 +++++- .../llama/atmosphere/LocalAgent.java | 54 +++++++++++++++---- .../ladenthin/llama/atmosphere/TaskLoop.java | 3 +- .../llama/atmosphere/TurnActivity.java | 41 ++++++++++++++ .../llama/atmosphere/ApprovalWireTest.java | 4 +- .../ConsoleApprovalStrategyTest.java | 16 +++++- 7 files changed, 127 insertions(+), 15 deletions(-) create mode 100644 llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/TurnActivity.java diff --git a/CLAUDE.md b/CLAUDE.md index c72686959..118a1e96f 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -2322,8 +2322,13 @@ are decisions, not details: user message**, which every template accepts. Pinned by `LocalAgentTest.aToolCallStaysInTheHistorySoTheNextTurnSeesItHappened`. **Live feedback while a turn runs** (`LocalAgent.activityLine`, `ShellTool`'s line-by-line output): - the pinned line names the running tool and its own elapsed time, and shell output is printed as it - arrives. Reading the pipe incrementally is not only cosmetic — an unread pipe blocks the child once + the block's first row names the running tool and its own elapsed time, and shell output is printed + as it arrives. **The turn runs on its own thread, and it has to:** Atmosphere's `execute()` is + synchronous — it returns only once the whole turn including every tool round is done — so running + it on the console thread leaves nobody to refresh the line, and the block sits on + "… waiting for input …" for the entire turn (exactly the symptom that was reported). The approval + prompt then reads a key in raw mode on that worker thread while the console thread redraws four + times a second, so `TurnActivity` pauses the redraw for as long as the question is open. Reading the pipe incrementally is not only cosmetic — an unread pipe blocks the child once it is full, which on Windows is roughly 4 KB. 8. **The context number in the status line is an estimate, marked `~`.** llama.cpp emits its usage chunk only when the client sets `stream_options.include_usage`, and Atmosphere's client does not; diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleApprovalStrategy.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleApprovalStrategy.java index 8b0828de5..e7f806b3b 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleApprovalStrategy.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleApprovalStrategy.java @@ -48,6 +48,7 @@ public final class ConsoleApprovalStrategy implements ApprovalStrategy { private final AgentTerminal terminal; private final boolean interactive; private final Ansi ansi; + private final TurnActivity activity; /** * Create the strategy. @@ -56,12 +57,15 @@ public final class ConsoleApprovalStrategy implements ApprovalStrategy { * {@code [a]} answer) * @param terminal where the question is asked * @param interactive whether anybody can answer at all ({@code false} for a one-shot run) + * @param activity paused while the question is open, so the spinner does not redraw over it */ - public ConsoleApprovalStrategy(AtomicReference mode, AgentTerminal terminal, boolean interactive) { + public ConsoleApprovalStrategy( + AtomicReference mode, AgentTerminal terminal, boolean interactive, TurnActivity activity) { this.mode = mode; this.terminal = terminal; this.interactive = interactive; this.ansi = terminal.ansi(); + this.activity = activity; } /** @@ -99,6 +103,15 @@ public ApprovalResolution awaitApprovalDetailed(PendingApproval approval, Stream return ApprovalResolution.deny(); } terminal.line(ansi.yellow("? " + approval.toolName()) + " " + ansi.dim(preview(approval))); + activity.pause(); + try { + return ask(approval); + } finally { + activity.resume(); + } + } + + private ApprovalResolution ask(PendingApproval approval) { while (true) { String answer = terminal.readKey(ansi.yellow(" allow? [y]es / [n]o / [a]uto (no more questions): ")); if (answer == null) { diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java index 85d1d1135..fb39e3109 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java @@ -195,13 +195,16 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream } // One-shot runs have nobody at the keyboard, so the strategy gets no console and denies // gated calls unless --auto was passed (see ConsoleApprovalStrategy). - runner.approval(new ConsoleApprovalStrategy(mode, terminal, interactive), ConsoleApprovalStrategy.policy()); + TurnActivity activity = new TurnActivity(); + runner.approval( + new ConsoleApprovalStrategy(mode, terminal, interactive, activity), + ConsoleApprovalStrategy.policy()); int contextSize = options.getModelPath() != null ? options.getCtxSize() : ServerProps.contextSize(baseUrl, options.getApiKey()); if (options.getPrompt() != null) { - return turn(runner, fileSystem, options.getPrompt(), history, terminal, callLog, 1, "") + return turn(runner, fileSystem, options.getPrompt(), history, terminal, callLog, 1, "", activity) .failure() == null ? 0 @@ -270,7 +273,8 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream terminal, callLog, turnNumber, - status); + status, + activity); pendingNote = toolNote(completed.rounds()); estimated = completed.inputTokens() == 0; inputTokens = estimated ? estimateTokens(systemPrompt(options), history) : completed.inputTokens(); @@ -358,11 +362,27 @@ static ConsoleSession turn( AgentTerminal terminal, ToolCallLog callLog, int turnNumber, - String stateLine) + String stateLine, + TurnActivity activity) throws InterruptedException { ConsoleSession session = new ConsoleSession(terminal, fileSystem); - runner.run(message, history, session); - boolean finished = awaitWithActivity(session, terminal, stateLine); + // Atmosphere's execute() is synchronous: it returns only once the whole turn, tool rounds + // included, has finished. Running it on this thread would leave nobody to drive the activity + // line -- which is exactly what happened: the block sat on "… waiting for input …" for the + // entire turn, so the spinner and its word were never seen. + Thread worker = new Thread( + () -> { + try { + runner.run(message, history, session); + } catch (RuntimeException e) { + session.error(e); + } + }, + "agent-turn"); + worker.setDaemon(true); + worker.start(); + boolean finished = awaitWithActivity(session, terminal, stateLine, activity); + worker.join(Duration.ofSeconds(5).toMillis()); history.add(ChatMessage.user(message)); callLog.add(turnNumber, session.rounds()); if (!session.text().isEmpty()) { @@ -510,10 +530,12 @@ private static long handleCommand( * @param session the running turn * @param terminal the console * @param stateLine the second line of the block, kept in place so it does not change height + * @param activity paused while an approval question is open * @return {@code true} when the turn finished within {@link #TURN_TIMEOUT} * @throws InterruptedException if interrupted while waiting */ - private static boolean awaitWithActivity(ConsoleSession session, AgentTerminal terminal, String stateLine) + private static boolean awaitWithActivity( + ConsoleSession session, AgentTerminal terminal, String stateLine, TurnActivity activity) throws InterruptedException { long start = System.nanoTime(); int frame = 0; @@ -524,6 +546,9 @@ private static boolean awaitWithActivity(ConsoleSession session, AgentTerminal t if (seconds > TURN_TIMEOUT.toSeconds()) { return false; } + if (activity.isPaused()) { + continue; // an approval question owns the terminal until it is answered + } terminal.status(List.of( activityLine( ACTIVITY_FRAMES.charAt(frame++ % ACTIVITY_FRAMES.length()), @@ -632,8 +657,19 @@ private static long compact( int before = history.size(); terminal.line("(compacting " + before + " messages …)"); ConsoleSession session = new ConsoleSession(terminal, fileSystem); - runner.runWithoutTools(instructions, List.copyOf(history), session, COMPACT_SYSTEM_PROMPT); - if (!awaitWithActivity(session, terminal, "/compact") || session.text().isBlank()) { + Thread worker = new Thread( + () -> { + try { + runner.runWithoutTools(instructions, List.copyOf(history), session, COMPACT_SYSTEM_PROMPT); + } catch (RuntimeException e) { + session.error(e); + } + }, + "agent-compact"); + worker.setDaemon(true); + worker.start(); + if (!awaitWithActivity(session, terminal, "/compact", new TurnActivity()) + || session.text().isBlank()) { terminal.line("(compact failed; history kept)"); return 0; } diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/TaskLoop.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/TaskLoop.java index 1c8168af1..a9f9c6866 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/TaskLoop.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/TaskLoop.java @@ -189,7 +189,8 @@ public static Outcome run( terminal, callLog, step, - "loop step " + step + "/" + options.maxSteps() + " · " + options.task()); + "loop step " + step + "/" + options.maxSteps() + " · " + options.task(), + new TurnActivity()); extra = ""; if (session.failure() != null) { return new Outcome("step " + step + " failed: " + session.failure(), false); diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/TurnActivity.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/TurnActivity.java new file mode 100644 index 000000000..18b18d1a7 --- /dev/null +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/TurnActivity.java @@ -0,0 +1,41 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import java.util.concurrent.atomic.AtomicBoolean; + +/** + * A switch that stops the activity line from redrawing while something else owns the terminal. + * + *

It exists because a turn runs on its own thread. Atmosphere's {@code execute} is synchronous — + * it returns only when the whole turn including every tool round is done — so the console thread has + * to drive the spinner while a second thread runs the turn. That is fine until the approval prompt + * appears: it reads a single key in raw mode on the turn's thread, and a status redraw arriving from + * the console thread in the middle of that writes escape sequences across the question. So the prompt + * pauses the redraw for as long as it is waiting for an answer. + */ +public final class TurnActivity { + + private final AtomicBoolean paused = new AtomicBoolean(); + + /** Stop redrawing the activity line. */ + public void pause() { + paused.set(true); + } + + /** Redraw it again. */ + public void resume() { + paused.set(false); + } + + /** + * Whether redrawing is currently suspended. + * + * @return {@code true} while something else owns the terminal + */ + public boolean isPaused() { + return paused.get(); + } +} diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ApprovalWireTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ApprovalWireTest.java index b230fd21b..10df5ee0e 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ApprovalWireTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ApprovalWireTest.java @@ -71,7 +71,9 @@ private AgentRunner runner(OpenAiCompatServer server, List tools 64, 10) .retryPolicy(RetryPolicy.NONE) - .approval(new ConsoleApprovalStrategy(mode, terminal(typed), true), ConsoleApprovalStrategy.policy()); + .approval( + new ConsoleApprovalStrategy(mode, terminal(typed), true, new TurnActivity()), + ConsoleApprovalStrategy.policy()); } private ConsoleSession session() { diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleApprovalStrategyTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleApprovalStrategyTest.java index c223b9ee5..4b99f0b78 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleApprovalStrategyTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleApprovalStrategyTest.java @@ -26,6 +26,7 @@ class ConsoleApprovalStrategyTest { private final ByteArrayOutputStream console = new ByteArrayOutputStream(); + private final TurnActivity activity = new TurnActivity(); private PendingApproval approval() { return new PendingApproval( @@ -42,7 +43,7 @@ private ApprovalOutcome ask(AtomicReference mode, String typed) { AgentTerminal terminal = new PlainTerminal(new PrintStream(console, true, StandardCharsets.UTF_8), reader, Ansi.PLAIN); // The strategy never touches the session; Atmosphere passes it only so a UI can emit events. - return new ConsoleApprovalStrategy(mode, terminal, typed != null).awaitApproval(approval(), null); + return new ConsoleApprovalStrategy(mode, terminal, typed != null, activity).awaitApproval(approval(), null); } private String consoleText() { @@ -118,6 +119,19 @@ void writingToolsAndTheShellAreGatedReadingToolsAreNot() { is(java.util.List.of("write_file", ShellTool.TOOL_NAME))); } + @Test + void theSpinnerIsPausedWhileTheQuestionIsOpen() { + // The turn runs on its own thread while the console thread redraws the status block four times + // a second. A redraw arriving in the middle of a raw-mode key read writes escape sequences + // across the question, so the prompt owns the terminal until it has an answer. + assertThat(activity.isPaused(), is(false)); + + ask(new AtomicReference<>(ApprovalMode.MANUAL), "y" + System.lineSeparator()); + + assertThat("and hands it back afterwards", activity.isPaused(), is(false)); + assertThat(consoleText(), containsString("allow?")); + } + private static ToolDefinition stub(String name) { return ToolDefinition.builder(name, "test").executor(args -> "").build(); } From 18ed0f2264bdb0163485f9c605388a84d2b68bdb Mon Sep 17 00:00:00 2001 From: Bernard Ladenthin Date: Wed, 23 Sep 2026 16:07:49 +0200 Subject: [PATCH 14/36] llama-atmosphere-agent: keep the pinned block intact under long output Reported: a /compact summary tore the bottom block apart -- the state row appeared in the middle of the summary text, and "you>" was printed five times in a row. Two causes, both here: - A status row longer than the window wraps onto two screen lines, while JLine's reserved region is sized in lines. Everything below it is then drawn in the wrong place. Rows are now cut to the window width (and the rule is width-1, so it cannot wrap either). - line() called reader.printAbove() unconditionally. That scrolls text in above the prompt AND redraws the prompt -- so every streamed line while nobody was reading redrew a prompt that was not there, which is where the repeated "you>" came from. It now writes directly unless the line reader really owns the screen. Tests: 108 green (was 107). Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01E2h8gXyyE5UeimkL9vQv9G --- .../llama/atmosphere/JLineTerminal.java | 33 +++++++++++++++++-- .../atmosphere/ConsoleFormattingTest.java | 12 +++++++ 2 files changed, 42 insertions(+), 3 deletions(-) diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java index 0abe0e2b7..a3f8c8a73 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java @@ -40,6 +40,7 @@ public final class JLineTerminal implements AgentTerminal { private final LineReader reader; private final Status status; private final Ansi ansi; + private volatile boolean reading; private JLineTerminal(Terminal terminal, LineReader reader, Status status, Ansi ansi) { this.terminal = terminal; @@ -74,17 +75,28 @@ private JLineTerminal(Terminal terminal, LineReader reader, Status status, Ansi @Override public void line(String text) { - reader.printAbove(text); + if (reading) { + // Only while the line reader owns the screen: printAbove scrolls the text in above the + // prompt and redraws that prompt afterwards. Calling it when nobody is reading redraws a + // prompt that is not there, which is where the repeated "you>" lines came from. + reader.printAbove(text); + } else { + terminal.writer().println(text); + terminal.writer().flush(); + } } @Override public @Nullable String readLine(String prompt) { + reading = true; try { return reader.readLine(prompt); } catch (UserInterruptException e) { return ""; // Ctrl-C: drop the line, ask again } catch (EndOfFileException e) { return null; // Ctrl-D + } finally { + reading = false; } } @@ -122,15 +134,30 @@ public void status(List lines) { } // A rule above the block separates it from the scrollback, the way the established terminal // agents frame their input. - int width = Math.max(10, terminal.getSize().getColumns()); + // One row must never wrap: a wrapped row occupies two screen lines, the reserved region is + // sized in lines, and everything below it is then drawn in the wrong place -- which is how a + // long summary tore the block apart. + int width = Math.max(10, terminal.getSize().getColumns() - 1); List block = new java.util.ArrayList<>(); block.add(new AttributedString("─".repeat(width), AttributedStyle.DEFAULT.foreground(AttributedStyle.BRIGHT))); for (String line : lines) { - block.add(new AttributedString(line, AttributedStyle.DEFAULT.foreground(AttributedStyle.BRIGHT))); + block.add( + new AttributedString(fit(line, width), AttributedStyle.DEFAULT.foreground(AttributedStyle.BRIGHT))); } status.update(block); } + /** + * Cut a status row to the window width. + * + * @param text the row + * @param width how many characters fit + * @return the row, ending in {@code …} when it had to be cut + */ + static String fit(String text, int width) { + return text.length() <= width ? text : text.substring(0, Math.max(1, width - 1)) + "…"; + } + @Override public boolean pinsStatus() { return true; diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleFormattingTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleFormattingTest.java index f573d32f1..752d73d81 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleFormattingTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleFormattingTest.java @@ -138,4 +138,16 @@ void anUnclosedMarkerIsLeftAsTyped() { assertThat(render("a **b\n", Ansi.PLAIN), is("a **b" + System.lineSeparator())); assertThat(render("2 * 3 * 4\n", Ansi.PLAIN), is("2 * 3 * 4" + System.lineSeparator())); } + + // ----- the pinned block ----- + + @Test + void aStatusRowIsCutToTheWindowWidthBecauseAWrappedRowBreaksTheBlock() { + // A wrapped row takes two screen lines while the reserved region is sized in lines, so + // everything below it lands in the wrong place -- a long /compact summary tore the block apart. + assertThat(JLineTerminal.fit("short", 20), is("short")); + assertThat(JLineTerminal.fit("0123456789", 10), is("0123456789")); + assertThat(JLineTerminal.fit("0123456789x", 10), is("012345678…")); + assertThat(JLineTerminal.fit("0123456789x", 10).length(), is(10)); + } } From 6e6735d996d92c76cfb0f3aa4fdea9eb4749c468 Mon Sep 17 00:00:00 2001 From: Bernard Ladenthin Date: Wed, 23 Sep 2026 16:26:33 +0200 Subject: [PATCH 15/36] llama-atmosphere-agent: auto-compaction before the request that would overflow --auto-compact (default true) and --compact-at (default 70): when the next request would fill more than that share of the context, the history is summarized first and the user's message is then answered with the summary as its context. Before rather than after, because afterwards the oversized request has already been sent -- which is the one case worth preventing. The threshold is lower than the ~85% a hosted agent uses: our token number is usually an estimate (llama.cpp reports its own count only to clients that ask for it) and the reply still has to fit next to the prompt. With an unknown context size nothing triggers at all rather than guessing. Tests: 110 green (was 108). Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01E2h8gXyyE5UeimkL9vQv9G --- llama-atmosphere-agent/README.md | 10 +++ .../llama/atmosphere/AgentOptions.java | 63 ++++++++++++++++++- .../llama/atmosphere/LocalAgent.java | 28 +++++++++ .../net/ladenthin/llama/atmosphere/help.txt | 2 + .../llama/atmosphere/LocalAgentTest.java | 38 +++++++++++ 5 files changed, 140 insertions(+), 1 deletion(-) diff --git a/llama-atmosphere-agent/README.md b/llama-atmosphere-agent/README.md index 0a2d206ac..71865b10f 100644 --- a/llama-atmosphere-agent/README.md +++ b/llama-atmosphere-agent/README.md @@ -152,6 +152,7 @@ irrelevant: inference stays in the running server, the agent's JVM loads no mode | `--workspace

` | directory the file tools are confined to, and where `run_command` starts | cwd | | `--allow-shell` | register `run_command`: any command line, starting in the workspace | off | | `--auto` | run tools without asking (otherwise every write and command is confirmed) | off | +| `--auto-compact ` / `--compact-at ` | summarize the history before it overflows the context, and how full it may get first | `true` / `70` | | `--system ` | replace the default system prompt | built-in | | `--prompt `, `-p` | one turn, then exit | interactive | | `--temperature ` / `--max-tokens ` | sampling / per-call budget | `0.2` / `2048` | @@ -233,6 +234,15 @@ entry as `new ChatMessage(role, content)` and drops the rest). A system message cleaner, but Mistral's template requires strict user/assistant alternation and Gemma has no system role at all. +**Auto-compaction.** Once the next request would fill more than `--compact-at` percent of the context +(70 by default), the history is summarized **before that request is sent** rather than after it — the +oversized request is the one thing worth avoiding, and afterwards it has already gone out. You see +`(context nearly full — compacting first)`, then your message is answered with the summary as its +context. `--auto-compact false` turns it off; `/compact` remains available at any time. With an +unknown context size — a foreign endpoint whose `/props` answers nothing — nothing is triggered at +all rather than guessed. The threshold sits below the ~85 % a hosted agent uses because our token +number is usually an estimate and the reply still has to fit next to the prompt. + **`/compact`** asks the model to summarize the conversation (goal, facts, work done, problems, state, next step; `/compact ` adds an emphasis), then replaces the history with that summary. Use it when the context fills up. Note the history only ever held the user texts and the final answers — diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentOptions.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentOptions.java index 64634c66b..769adeb88 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentOptions.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentOptions.java @@ -37,6 +37,18 @@ public final class AgentOptions { /** Context size for the in-process model ({@code --model}). */ public static final int DEFAULT_CTX_SIZE = 8192; + /** Whether the history is summarized on its own before it overflows the context. */ + public static final boolean DEFAULT_AUTO_COMPACT = true; + + /** + * How full the context may get before that happens, in percent. + * + *

Lower than the ~85 % a hosted agent uses, and deliberately so: this number is usually an + * estimate from the text length (llama.cpp reports its own count only to clients that ask for it), + * and the reply still has to fit next to the prompt. + */ + public static final int DEFAULT_COMPACT_AT = 70; + /** * Log verbosity threshold of the in-process model ({@code --model}): llama.cpp's {@code -lv} * scale, {@code 0} output only, {@code 1} errors, {@code 2} warnings, {@code 3} info, {@code 4} @@ -57,6 +69,8 @@ public final class AgentOptions { private final Path workspace; private final boolean allowShell; private final boolean auto; + private final boolean autoCompact; + private final int compactAt; private final double temperature; private final int maxTokens; private final int maxToolRounds; @@ -76,6 +90,8 @@ private AgentOptions(Builder b) { this.workspace = b.workspace; this.allowShell = b.allowShell; this.auto = b.auto; + this.autoCompact = b.autoCompact; + this.compactAt = b.compactAt; this.temperature = b.temperature; this.maxTokens = b.maxTokens; this.maxToolRounds = b.maxToolRounds; @@ -100,6 +116,8 @@ public static AgentOptions parse(String[] args) { case "-h", "--help" -> b.help = true; case "--allow-shell" -> b.allowShell = true; case "--auto" -> b.auto = true; + case "--auto-compact" -> b.autoCompact = booleanValue(args, ++i, a); + case "--compact-at" -> b.compactAt = percentValue(args, ++i, a); case "--base-url" -> b.baseUrl = stripTrailingSlash(value(args, ++i, a)); case "--model" -> b.modelPath = value(args, ++i, a); case "--ngl", "--gpu-layers" -> b.gpuLayers = intValue(args, ++i, a); @@ -134,6 +152,25 @@ private static String value(String[] args, int index, String flag) { return args[index]; } + private static boolean booleanValue(String[] args, int index, String flag) { + String raw = value(args, index, flag).trim(); + if ("true".equalsIgnoreCase(raw) || "yes".equalsIgnoreCase(raw) || "on".equalsIgnoreCase(raw)) { + return true; + } + if ("false".equalsIgnoreCase(raw) || "no".equalsIgnoreCase(raw) || "off".equalsIgnoreCase(raw)) { + return false; + } + throw new IllegalArgumentException("Expected true or false for " + flag + ", got: " + raw); + } + + private static int percentValue(String[] args, int index, String flag) { + int percent = intValue(args, index, flag); + if (percent < 10 || percent > 95) { + throw new IllegalArgumentException(flag + " must be between 10 and 95, got: " + percent); + } + return percent; + } + private static int intValue(String[] args, int index, String flag) { String raw = value(args, index, flag); try { @@ -172,6 +209,9 @@ public static String usage() { " --workspace

directory the file tools are confined to (default: cwd)", " --allow-shell add the run_command tool (runs any command line, starting in the workspace)", " --auto run tools without asking (default: ask before writes and commands)", + " --auto-compact summarize the history before it overflows the context (default " + + DEFAULT_AUTO_COMPACT + ")", + " --compact-at how full the context may get first (default " + DEFAULT_COMPACT_AT + ")", " --system replace the default system prompt", " --prompt , -p run one turn and exit (default: interactive; /exit to quit)", " --temperature sampling temperature (default " + DEFAULT_TEMPERATURE + ")", @@ -263,6 +303,24 @@ public Path getWorkspace() { return workspace; } + /** + * Whether the history is summarized before it overflows the context. + * + * @return {@code true} when auto-compaction is on + */ + public boolean isAutoCompact() { + return autoCompact; + } + + /** + * How full the context may get before the history is summarized. + * + * @return the threshold in percent + */ + public int getCompactAt() { + return compactAt; + } + /** * Whether tool calls run without asking. * @@ -341,7 +399,8 @@ public String toString() { return "AgentOptions{baseUrl=" + baseUrl + ", modelPath=" + modelPath + ", gpuLayers=" + gpuLayers + ", ctxSize=" + ctxSize + ", logVerbosity=" + (verbose ? "verbose" : logVerbosity) + ", modelId=" + modelId + ", workspace=" + workspace - + ", allowShell=" + allowShell + ", auto=" + auto + ", temperature=" + temperature + ", maxTokens=" + + ", allowShell=" + allowShell + ", auto=" + auto + ", autoCompact=" + autoCompact + ", temperature=" + + temperature + ", maxTokens=" + maxTokens + ", maxToolRounds=" + maxToolRounds + ", prompt=" + (prompt == null ? "" : "") + "}"; @@ -363,6 +422,8 @@ private static final class Builder { Path workspace = Paths.get("").toAbsolutePath().normalize(); boolean allowShell; boolean auto; + boolean autoCompact = DEFAULT_AUTO_COMPACT; + int compactAt = DEFAULT_COMPACT_AT; double temperature = DEFAULT_TEMPERATURE; int maxTokens = DEFAULT_MAX_TOKENS; int maxToolRounds = DEFAULT_MAX_TOOL_ROUNDS; diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java index fb39e3109..ed14a2022 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java @@ -263,6 +263,14 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream continue; } turnNumber++; + // Compact BEFORE the request that would overflow, not after it: afterwards the + // oversized request has already been sent, which is the one thing to avoid. + if (needsCompaction( + options, contextSize, estimateTokens(systemPrompt(options), history) + line.length() / 4)) { + terminal.line(terminal.ansi().yellow("(context nearly full — compacting first)")); + inputTokens = compact(runner, fileSystem, history, "", terminal); + estimated = true; + } ConsoleSession completed = turn( runner, fileSystem, @@ -604,6 +612,26 @@ static String spinnerWord() { : words.get(java.util.concurrent.ThreadLocalRandom.current().nextInt(words.size())); } + /** + * Whether the history has to be summarized before the next request is sent. + * + *

The threshold is on the low side ({@link AgentOptions#DEFAULT_COMPACT_AT} %) because the + * number it is compared against is usually an estimate, and because the model's reply has to fit + * next to the prompt. With an unknown context size — a foreign endpoint whose {@code /props} says + * nothing — nothing is decided at all rather than guessed. + * + * @param options the options, for the switch and the threshold + * @param contextSize the context window, or {@link StatusLine#UNKNOWN_CONTEXT} + * @param estimatedTokens what the next request is expected to carry + * @return {@code true} when the history should be summarized first + */ + static boolean needsCompaction(AgentOptions options, int contextSize, long estimatedTokens) { + if (!options.isAutoCompact() || contextSize <= StatusLine.UNKNOWN_CONTEXT) { + return false; + } + return estimatedTokens * 100 >= (long) contextSize * options.getCompactAt(); + } + /** * A rough token count of what the next request will carry. * diff --git a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt index 3a281efc9..939121e80 100644 --- a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt +++ b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt @@ -6,6 +6,8 @@ Commands (everything else is sent to the model): /calls every tool call of this session, with its result (/log) /mode [manual|auto] show or set the approval mode (/approve) /compact [focus] summarize the history and continue with the summary + (happens on its own once the context is ~70% full; --auto-compact false + turns that off, --compact-at moves the threshold) /loop [--every 5m] [--max 20] [--check ''] keep working on until the model answers <>; the state lives in AGENT-LOOP.md, not in the conversation diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java index d53f4c08f..1daf6722e 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java @@ -287,4 +287,42 @@ void theToolNoteIsAddressedToTheModelAndNotWrittenAsItsOwnWords() { assertThat(note, containsString("3 matches in 2 files")); assertThat(LocalAgent.toolNote(List.of()), is("")); } + + @Test + void compactionIsDecidedBeforeTheRequestAndOnlyWhenTheWindowIsKnown() { + AgentOptions on = AgentOptions.parse(new String[] {"--base-url", "u"}); + assertThat("on by default", on.isAutoCompact(), is(true)); + assertThat(on.getCompactAt(), is(AgentOptions.DEFAULT_COMPACT_AT)); + + // 70 % of 1000 tokens: 699 still fits, 700 does not + assertThat(LocalAgent.needsCompaction(on, 1000, 699), is(false)); + assertThat(LocalAgent.needsCompaction(on, 1000, 700), is(true)); + + // an unknown window is never guessed at + assertThat(LocalAgent.needsCompaction(on, StatusLine.UNKNOWN_CONTEXT, 1_000_000), is(false)); + + AgentOptions off = AgentOptions.parse(new String[] {"--base-url", "u", "--auto-compact", "false"}); + assertThat(off.isAutoCompact(), is(false)); + assertThat(LocalAgent.needsCompaction(off, 1000, 999), is(false)); + + AgentOptions early = AgentOptions.parse(new String[] {"--base-url", "u", "--compact-at", "50"}); + assertThat(LocalAgent.needsCompaction(early, 1000, 500), is(true)); + assertThat(LocalAgent.needsCompaction(early, 1000, 499), is(false)); + } + + @Test + void aMalformedCompactionFlagIsRejectedWithItsReason() { + assertThat( + org.junit.jupiter.api.Assertions.assertThrows( + IllegalArgumentException.class, + () -> AgentOptions.parse(new String[] {"--base-url", "u", "--auto-compact", "maybe"})) + .getMessage(), + containsString("Expected true or false")); + assertThat( + org.junit.jupiter.api.Assertions.assertThrows( + IllegalArgumentException.class, + () -> AgentOptions.parse(new String[] {"--base-url", "u", "--compact-at", "99"})) + .getMessage(), + containsString("between 10 and 95")); + } } From 0ab8010ea7926a6ec6b4bad7ab08db59cd8bef1f Mon Sep 17 00:00:00 2001 From: Bernard Ladenthin Date: Wed, 23 Sep 2026 16:58:41 +0200 Subject: [PATCH 16/36] llama-atmosphere-agent: refuse to compact an already compacted history Reported: calling /compact repeatedly produced the same summary every time and made llama.cpp log "need to evaluate at least 1 token for each active slot". After a compaction the history is exactly the summary plus its acknowledgement, so compacting again summarizes a summary: same text, another model call, and a byte-identical prompt -- which is what the llama.cpp warning is about (the prompt is fully cached, so there is nothing left to evaluate; the server rolls n_past back by one). The warning is the server's, not a fault here, and it is documented as such. Detection is the summary's own first line rather than "history has two messages": one ordinary turn also leaves two messages, and compacting that is legitimate. Tests: 111 green (was 110). Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01E2h8gXyyE5UeimkL9vQv9G --- llama-atmosphere-agent/README.md | 6 ++++ .../llama/atmosphere/LocalAgent.java | 26 +++++++++++++++-- .../llama/atmosphere/LocalAgentTest.java | 29 +++++++++++++++++++ 3 files changed, 59 insertions(+), 2 deletions(-) diff --git a/llama-atmosphere-agent/README.md b/llama-atmosphere-agent/README.md index 71865b10f..2d68b9670 100644 --- a/llama-atmosphere-agent/README.md +++ b/llama-atmosphere-agent/README.md @@ -243,6 +243,12 @@ unknown context size — a foreign endpoint whose `/props` answers nothing — n all rather than guessed. The threshold sits below the ~85 % a hosted agent uses because our token number is usually an estimate and the reply still has to fit next to the prompt. +Calling `/compact` twice in a row answers `(the history is already a summary — nothing to compact)`: +after a compaction the history *is* the summary plus its acknowledgement, so summarizing it again +returns the same text for another model call. It also re-sends a byte-identical prompt, which is what +makes llama.cpp log `need to evaluate at least 1 token for each active slot` — a harmless note from +the server about a prompt it has already cached in full, not an error on our side. + **`/compact`** asks the model to summarize the conversation (goal, facts, work done, problems, state, next step; `/compact ` adds an emphasis), then replaces the history with that summary. Use it when the context fills up. Note the history only ever held the user texts and the final answers — diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java index ed14a2022..12e5dc8e2 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java @@ -75,6 +75,9 @@ public final class LocalAgent { "You summarize a conversation between a user and a coding assistant. Follow the user's" + " instructions exactly and answer with the summary only."; + /** How the user message of a compacted history begins; also how a repeat compaction is detected. */ + static final String SUMMARY_PREFIX = "Summary of the conversation so far:"; + /** How much of a tool result is kept in the history of later turns. */ private static final int HISTORY_RESULT_CHARS = 400; @@ -632,6 +635,18 @@ static boolean needsCompaction(AgentOptions options, int contextSize, long estim return estimatedTokens * 100 >= (long) contextSize * options.getCompactAt(); } + /** + * Whether the history is the untouched result of a compaction. + * + * @param history the conversation + * @return {@code true} when it is exactly the summary and its acknowledgement + */ + static boolean isCompacted(List history) { + return history.size() == 2 + && history.get(0).content() != null + && history.get(0).content().startsWith(SUMMARY_PREFIX); + } + /** * A rough token count of what the next request will carry. * @@ -680,6 +695,13 @@ private static long compact( terminal.line("(nothing to compact)"); return 0; } + if (isCompacted(history)) { + // After a compaction the history IS the summary plus its acknowledgement. Summarizing that + // again returns the same text for another model call -- and re-sends a byte-identical + // prompt, which is what makes llama.cpp log "need to evaluate at least 1 token". + terminal.line("(the history is already a summary — nothing to compact)"); + return estimateTokens("", history); + } String instructions = prompt(COMPACT_PROMPT) .replace("{focus}", focus.isEmpty() ? "" : System.lineSeparator() + "Focus on: " + focus); int before = history.size(); @@ -702,8 +724,8 @@ private static long compact( return 0; } history.clear(); - history.add(ChatMessage.user("Summary of the conversation so far:" + System.lineSeparator() - + session.text().strip())); + history.add(ChatMessage.user( + SUMMARY_PREFIX + System.lineSeparator() + session.text().strip())); history.add(ChatMessage.assistant("Understood, I will continue from that summary.")); terminal.line("(compacted " + before + " messages into a summary of " + session.text().strip().length() + " characters)"); diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java index 1daf6722e..ee5efdfe6 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java @@ -325,4 +325,33 @@ void aMalformedCompactionFlagIsRejectedWithItsReason() { .getMessage(), containsString("between 10 and 95")); } + + @Test + void compactingAnAlreadyCompactedHistoryIsRefusedInsteadOfRepeated() throws Exception { + // A compacted history is the summary plus its acknowledgement. Summarizing that again returns + // the same text for another model call -- and re-sends a byte-identical prompt, which llama.cpp + // answers with "need to evaluate at least 1 token for each active slot". + ScriptedBackend backend = new ScriptedBackend((call, request) -> ScriptedBackend.textTurn("a summary")); + ByteArrayOutputStream out = new ByteArrayOutputStream(); + try (OpenAiCompatServer server = server(backend)) { + AgentOptions options = AgentOptions.parse(new String[] { + "--base-url", "http://127.0.0.1:" + server.getPort() + "/v1", "--workspace", workspace.toString() + }); + + LocalAgent.run( + options, + new StringReader("hello" + System.lineSeparator() + "/compact" + System.lineSeparator() + "/compact" + + System.lineSeparator()), + new PrintStream(out, true, StandardCharsets.UTF_8), + new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8)); + } + + // one turn + one compaction = two requests; the second /compact must not add a third + assertThat(backend.requests(), hasSize(2)); + assertThat( + "the first compaction really happened", + out.toString(StandardCharsets.UTF_8), + containsString("compacted")); + assertThat(out.toString(StandardCharsets.UTF_8), containsString("already a summary")); + } } From e8cbb6cb1ee48d3878e978d143e0713fdb109e67 Mon Sep 17 00:00:00 2001 From: Bernard Ladenthin Date: Wed, 23 Sep 2026 17:26:33 +0200 Subject: [PATCH 17/36] llama-atmosphere-agent: one printed line is one screen line, a live ctx figure, mode glyphs Three console fixes, all reported from real sessions. A tool call whose argument carried a whole file tore the pinned block apart. The block is reserved in *lines*, so a single "line" holding twenty newlines moves the screen twenty rows further than the terminal accounted for and the block is then drawn across the output. ConsoleSession now folds every value onto one line and cuts each argument *on its own* (80 chars) before cutting the whole rendering (200), so a write_file call still shows the file name rather than half the file; results and errors fold the same way, and JLineTerminal.line splits a multi-line string as a backstop. Only the console is cut -- the model still receives everything, and rounds() keeps the full arguments for the history note and /calls. ConsoleSessionTest drives a recording terminal and asserts the rule directly. The ctx figure only moved at the next you> prompt. The state row was rendered once before the turn and handed to the redraw loop as a fixed string, so it stood still through every tool round -- exactly when knowing the context is filling would be useful. It is a Function now, asked again on every redraw; liveTokens adds what the turn has produced (streamed text plus every call and result, all of which is in the next model call's prompt) and yields to the server's own count as soon as one arrives. The approval mode carries a glyph: the status line and /mode read "pause manual" and "play-play auto", the transport symbols the established terminal agents use for the same distinction. The word stays next to it. 117 tests green. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01E2h8gXyyE5UeimkL9vQv9G --- CLAUDE.md | 32 +++- llama-atmosphere-agent/README.md | 34 ++-- .../llama/atmosphere/ApprovalMode.java | 33 +++- .../llama/atmosphere/ConsoleSession.java | 78 ++++++++- .../llama/atmosphere/JLineTerminal.java | 7 + .../llama/atmosphere/LocalAgent.java | 66 ++++++-- .../llama/atmosphere/StatusLine.java | 2 +- .../ladenthin/llama/atmosphere/TaskLoop.java | 4 +- .../net/ladenthin/llama/atmosphere/help.txt | 3 +- .../atmosphere/ConsoleFormattingTest.java | 12 +- .../llama/atmosphere/ConsoleSessionTest.java | 149 ++++++++++++++++++ 11 files changed, 383 insertions(+), 37 deletions(-) create mode 100644 llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleSessionTest.java diff --git a/CLAUDE.md b/CLAUDE.md index 118a1e96f..d8cbcd334 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -2330,11 +2330,33 @@ are decisions, not details: prompt then reads a key in raw mode on that worker thread while the console thread redraws four times a second, so `TurnActivity` pauses the redraw for as long as the question is open. Reading the pipe incrementally is not only cosmetic — an unread pipe blocks the child once it is full, which on Windows is roughly 4 KB. -8. **The context number in the status line is an estimate, marked `~`.** llama.cpp emits its usage - chunk only when the client sets `stream_options.include_usage`, and Atmosphere's client does not; - `ConsoleSession.usage()` takes the real count when one arrives, otherwise `LocalAgent.estimateTokens` - uses four characters per token. The window size is `--ctx-size` (in-process) or the server's - `/props` (`ServerProps`), and is omitted rather than guessed when neither answers. +8. **The context number in the status line is an estimate, marked `~`, and it moves during the turn.** + llama.cpp emits its usage chunk only when the client sets `stream_options.include_usage`, and + Atmosphere's client does not; `ConsoleSession.usage()` takes the real count when one arrives, + otherwise `LocalAgent.estimateTokens` uses four characters per token. The window size is + `--ctx-size` (in-process) or the server's `/props` (`ServerProps`), and is omitted rather than + guessed when neither answers. **The state row is a `Function`, not a + string**, and `awaitWithActivity` asks it again on every redraw: it used to be rendered once before + the turn and handed over fixed, so the figure stood still through every tool round and only moved + at the next `you>` — which is exactly when it no longer helps anyone decide whether to `/compact`. + `LocalAgent.liveTokens` adds `ConsoleSession.producedChars()` (streamed text **plus** every tool + call and result — all of it is in the prompt of the next model call of the *same* turn) to what the + request carried when it was sent, and yields to the server's own count as soon as one arrives. + `TaskLoop` passes a constant function, and its step label must be copied into a local first: a + lambda may not close over the loop counter. +9. **One call to `AgentTerminal.line` is one screen line**, and `ConsoleSessionTest` is what defends + it. The pinned block is reserved in **lines**, so a single "line" carrying twenty newlines moves the + screen twenty rows further than the terminal accounted for and the block is then drawn across the + output — reported twice, both times from a `write_file` call whose `content` argument was the file. + `ConsoleSession.describeArguments` folds and cuts **each argument value on its own** (80 chars) + before cutting the whole rendering (200), so a call carrying a whole file still shows the file + *name*; results and errors are folded the same way. `JLineTerminal.line` splits a multi-line string + as a backstop for a caller that forgets. Only the console is cut — the model gets everything, and + `ConsoleSession.rounds()` keeps the full arguments for the history note and `/calls`. +10. **The approval mode carries a glyph**: `ApprovalMode.symbol()` / `badge()` render `⏸ manual` and + `⏵⏵ auto` on the status line and in `/mode`, the transport symbols the established terminal agents + use for the same distinction. The word stays next to it; the glyph is what makes the one setting + that decides whether the next command asks first findable at a glance. **The default system prompt is general-purpose on purpose — do not narrow it back.** Every model-facing text is a resource, not a Java literal: `src/main/resources/net/ladenthin/llama/atmosphere/` holds diff --git a/llama-atmosphere-agent/README.md b/llama-atmosphere-agent/README.md index 2d68b9670..5fb1d24ce 100644 --- a/llama-atmosphere-agent/README.md +++ b/llama-atmosphere-agent/README.md @@ -186,7 +186,7 @@ unknown `/command` included — goes to the model: | `/status` | mode, context use, tools, model, workspace, history size | | `/tools` | the tools offered, and which of them ask first | | `/calls` (`/log`) | every tool call of this session with its result — the receipt | -| `/mode [manual\|auto]` (`/approve`) | show or set the approval mode | +| `/mode [manual\|auto]` (`/approve`) | show or set the approval mode (`⏸ manual` / `⏵⏵ auto`) | | `/compact [focus]` | summarize the conversation and continue from the summary | | `/loop [--every 5m] [--max 20] [--check ''] ` | keep working on one task until it is done | | `/clear` (`/reset`, `/new`) | drop the history | @@ -299,7 +299,7 @@ on Windows that buffer is about 4 KB. ``` ──────────────────────────────────────────────────────────────────── ⠙ Fettling… (run_command 47s of 61s · 2 tool calls) -[/path/to/project · manual · ctx ~3.1k/16k · 9 tools · local-model] +[/path/to/project · ⏸ manual · ctx ~3.1k/16k · 9 tools · local-model] ``` The first row is what the agent is doing — `… waiting for input …` when it is your turn, the spinner @@ -313,12 +313,28 @@ stream (piped input, a one-shot run) nothing can be pinned, so the state line is prompt instead and the activity row is dropped rather than repeated into the log. **The status line** above the prompt reads -`[manual · ctx ~3.1k/16k · 9 tools · local-model]`: the approval mode, the context used out of the -window, the number of tools and the model id. A `~` means the number is an estimate from the text -length: llama.cpp reports token counts only to clients that ask for them -(`stream_options.include_usage`), which Atmosphere's client does not. The window size comes from -`--ctx-size` with `--model`, and from the server's `/props` with `--base-url`; when neither answers, -the line shows the count alone. +`[/path/to/project · ⏸ manual · ctx ~3.1k/16k · 9 tools · local-model]`: the workspace, the approval +mode, the context used out of the window, the number of tools and the model id. + +The mode carries a glyph as well as its name — **`⏸ manual`** stops at every gated call, **`⏵⏵ auto`** +runs through — so the one thing that decides whether the next command asks first is findable without +reading the line. + +The context figure **moves while the turn runs**, not only at the next prompt: every tool round appends +the call and its output to the conversation the next model call of the same turn is sent, so a turn that +reads three files and runs a build can add thousands of tokens before you get the prompt back. A `~` +means the number is an estimate from the text length: llama.cpp reports token counts only to clients +that ask for them (`stream_options.include_usage`), which Atmosphere's client does not. The window size +comes from `--ctx-size` with `--model`, and from the server's `/props` with `--base-url`; when neither +answers, the line shows the count alone. + +**One printed line is one screen line.** A tool call and its result are shown as +`● write_file {file_path=notes.md, content=# Notes ## Build … (4812 chars)}` — every argument is +folded onto one line and cut *on its own* before the whole thing is cut, so a call carrying a whole +file still shows the file *name*. The reason is not tidiness: the block at the bottom is reserved in +*lines*, so a single "line" carrying twenty newlines moves the screen twenty rows further than the +terminal accounted for and the block ends up drawn across the output — which is what a `write_file` +call did. The model still receives every argument and every result in full; only the console is cut. **Colours and Markdown.** The answer is rendered line by line as it streams: headings, bullets, fenced code blocks and inline `**bold**` / `` `code` ``. Nothing is ever redrawn, so piping the output @@ -346,7 +362,7 @@ Then, in this order: | `docker is running locally, list the images` | `? run_command {command=docker images}` and the prompt `[y]es / [n]o / [a]uto` | | answer `n` | the command does **not** run; the model is told it was cancelled and offers an alternative | | ask again, answer `y` | the command runs and its output goes back to the model | -| `/mode auto` | the status line flips to `auto`; nothing asks any more | +| `/mode auto` | the status line flips to `⏵⏵ auto`; nothing asks any more | | `explain Markdown with a heading, a list, bold text and a code block` | the answer arrives rendered: heading bold, `•` bullets, code in colour | | press ↑ | the previous line comes back; Tab after `/` completes the commands | | `/compact` | the conversation is summarized and replaces the history; `ctx` drops | diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ApprovalMode.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ApprovalMode.java index a179f588a..dec2efffc 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ApprovalMode.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ApprovalMode.java @@ -17,10 +17,16 @@ public enum ApprovalMode { /** Ask before every gated tool call. */ - MANUAL, + MANUAL("⏸"), /** Run every tool call without asking. */ - AUTO; + AUTO("⏵⏵"); + + private final String symbol; + + ApprovalMode(String symbol) { + this.symbol = symbol; + } /** * The lower-case name used on the console and in {@code /mode}. @@ -31,6 +37,29 @@ public String label() { return name().toLowerCase(Locale.ROOT); } + /** + * The glyph shown in front of the name on the status line. + * + *

Two transport symbols, the way the established terminal agents mark the same distinction: + * {@code ⏸} for a session that stops at every gated call, {@code ⏵⏵} for one that runs through. The + * name stays next to it — the glyph makes the mode findable at a glance, it does not replace the + * word. + * + * @return {@code "⏸"} or {@code "⏵⏵"} + */ + public String symbol() { + return symbol; + } + + /** + * Symbol and name together, as the status line and {@code /mode} print them. + * + * @return e.g. {@code "⏸ manual"} + */ + public String badge() { + return symbol + " " + label(); + } + /** * Parse a mode name as typed by the user. * diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleSession.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleSession.java index 53d2f38c9..63acb8835 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleSession.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleSession.java @@ -29,6 +29,12 @@ public final class ConsoleSession implements StreamingSession { private static final int RESULT_PREVIEW_CHARS = 400; + /** How much of a call's arguments the console shows; the model still gets them in full. */ + private static final int ARGUMENT_PREVIEW_CHARS = 200; + + /** How much of a single argument value survives, so one big one cannot hide the others. */ + private static final int ARGUMENT_VALUE_PREVIEW_CHARS = 80; + private final AgentTerminal terminal; private final Ansi ansi; private final MarkdownConsole markdown; @@ -107,6 +113,25 @@ public void usage(TokenUsage usage) { } } + /** + * Everything this turn has produced so far, in characters: the streamed text plus every tool call + * with its result. + * + *

All of it is in the prompt of the next model call of the same turn — a tool round + * appends the call and its output to the conversation the server is sent. So this is what makes the + * context grow while the turn runs, and the status line adds it to the count it showed before the + * turn started instead of standing still until the next prompt. + * + * @return the character count + */ + public long producedChars() { + long chars = text.length(); + for (ToolRound round : rounds) { + chars += round.argumentsJson().length() + round.result().length(); + } + return chars; + } + /** * The input tokens of the last model call of this turn. * @@ -158,7 +183,7 @@ public void emit(AiEvent event) { rounds.add(new ToolRound(start.toolName(), String.valueOf(start.arguments()), "")); markdown.flush(); terminal.line(ansi.green("●") + " " + ansi.bold(start.toolName()) + " " - + ansi.dim(String.valueOf(start.arguments()))); + + ansi.dim(describeArguments(start.arguments()))); } case AiEvent.ToolResult result -> { runningTool = null; @@ -168,7 +193,7 @@ public void emit(AiEvent event) { case AiEvent.ToolError error -> { runningTool = null; recordResult("error: " + error.error()); - terminal.line(ansi.red(" ↳ error: " + error.error())); + terminal.line(ansi.red(" ↳ error: " + cut(String.valueOf(error.error()), RESULT_PREVIEW_CHARS))); } default -> StreamingSession.super.emit(event); } @@ -204,10 +229,51 @@ private void recordResult(String result) { } private static String preview(String value) { - String oneLine = value.replace("\r\n", "\n").replace('\n', ' '); - return oneLine.length() <= RESULT_PREVIEW_CHARS - ? oneLine - : oneLine.substring(0, RESULT_PREVIEW_CHARS) + " … (" + value.length() + " chars)"; + return cut(value, RESULT_PREVIEW_CHARS); + } + + /** + * The arguments of a call, short enough for one console line. + * + *

Every value is cut on its own before the whole thing is. Cutting only the rendered + * map would let one big argument push the others out of the line — a {@code write_file} call would + * then show half of the file and not the name of the file, which is the one thing worth seeing. + * + * @param arguments what the model passed, usually a map + * @return one line + */ + private static String describeArguments(Object arguments) { + if (!(arguments instanceof Map map)) { + return cut(String.valueOf(arguments), ARGUMENT_PREVIEW_CHARS); + } + StringBuilder rendered = new StringBuilder("{"); + for (Map.Entry entry : map.entrySet()) { + if (rendered.length() > 1) { + rendered.append(", "); + } + rendered.append(entry.getKey()) + .append('=') + .append(cut(String.valueOf(entry.getValue()), ARGUMENT_VALUE_PREVIEW_CHARS)); + } + return cut(rendered.append('}').toString(), ARGUMENT_PREVIEW_CHARS); + } + + /** + * Fold a value onto one line and cut it. + * + *

Both halves matter. The cut keeps a whole file out of the scrollback, and the folding keeps + * the pinned block intact: that block is sized in lines, so a single printed "line" + * carrying twenty newlines moves the screen twenty rows further than the terminal accounted for, + * and the block ends up drawn across the output. A {@code write_file} call whose arguments contain + * the file did exactly that. + * + * @param value the raw text + * @param max how many characters survive + * @return one line, with a note about what was left out + */ + private static String cut(String value, int max) { + String oneLine = value.replace("\r\n", " ").replace('\n', ' ').replace('\r', ' '); + return oneLine.length() <= max ? oneLine : oneLine.substring(0, max) + " … (" + value.length() + " chars)"; } /** diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java index a3f8c8a73..5737d8989 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java @@ -75,6 +75,13 @@ private JLineTerminal(Terminal terminal, LineReader reader, Status status, Ansi @Override public void line(String text) { + if (text.indexOf('\n') >= 0 || text.indexOf('\r') >= 0) { + // One call must be one screen line: the status block is sized in lines, so a multi-line + // string handed over as "a line" desynchronises the reserved region. Callers fold their + // text themselves; this is the backstop for the ones that forget. + text.lines().forEach(this::line); + return; + } if (reading) { // Only while the line reader owns the screen: printAbove scrolls the text in above the // prompt and redraws that prompt afterwards. Calling it when nobody is reading redraws a diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java index 12e5dc8e2..c10882dd7 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java @@ -84,6 +84,9 @@ public final class LocalAgent { /** How often the activity line is refreshed while a turn runs. */ private static final Duration ACTIVITY_INTERVAL = Duration.ofMillis(250); + /** The usual rule of thumb, used wherever a token count has to be guessed from text. */ + private static final int CHARS_PER_TOKEN = 4; + /** The spinner shown in the activity line. */ private static final String ACTIVITY_FRAMES = "⠋⠙⠹⠸⠼⠴⠦⠧⠇⠏"; @@ -207,7 +210,16 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream : ServerProps.contextSize(baseUrl, options.getApiKey()); if (options.getPrompt() != null) { - return turn(runner, fileSystem, options.getPrompt(), history, terminal, callLog, 1, "", activity) + return turn( + runner, + fileSystem, + options.getPrompt(), + history, + terminal, + callLog, + 1, + ignored -> "", + activity) .failure() == null ? 0 @@ -274,6 +286,8 @@ options, contextSize, estimateTokens(systemPrompt(options), history) + line.leng inputTokens = compact(runner, fileSystem, history, "", terminal); estimated = true; } + // What the request carries before the model has answered anything; the turn adds to it. + long baseTokens = estimateTokens(systemPrompt(options), history) + line.length() / CHARS_PER_TOKEN; ConsoleSession completed = turn( runner, fileSystem, @@ -284,7 +298,14 @@ options, contextSize, estimateTokens(systemPrompt(options), history) + line.leng terminal, callLog, turnNumber, - status, + running -> StatusLine.render( + options.getWorkspace(), + mode.get(), + liveTokens(baseTokens, running), + running.inputTokens() == 0, + contextSize, + tools.size(), + options.getModelId()), activity); pendingNote = toolNote(completed.rounds()); estimated = completed.inputTokens() == 0; @@ -373,7 +394,7 @@ static ConsoleSession turn( AgentTerminal terminal, ToolCallLog callLog, int turnNumber, - String stateLine, + java.util.function.Function stateLine, TurnActivity activity) throws InterruptedException { ConsoleSession session = new ConsoleSession(terminal, fileSystem); @@ -392,7 +413,7 @@ static ConsoleSession turn( "agent-turn"); worker.setDaemon(true); worker.start(); - boolean finished = awaitWithActivity(session, terminal, stateLine, activity); + boolean finished = awaitWithActivity(session, terminal, () -> stateLine.apply(session), activity); worker.join(Duration.ofSeconds(5).toMillis()); history.add(ChatMessage.user(message)); callLog.add(turnNumber, session.rounds()); @@ -506,7 +527,7 @@ private static long handleCommand( return inputTokens; } } - terminal.line("approval mode: " + mode.get().label()); + terminal.line("approval mode: " + mode.get().badge()); } case STATUS -> { terminal.line(StatusLine.render( @@ -540,13 +561,17 @@ private static long handleCommand( * * @param session the running turn * @param terminal the console - * @param stateLine the second line of the block, kept in place so it does not change height + * @param stateLine the second row of the block, asked again on every redraw so the context figure + * moves while the turn runs rather than standing still until the next prompt * @param activity paused while an approval question is open * @return {@code true} when the turn finished within {@link #TURN_TIMEOUT} * @throws InterruptedException if interrupted while waiting */ private static boolean awaitWithActivity( - ConsoleSession session, AgentTerminal terminal, String stateLine, TurnActivity activity) + ConsoleSession session, + AgentTerminal terminal, + java.util.function.Supplier stateLine, + TurnActivity activity) throws InterruptedException { long start = System.nanoTime(); int frame = 0; @@ -568,9 +593,9 @@ private static boolean awaitWithActivity( session.runningTool(), session.runningSeconds(), session.toolCalls()), - stateLine)); + stateLine.get())); } - terminal.status(List.of(IDLE_LINE, stateLine)); + terminal.status(List.of(IDLE_LINE, stateLine.get())); return true; } @@ -664,7 +689,26 @@ static long estimateTokens(String systemPrompt, List history) { for (ChatMessage message : history) { characters += message.content() == null ? 0 : message.content().length(); } - return characters / 4; + return characters / CHARS_PER_TOKEN; + } + + /** + * The context figure while a turn is running. + * + *

The count shown before the turn started is not the count during it: every tool round appends + * the call and its output to the conversation the next model call of the same turn is sent, so a + * turn that reads three files and runs a build can add thousands of tokens before the prompt comes + * back. The status row used to be rendered once and handed to the redraw loop as a fixed string, so + * it stood still for the whole turn and only moved at the next {@code you>} — which is precisely + * when it no longer matters. + * + * @param baseTokens what the request carried when it was sent + * @param session the running turn + * @return the server's own count once it reported one, else the base plus what the turn produced + */ + static long liveTokens(long baseTokens, ConsoleSession session) { + long reported = session.inputTokens(); + return reported > 0 ? reported : baseTokens + session.producedChars() / CHARS_PER_TOKEN; } /** @@ -718,7 +762,7 @@ private static long compact( "agent-compact"); worker.setDaemon(true); worker.start(); - if (!awaitWithActivity(session, terminal, "/compact", new TurnActivity()) + if (!awaitWithActivity(session, terminal, () -> "/compact", new TurnActivity()) || session.text().isBlank()) { terminal.line("(compact failed; history kept)"); return 0; diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/StatusLine.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/StatusLine.java index 1dff4fd9b..f4ba96a4f 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/StatusLine.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/StatusLine.java @@ -54,7 +54,7 @@ public static String render( int contextSize, int tools, String modelId) { - return "[" + shorten(workspace) + " · " + mode.label() + " · " + context(inputTokens, estimated, contextSize) + return "[" + shorten(workspace) + " · " + mode.badge() + " · " + context(inputTokens, estimated, contextSize) + " · " + tools + " tools · " + modelId + "]"; } diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/TaskLoop.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/TaskLoop.java index a9f9c6866..099a32a02 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/TaskLoop.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/TaskLoop.java @@ -179,6 +179,8 @@ public static Outcome run( "budget of " + budget.toMinutes() + " min used up after " + (step - 1) + " steps", false); } terminal.line(terminal.ansi().dim("── loop step " + step + "/" + options.maxSteps() + " ──")); + // the status row is rendered on every redraw, and a lambda may not close over the counter + String stepLabel = "loop step " + step + "/" + options.maxSteps() + " · " + options.task(); // A fresh history every step: the file is the memory, so the context cannot grow. ConsoleSession session = LocalAgent.turn( @@ -189,7 +191,7 @@ public static Outcome run( terminal, callLog, step, - "loop step " + step + "/" + options.maxSteps() + " · " + options.task(), + ignored -> stepLabel, new TurnActivity()); extra = ""; if (session.failure() != null) { diff --git a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt index 939121e80..e8b94d4d7 100644 --- a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt +++ b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt @@ -4,7 +4,8 @@ Commands (everything else is sent to the model): /status endpoint, model, tools, mode, context use /tools the tools offered to the model /calls every tool call of this session, with its result (/log) - /mode [manual|auto] show or set the approval mode (/approve) + /mode [manual|auto] show or set the approval mode: manual asks first (shown as + the mode symbol on the status line), auto runs through (/approve) /compact [focus] summarize the history and continue with the summary (happens on its own once the context is ~70% full; --auto-compact false turns that off, --compact-at moves the threshold) diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleFormattingTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleFormattingTest.java index 752d73d81..4735d81fd 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleFormattingTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleFormattingTest.java @@ -53,7 +53,17 @@ void theStatusLineShowsWorkspaceModeContextToolsAndModel() { String line = StatusLine.render( java.nio.file.Path.of("/tmp/ws"), ApprovalMode.MANUAL, 1234, false, 16384, 9, "local-model"); - assertThat(line, containsString("ws · manual · ctx 1.2k/16k · 9 tools · local-model]")); + assertThat(line, containsString("ws · ⏸ manual · ctx 1.2k/16k · 9 tools · local-model]")); + } + + @Test + void eachModeCarriesItsOwnSymbol() { + // the glyph is what makes the mode findable at a glance; the word stays next to it + assertThat(ApprovalMode.MANUAL.badge(), is("⏸ manual")); + assertThat(ApprovalMode.AUTO.badge(), is("⏵⏵ auto")); + assertThat( + StatusLine.render(java.nio.file.Path.of("/tmp/ws"), ApprovalMode.AUTO, 0, true, 0, 1, "m"), + containsString("⏵⏵ auto")); } @Test diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleSessionTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleSessionTest.java new file mode 100644 index 000000000..a58f6f99b --- /dev/null +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleSessionTest.java @@ -0,0 +1,149 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import static org.hamcrest.MatcherAssert.assertThat; +import static org.hamcrest.Matchers.containsString; +import static org.hamcrest.Matchers.is; +import static org.hamcrest.Matchers.not; + +import java.nio.file.Path; +import java.util.ArrayList; +import java.util.List; +import java.util.Map; +import org.atmosphere.ai.AiEvent; +import org.atmosphere.ai.fs.AgentFileSystem; +import org.atmosphere.ai.fs.WorkspaceAgentFileSystem; +import org.junit.jupiter.api.Test; +import org.junit.jupiter.api.io.TempDir; + +/** + * What the console is handed for one turn, with a terminal that only records. + * + *

The rule every test here defends: one call to the terminal is one screen line. The pinned + * block at the bottom is sized in lines, so a single "line" carrying ten newlines pushes the screen + * ten rows further than the terminal accounted for and the block is drawn across the output — which is + * what a {@code write_file} call with a whole file in its arguments did. + */ +class ConsoleSessionTest { + + @TempDir + Path workspace; + + /** A terminal that records what it was told to print. */ + private static final class RecordingTerminal implements AgentTerminal { + private final List lines = new ArrayList<>(); + + @Override + public void line(String text) { + lines.add(text); + } + + @Override + public String readLine(String prompt) { + return null; + } + + @Override + public String readKey(String prompt) { + return null; + } + + @Override + public void status(List statusLines) {} + + @Override + public boolean pinsStatus() { + return false; + } + + @Override + public Ansi ansi() { + return Ansi.PLAIN; + } + + @Override + public void close() {} + } + + private final RecordingTerminal terminal = new RecordingTerminal(); + + private ConsoleSession session() { + AgentFileSystem fs = new WorkspaceAgentFileSystem(workspace, AgentFileSystem.Limits.defaults()); + return new ConsoleSession(terminal, fs); + } + + private void assertEveryLineIsOneLine() { + for (String line : terminal.lines) { + assertThat("a printed line must not contain a newline: " + line, line.contains("\n"), is(false)); + assertThat(line.contains("\r"), is(false)); + } + } + + @Test + void aToolCallWithAWholeFileInItsArgumentsStaysOnOneLine() { + String fileContent = "# Project Summary\n\n## Build\n\nmvn package\n".repeat(20); + ConsoleSession session = session(); + + session.emit(new AiEvent.ToolStart("write_file", Map.of("path", "project_summary.md", "content", fileContent))); + + assertEveryLineIsOneLine(); + assertThat(terminal.lines, is(not(List.of()))); + assertThat(terminal.lines.get(0), containsString("write_file")); + assertThat(terminal.lines.get(0), containsString("project_summary.md")); + assertThat("and it is cut, not merely joined", terminal.lines.get(0).length() < 400, is(true)); + } + + @Test + void aMultiLineToolResultStaysOnOneLineToo() { + ConsoleSession session = session(); + + session.emit(new AiEvent.ToolStart("run_command", Map.of("command", "mvn -version"))); + session.emit(new AiEvent.ToolResult("run_command", "exit code: 0\nline one\r\nline two\n")); + session.emit(new AiEvent.ToolError("run_command", "boom\nand more")); + + assertEveryLineIsOneLine(); + } + + @Test + void theContextFigureGrowsWhileTheTurnRuns() { + // it used to be rendered once before the turn and handed over as a fixed string, so it stood + // still through every tool round and only moved at the next prompt + ConsoleSession session = session(); + assertThat(LocalAgent.liveTokens(1000, session), is(1000L)); + + session.send("a".repeat(400)); + session.emit(new AiEvent.ToolStart("read_file", Map.of("file_path", "x"))); + session.emit(new AiEvent.ToolResult("read_file", "b".repeat(4000))); + + assertThat( + "the tool output counts too, it is in the next call's prompt", + LocalAgent.liveTokens(1000, session) > 2000L, + is(true)); + } + + @Test + void aReportedCountWinsOverTheEstimate() { + ConsoleSession session = session(); + session.send("a".repeat(4000)); + session.usage(new org.atmosphere.ai.TokenUsage(7777, 10, 0, 7787, "m")); + + assertThat(LocalAgent.liveTokens(1000, session), is(7777L)); + } + + @Test + void theRecordedRoundKeepsTheFullArgumentsEvenThoughTheConsoleShowsLess() { + String fileContent = "a".repeat(5000); + ConsoleSession session = session(); + + session.emit(new AiEvent.ToolStart("write_file", Map.of("content", fileContent))); + + assertThat("the console is cut …", terminal.lines.get(0).length() < 400, is(true)); + assertThat( + "… but what the model is told later is not cut here", + session.rounds().get(0).argumentsJson().length() > 4000, + is(true)); + } +} From 35a072e24510e70b4c00acef5fb3eb086f5d6c3f Mon Sep 17 00:00:00 2001 From: Bernard Ladenthin Date: Wed, 23 Sep 2026 18:08:13 +0200 Subject: [PATCH 18/36] llama-atmosphere-agent: shift+tab switches the approval mode The glyph made the mode visible; this makes it reachable without typing a command, the way the established terminal agents bind it. ApprovalMode.next() is the cycle. The binding is AgentTerminal.onCycleMode, which defaults to declining -- only JLineTerminal overrides it, and the startup line advertises the key only when the bind succeeded, so a plain stream and a terminal that cannot send backtab are simply quiet about it and keep /mode. Two details that are not obvious. The key is bound both through terminfo (key_btab) and to the literal ESC [ Z, because JLine's windows-vtp.caps declares no key_btab at all while the terminal, in virtual-terminal input mode, does send the sequence. And a widget only runs while a line is being read, so the switch happens between turns -- which is when the mode is decided anyway. The REPL therefore holds its two status numbers in an AtomicLong/AtomicBoolean rather than locals, so the widget can re-render the pinned row from inside the reader with what the last turn left behind. 118 tests green. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01E2h8gXyyE5UeimkL9vQv9G --- CLAUDE.md | 16 +++++-- llama-atmosphere-agent/README.md | 6 ++- .../llama/atmosphere/AgentTerminal.java | 17 +++++++ .../llama/atmosphere/ApprovalMode.java | 14 ++++++ .../llama/atmosphere/JLineTerminal.java | 36 ++++++++++++++ .../llama/atmosphere/LocalAgent.java | 47 ++++++++++++++----- .../net/ladenthin/llama/atmosphere/help.txt | 2 + .../atmosphere/ConsoleFormattingTest.java | 19 ++++++++ 8 files changed, 139 insertions(+), 18 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index d8cbcd334..37f0f12e3 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -2353,10 +2353,18 @@ are decisions, not details: *name*; results and errors are folded the same way. `JLineTerminal.line` splits a multi-line string as a backstop for a caller that forgets. Only the console is cut — the model gets everything, and `ConsoleSession.rounds()` keeps the full arguments for the history note and `/calls`. -10. **The approval mode carries a glyph**: `ApprovalMode.symbol()` / `badge()` render `⏸ manual` and - `⏵⏵ auto` on the status line and in `/mode`, the transport symbols the established terminal agents - use for the same distinction. The word stays next to it; the glyph is what makes the one setting - that decides whether the next command asks first findable at a glance. +10. **The approval mode carries a glyph, and shift+tab switches it**: `ApprovalMode.symbol()` / + `badge()` render `⏸ manual` and `⏵⏵ auto` on the status line and in `/mode`, the transport symbols + the established terminal agents use for the same distinction; `ApprovalMode.next()` is the cycle + the key walks. The binding is `AgentTerminal.onCycleMode(Runnable)`, which **defaults to declining** + — only `JLineTerminal` overrides it, and the startup line advertises the key only when the bind + succeeded. Two details are not obvious: it is bound **both** through terminfo + (`InfoCmp.Capability.key_btab`) **and** to the literal `ESC [ Z`, because JLine's + `windows-vtp.caps` declares no `key_btab` at all while the terminal in virtual-terminal input mode + does send the sequence; and the shortcut fires **only while a line is being read**, so it switches + the mode between turns — which is when it is decided anyway. The REPL therefore holds the two + status numbers in an `AtomicLong`/`AtomicBoolean` rather than locals, so the widget (which runs + inside the reader) can re-render the pinned row with what the last turn left behind. **The default system prompt is general-purpose on purpose — do not narrow it back.** Every model-facing text is a resource, not a Java literal: `src/main/resources/net/ladenthin/llama/atmosphere/` holds diff --git a/llama-atmosphere-agent/README.md b/llama-atmosphere-agent/README.md index 5fb1d24ce..31a709564 100644 --- a/llama-atmosphere-agent/README.md +++ b/llama-atmosphere-agent/README.md @@ -318,7 +318,10 @@ mode, the context used out of the window, the number of tools and the model id. The mode carries a glyph as well as its name — **`⏸ manual`** stops at every gated call, **`⏵⏵ auto`** runs through — so the one thing that decides whether the next command asks first is findable without -reading the line. +reading the line. **Shift+Tab at the prompt switches it**, without typing `/mode`; the status line +updates on the key. The shortcut needs a real terminal and works between turns (while the prompt is +waiting), which is when the mode matters — during a turn nobody is reading keys. Where the terminal +does not send backtab, the startup line simply does not offer it and `/mode` still works. The context figure **moves while the turn runs**, not only at the next prompt: every tool round appends the call and its output to the conversation the next model call of the same turn is sent, so a turn that @@ -363,6 +366,7 @@ Then, in this order: | answer `n` | the command does **not** run; the model is told it was cancelled and offers an alternative | | ask again, answer `y` | the command runs and its output goes back to the model | | `/mode auto` | the status line flips to `⏵⏵ auto`; nothing asks any more | +| press Shift+Tab at the prompt | the same switch without a command; the status line updates immediately | | `explain Markdown with a heading, a list, bold text and a code block` | the answer arrives rendered: heading bold, `•` bullets, code in colour | | press ↑ | the previous line comes back; Tab after `/` completes the commands | | `/compact` | the conversation is summarized and replaces the history; `ctx` drops | diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentTerminal.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentTerminal.java index 8da5e07e8..c9f4efe57 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentTerminal.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentTerminal.java @@ -67,6 +67,23 @@ public interface AgentTerminal extends AutoCloseable { */ boolean pinsStatus(); + /** + * Ask the terminal to run {@code action} when the user presses shift+tab at the prompt. + * + *

Only a real terminal can offer this: it needs to own the keyboard and see a key that is not + * a line of text. It also only fires **while a line is being read** — during a turn nobody is + * reading keys, so the mode is switched between turns, which is when it matters. + * + *

The default is to decline, which every non-interactive console does; the caller uses that + * answer to decide whether to advertise the shortcut, and {@code /mode} remains either way. + * + * @param action what to run on the key; it must not block + * @return {@code true} when the key was bound + */ + default boolean onCycleMode(Runnable action) { + return false; + } + /** * The styles to use for this console. * diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ApprovalMode.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ApprovalMode.java index dec2efffc..9ef88ca99 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ApprovalMode.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ApprovalMode.java @@ -60,6 +60,20 @@ public String badge() { return symbol + " " + label(); } + /** + * The next mode in the cycle, which is what the shift+tab shortcut switches to. + * + *

With two modes this is a toggle; it is written as a cycle so a third mode would need no + * change here. The order follows the declaration order, so it is the same order {@code /mode} + * lists. + * + * @return the following mode, wrapping around at the end + */ + public ApprovalMode next() { + ApprovalMode[] all = values(); + return all[(ordinal() + 1) % all.length]; + } + /** * Parse a mode name as typed by the user. * diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java index 5737d8989..32e77fd55 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java @@ -7,9 +7,12 @@ import java.io.IOException; import java.util.List; import java.util.Locale; +import org.jline.keymap.KeyMap; +import org.jline.reader.Binding; import org.jline.reader.EndOfFileException; import org.jline.reader.LineReader; import org.jline.reader.LineReaderBuilder; +import org.jline.reader.Reference; import org.jline.reader.UserInterruptException; import org.jline.reader.impl.completer.StringsCompleter; import org.jline.terminal.Attributes; @@ -17,6 +20,7 @@ import org.jline.terminal.TerminalBuilder; import org.jline.utils.AttributedString; import org.jline.utils.AttributedStyle; +import org.jline.utils.InfoCmp; import org.jline.utils.Status; import org.jspecify.annotations.Nullable; @@ -165,6 +169,38 @@ static String fit(String text, int width) { return text.length() <= width ? text : text.substring(0, Math.max(1, width - 1)) + "…"; } + /** + * What a terminal sends for shift+tab: {@code ESC [ Z}, "backtab" (CSI Z). + * + *

It is bound literally as well as through terminfo, because JLine's Windows terminfo + * (windows-vtp.caps) declares no key_btab at all — so the capability + * lookup yields nothing there while the terminal itself, in virtual-terminal input mode, does + * send the sequence. + */ + private static final String BACKTAB = "\033[Z"; + + /** The name the cycle action is registered under; a widget is addressed by name, not by object. */ + private static final String CYCLE_MODE_WIDGET = "jllama-cycle-approval-mode"; + + @Override + public boolean onCycleMode(Runnable action) { + KeyMap keys = reader.getKeyMaps().get(LineReader.MAIN); + if (keys == null) { + return false; + } + reader.getWidgets().put(CYCLE_MODE_WIDGET, () -> { + action.run(); + return true; + }); + Reference widget = new Reference(CYCLE_MODE_WIDGET); + String fromTerminfo = KeyMap.key(terminal, InfoCmp.Capability.key_btab); + if (fromTerminfo != null) { + keys.bind(widget, fromTerminfo); + } + keys.bind(widget, BACKTAB); + return true; + } + @Override public boolean pinsStatus() { return true; diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java index c10882dd7..852bab229 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java @@ -229,9 +229,29 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream err.println("No interactive input available; pass --prompt ."); return 2; } - err.println("Interactive mode: type a request, /help for the commands."); - long inputTokens = 0; - boolean estimated = false; + // Held rather than kept in locals so the shift+tab shortcut, which fires from inside the + // line reader, can re-render the status row with the numbers the last turn left behind. + java.util.concurrent.atomic.AtomicLong inputTokens = new java.util.concurrent.atomic.AtomicLong(); + java.util.concurrent.atomic.AtomicBoolean estimated = new java.util.concurrent.atomic.AtomicBoolean(); + AgentTerminal console = terminal; + java.util.function.Supplier idleStatus = () -> StatusLine.render( + options.getWorkspace(), + mode.get(), + inputTokens.get(), + estimated.get(), + contextSize, + tools.size(), + options.getModelId()); + boolean shortcut = console.onCycleMode(() -> { + mode.set(mode.get().next()); + if (console.pinsStatus()) { + console.status(List.of(IDLE_LINE, idleStatus.get())); + } else { + console.line(console.ansi().dim(idleStatus.get())); + } + }); + err.println("Interactive mode: type a request, /help for the commands." + + (shortcut ? " shift+tab switches the approval mode." : "")); int turnNumber = 0; String pendingNote = ""; while (true) { @@ -241,8 +261,8 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream String status = StatusLine.render( options.getWorkspace(), mode.get(), - inputTokens, - estimated, + inputTokens.get(), + estimated.get(), contextSize, tools.size(), options.getModelId()); @@ -263,7 +283,7 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream if (command.get().command() == SlashCommands.Command.EXIT) { return 0; } - inputTokens = handleCommand( + inputTokens.set(handleCommand( command.get(), runner, fileSystem, @@ -271,10 +291,10 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream mode, options, contextSize, - inputTokens, - estimated, + inputTokens.get(), + estimated.get(), terminal, - callLog); + callLog)); continue; } turnNumber++; @@ -283,8 +303,8 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream if (needsCompaction( options, contextSize, estimateTokens(systemPrompt(options), history) + line.length() / 4)) { terminal.line(terminal.ansi().yellow("(context nearly full — compacting first)")); - inputTokens = compact(runner, fileSystem, history, "", terminal); - estimated = true; + inputTokens.set(compact(runner, fileSystem, history, "", terminal)); + estimated.set(true); } // What the request carries before the model has answered anything; the turn adds to it. long baseTokens = estimateTokens(systemPrompt(options), history) + line.length() / CHARS_PER_TOKEN; @@ -308,8 +328,9 @@ options, contextSize, estimateTokens(systemPrompt(options), history) + line.leng options.getModelId()), activity); pendingNote = toolNote(completed.rounds()); - estimated = completed.inputTokens() == 0; - inputTokens = estimated ? estimateTokens(systemPrompt(options), history) : completed.inputTokens(); + estimated.set(completed.inputTokens() == 0); + inputTokens.set( + estimated.get() ? estimateTokens(systemPrompt(options), history) : completed.inputTokens()); } } finally { if (terminal != null) { diff --git a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt index e8b94d4d7..434009336 100644 --- a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt +++ b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt @@ -15,6 +15,8 @@ Commands (everything else is sent to the model): /clear drop the history (/reset, /new) /exit leave (/quit) +Shift+Tab at the prompt switches the approval mode without typing a command. + In manual mode every tool that writes or runs a command asks first: [y]es runs it once, [n]o tells the model the user cancelled it, [a]uto stops asking for the rest of the session. Reading tools never ask. An unknown /command is sent to the model, not rejected. diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleFormattingTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleFormattingTest.java index 4735d81fd..b06d06a5e 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleFormattingTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleFormattingTest.java @@ -56,6 +56,25 @@ void theStatusLineShowsWorkspaceModeContextToolsAndModel() { assertThat(line, containsString("ws · ⏸ manual · ctx 1.2k/16k · 9 tools · local-model]")); } + @Test + void shiftTabCyclesTheModeAndAPlainStreamDeclinesTheShortcut() { + // the key can only be seen by a console that owns the keyboard; everything else keeps /mode + assertThat(ApprovalMode.MANUAL.next(), is(ApprovalMode.AUTO)); + assertThat(ApprovalMode.AUTO.next(), is(ApprovalMode.MANUAL)); + assertThat( + "a cycle, so it always returns to where it started", + ApprovalMode.MANUAL.next().next(), + is(ApprovalMode.MANUAL)); + + AgentTerminal plain = + new PlainTerminal(new java.io.PrintStream(new java.io.ByteArrayOutputStream()), null, Ansi.PLAIN); + assertThat( + plain.onCycleMode(() -> { + throw new AssertionError("must not run"); + }), + is(false)); + } + @Test void eachModeCarriesItsOwnSymbol() { // the glyph is what makes the mode findable at a glance; the word stays next to it From 7366204b4682eb902bdd52bf47bcf7cefac52ad6 Mon Sep 17 00:00:00 2001 From: Bernard Ladenthin Date: Wed, 23 Sep 2026 18:16:15 +0200 Subject: [PATCH 19/36] llama-atmosphere-agent: the prompt stays at the bottom, and typing stops the turn You could not type while the agent was working: the line reader only ran between turns. Now one thread inside JLineTerminal owns the keyboard for the whole session and fills a queue, and every read in that class is served from it -- a terminal has one keyboard, and two threads reading it take turns at random. Everything else is written above the prompt through printAbove, which is what that method exists for. A line typed during a turn stops the turn. AgentRunner.start() uses Atmosphere's cancellable entry point, whose handle closes the in-flight SSE stream ("D-6 built-in hard-cancel"), so the stop is immediate rather than at the end of a tool loop that may run for minutes; the line stays queued and becomes the next message, with whatever the model had already produced kept in the history. executeWithHandle dispatches on a virtual thread and returns at once, so turn() no longer needs its own worker thread either. Two consequences worth stating. The approval question is answered in the same input line now, so it is y + Enter rather than a bare y: a single-key read needs a second reader on the same keyboard. And awaitWithActivity returns a three-valued TurnEnd, because an interrupted turn must not be reported as the timeout error the old boolean produced. This is not Claude Code's steering: AgentExecutionContext is a record whose request is built once from the message plus the history, with nothing to append to, so stop-and-resend is the achievable equivalent. 120 tests green. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01E2h8gXyyE5UeimkL9vQv9G --- CLAUDE.md | 21 +++ llama-atmosphere-agent/README.md | 15 ++ .../llama/atmosphere/AgentRunner.java | 32 +++- .../llama/atmosphere/AgentTerminal.java | 26 +++- .../llama/atmosphere/JLineTerminal.java | 140 +++++++++++++----- .../llama/atmosphere/LocalAgent.java | 74 +++++---- .../net/ladenthin/llama/atmosphere/help.txt | 3 +- .../llama/atmosphere/ConsoleSessionTest.java | 61 ++++++++ 8 files changed, 303 insertions(+), 69 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 37f0f12e3..230841be6 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -2366,6 +2366,27 @@ are decisions, not details: status numbers in an `AtomicLong`/`AtomicBoolean` rather than locals, so the widget (which runs inside the reader) can re-render the pinned row with what the last turn left behind. +11. **The prompt stays at the bottom during a turn, and typing stops the turn.** One thread inside + `JLineTerminal` (`startReading`) sits in `readLine` for the whole session and fills a queue; + **every** read in that class is served from it, because a terminal has one keyboard and two threads + reading it take turns at random. That is why `readKey` no longer reads a single key in raw mode: the + question is printed above the prompt and answered in the same input line (`y` + Enter). End of input + cannot be a queue value, so a sentinel is queued and **put back on every take** — otherwise the + second reader after Ctrl-D would see "nothing typed yet" instead of "no more input". + `AgentTerminal.hasPendingInput()` is the peek the turn loop peeks with; it deliberately does **not** + consume, so the line the interruption was triggered by is still there for the next `readLine` and + becomes the next message. The stop itself is `AgentRunner.start(...)` → + `runtime.executeWithHandle(...)`, whose handle closes the in-flight SSE stream (Atmosphere's own + "D-6 built-in hard-cancel"); that also replaced `turn()`'s hand-rolled worker thread, since + `executeWithHandle` dispatches on a virtual thread and returns at once. `awaitWithActivity` returns a + three-valued `TurnEnd` rather than a boolean, because *interrupted* must not be reported as the + *timed out* error the old `false` produced. **Order in that loop is load-bearing and a test pins the + behaviour**: the `activity.isPaused()` check comes first, so while an approval question is open a + typed line is its answer and not an interruption. **What this is not:** Claude Code injects a + mid-turn message into the running loop; `AgentExecutionContext` is a record whose request is built + once from `message()` + `history()`, with nothing to append to, so stop-and-resend is the achievable + equivalent — and it acts immediately instead of waiting out a tool loop. + **The default system prompt is general-purpose on purpose — do not narrow it back.** Every model-facing text is a resource, not a Java literal: `src/main/resources/net/ladenthin/llama/atmosphere/` holds `system-prompt.txt`, `system-prompt-shell.txt`, `system-prompt-no-shell.txt` and `run-command-tool.txt` diff --git a/llama-atmosphere-agent/README.md b/llama-atmosphere-agent/README.md index 31a709564..b7d71be17 100644 --- a/llama-atmosphere-agent/README.md +++ b/llama-atmosphere-agent/README.md @@ -331,6 +331,20 @@ that ask for them (`stream_options.include_usage`), which Atmosphere's client do comes from `--ctx-size` with `--model`, and from the server's `/props` with `--base-url`; when neither answers, the line shows the count alone. +**The prompt is at the bottom the whole time, and you can type while the agent works.** One thread owns +the keyboard and sits in the line reader for the entire session; everything else is written *above* the +prompt. A line typed **during** a turn **stops that turn** — Atmosphere's cancellable entry point closes +the HTTP stream the model is answering on — and is then sent as the next message, with whatever the model +had already produced kept in the history. That is not quite what Claude Code does (it feeds the message +into the running loop); Atmosphere builds its request once from the message plus the history and has no +place to append to, so stop-and-resend is the honest equivalent, and it is immediate rather than waiting +out a tool loop that may run for minutes. + +The cost is one keystroke: the approval question is answered in that same input line, so it is +`y` + Enter rather than a bare `y`. A single-key read needs a second reader on the same keyboard, and two +readers on one terminal take turns at random. While a question is open the typing-interrupts rule is +suspended, so an answer is an answer and not an interruption. + **One printed line is one screen line.** A tool call and its result are shown as `● write_file {file_path=notes.md, content=# Notes ## Build … (4812 chars)}` — every argument is folded onto one line and cut *on its own* before the whole thing is cut, so a call carrying a whole @@ -367,6 +381,7 @@ Then, in this order: | ask again, answer `y` | the command runs and its output goes back to the model | | `/mode auto` | the status line flips to `⏵⏵ auto`; nothing asks any more | | press Shift+Tab at the prompt | the same switch without a command; the status line updates immediately | +| type a sentence while it is still working, press Enter | the turn stops at once and your message is the next one | | `explain Markdown with a heading, a list, bold text and a code block` | the answer arrives rendered: heading bold, `•` bullets, code in colour | | press ↑ | the previous line comes back; Tab after `/` completes the commands | | `/compact` | the conversation is summarized and replaces the history; `ctx` drops | diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentRunner.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentRunner.java index de969f86f..574b420f0 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentRunner.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentRunner.java @@ -8,6 +8,7 @@ import java.util.Map; import org.atmosphere.ai.AgentExecutionContext; import org.atmosphere.ai.AiConfig; +import org.atmosphere.ai.ExecutionHandle; import org.atmosphere.ai.RetryPolicy; import org.atmosphere.ai.StreamingSession; import org.atmosphere.ai.approval.ApprovalStrategy; @@ -133,7 +134,25 @@ public List toolNames() { * @param session receives streamed text, tool events and the terminal complete/error */ public void run(String message, List history, StreamingSession session) { - run(message, history, session, tools, systemPrompt); + runtime.execute(context(message, history, tools, systemPrompt, session), session); + } + + /** + * Start one user turn and return at once, with a handle that can stop it. + * + *

This is the same turn {@link #run} performs, on Atmosphere's cancellation-aware entry point: + * the turn runs on a virtual thread of the framework's, and {@link ExecutionHandle#cancel()} + * closes the HTTP stream the model is answering on, which unblocks the read loop. That is what + * lets a request typed while the agent is working take effect immediately instead of at the end + * of a tool loop that may run for minutes. + * + * @param message the user message + * @param history prior turns, replayed before the message + * @param session receives streamed text, tool events and the terminal complete/error + * @return the handle; the session's own completion stays the signal that the turn is over + */ + public ExecutionHandle start(String message, List history, StreamingSession session) { + return runtime.executeWithHandle(context(message, history, tools, systemPrompt, session), session); } /** @@ -147,15 +166,15 @@ public void run(String message, List history, StreamingSession sess */ public void runWithoutTools( String message, List history, StreamingSession session, String systemPrompt) { - run(message, history, session, List.of(), systemPrompt); + runtime.execute(context(message, history, List.of(), systemPrompt, session), session); } - private void run( + private AgentExecutionContext context( String message, List history, - StreamingSession session, List tools, - String systemPrompt) { + String systemPrompt, + StreamingSession session) { AgentExecutionContext context = new AgentExecutionContext( message, systemPrompt, @@ -176,7 +195,6 @@ private void run( if (approvalStrategy != null && approvalPolicy != null) { context = context.withApprovalStrategy(approvalStrategy).withApprovalPolicy(approvalPolicy); } - context = ToolLoopPolicies.attach(context, ToolLoopPolicy.maxIterations(maxToolRounds)); - runtime.execute(context, session); + return ToolLoopPolicies.attach(context, ToolLoopPolicy.maxIterations(maxToolRounds)); } } diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentTerminal.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentTerminal.java index c9f4efe57..03a914150 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentTerminal.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentTerminal.java @@ -38,11 +38,15 @@ public interface AgentTerminal extends AutoCloseable { String readLine(String prompt); /** - * Read a single answer, without waiting for Enter where the terminal allows it. + * Ask a question and read the answer. + * + *

The answer is a line, terminated with Enter, on both consoles. A real terminal could read a + * single key instead, but only by opening a second reader on the keyboard, and the one reader it + * has is busy offering the prompt that stays visible while the agent works — which is worth more + * than saving an Enter on a question that is asked a few times a session. * * @param prompt the question to show - * @return the answer in lower case (a single key, or a whole line on a plain stream), or - * {@code null} at end of input + * @return the answer in lower case, or {@code null} at end of input */ @Nullable String readKey(String prompt); @@ -67,6 +71,22 @@ public interface AgentTerminal extends AutoCloseable { */ boolean pinsStatus(); + /** + * Whether the user has already typed a line that nobody has read yet. + * + *

This is what makes the prompt useful during a turn: a console that keeps reading while the + * agent works can say so, and the turn is then cut short and the line answered instead of being + * made to wait for an answer nobody wants any more. The line stays queued — the caller reads it + * with {@link #readLine} as usual. + * + *

A console that reads only when asked has nothing pending by definition, which is the default. + * + * @return {@code true} when a line is waiting + */ + default boolean hasPendingInput() { + return false; + } + /** * Ask the terminal to run {@code action} when the user presses shift+tab at the prompt. * diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java index 32e77fd55..a120a6865 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java @@ -7,6 +7,8 @@ import java.io.IOException; import java.util.List; import java.util.Locale; +import java.util.concurrent.BlockingQueue; +import java.util.concurrent.LinkedBlockingQueue; import org.jline.keymap.KeyMap; import org.jline.reader.Binding; import org.jline.reader.EndOfFileException; @@ -15,7 +17,6 @@ import org.jline.reader.Reference; import org.jline.reader.UserInterruptException; import org.jline.reader.impl.completer.StringsCompleter; -import org.jline.terminal.Attributes; import org.jline.terminal.Terminal; import org.jline.terminal.TerminalBuilder; import org.jline.utils.AttributedString; @@ -26,12 +27,15 @@ /** * An {@link AgentTerminal} on a real terminal, via JLine: line editing and history at the prompt, tab - * completion of the commands, a status line pinned to the bottom of the window, and single-key - * answers. + * completion of the commands, a status line pinned to the bottom of the window, and a prompt that is + * there at all times — including while the agent is working. * - *

Streamed output goes through {@link LineReader#printAbove(String)}, which scrolls it in above - * the prompt while the bottom block stays where it is — the one thing a plain {@code println} cannot - * do. Nothing is ever redrawn above that block, so the scrollback stays exactly as it was written. + *

That last part is why a thread of its own owns the keyboard (see {@code startReading}): it sits + * in {@code readLine} for the whole session and puts what is typed on a queue, and every read in this + * class is served from that queue. Streamed output goes through {@link LineReader#printAbove(String)}, + * which scrolls it in above the prompt while the bottom block stays where it is — the one thing a + * plain {@code println} cannot do. Nothing is ever redrawn above that block, so the scrollback stays + * exactly as it was written. * *

{@link #open} returns {@code null} instead of throwing when there is no usable terminal (piped * input, a "dumb" terminal, a missing native provider); the caller then uses {@link PlainTerminal}. @@ -40,11 +44,23 @@ */ public final class JLineTerminal implements AgentTerminal { + /** + * What is queued in place of a line when input ends. + * + *

A queue of lines cannot carry "no more lines" as a value, and the reader thread is the only + * one that learns it. The sentinel is put back on every take, so end of input stays end of input + * for every later caller instead of turning back into "nothing typed yet". + */ + private static final String END_OF_INPUT = "\u0000end-of-input"; + private final Terminal terminal; private final LineReader reader; private final Status status; private final Ansi ansi; + private final BlockingQueue typed = new LinkedBlockingQueue<>(); private volatile boolean reading; + private volatile boolean closed; + private @Nullable Thread input; private JLineTerminal(Terminal terminal, LineReader reader, Status status, Ansi ansi) { this.terminal = terminal; @@ -97,43 +113,96 @@ public void line(String text) { } } + /** + * Start the one thread that owns the keyboard, if it is not running yet. + * + *

Why a thread of its own. The prompt is supposed to be there at all times — while the + * agent is working, not only between turns — and only a thread that sits in {@code readLine} can + * offer that. Everything else then writes through {@link LineReader#printAbove}, which scrolls + * text in above the prompt and leaves it where it is. + * + *

Why exactly one. A terminal has one keyboard, and two threads reading it take turns at + * random. So every read in this class — a request, an approval answer — is served from the one + * queue this thread fills, and nothing else ever reads the terminal. + * + * @param prompt the prompt, kept for the whole session (the first caller decides it) + */ + private synchronized void startReading(String prompt) { + if (input != null) { + return; + } + input = new Thread( + () -> { + while (!closed) { + reading = true; + try { + typed.put(reader.readLine(prompt)); + } catch (UserInterruptException e) { + // Ctrl-C: drop what was typed and ask again, as before. + } catch (EndOfFileException e) { + typed.offer(END_OF_INPUT); // Ctrl-D + return; + } catch (InterruptedException e) { + Thread.currentThread().interrupt(); + return; + } catch (RuntimeException e) { + typed.offer(END_OF_INPUT); // the terminal is gone; stop reading it + return; + } finally { + reading = false; + } + } + }, + "agent-input"); + input.setDaemon(true); + input.start(); + } + @Override public @Nullable String readLine(String prompt) { - reading = true; + startReading(prompt); try { - return reader.readLine(prompt); - } catch (UserInterruptException e) { - return ""; // Ctrl-C: drop the line, ask again - } catch (EndOfFileException e) { - return null; // Ctrl-D - } finally { - reading = false; + return take(typed.take()); + } catch (InterruptedException e) { + Thread.currentThread().interrupt(); + return null; } } + /** + * Hand out a queued line, keeping the end-of-input marker in the queue. + * + * @param line what came off the queue + * @return the line, or {@code null} when input has ended + */ + private @Nullable String take(String line) { + if (END_OF_INPUT.equals(line)) { + typed.offer(END_OF_INPUT); + return null; + } + return line; + } + + @Override + public boolean hasPendingInput() { + return !typed.isEmpty(); + } + @Override public @Nullable String readKey(String prompt) { - terminal.writer().print(prompt); - terminal.writer().flush(); - Attributes saved = terminal.enterRawMode(); + // The prompt belongs to the input thread and cannot be changed while it is waiting, so the + // question is printed as an ordinary line above it and answered in the same input line as + // everything else. That costs an Enter, and buys the one thing worth more: a prompt that is + // there while the agent works. A single-key read here would need a second reader on the same + // terminal, and the two would take turns at random. + line(prompt); + startReading("you> "); try { - int key = terminal.reader().read(); - if (key < 0 || key == 4) { // end of input, Ctrl-D - return null; - } - if (key == 3) { // Ctrl-C: treat as "no", the safe answer - terminal.writer().print("^C" + System.lineSeparator()); - terminal.writer().flush(); - return "n"; - } - String answer = String.valueOf((char) key).toLowerCase(Locale.ROOT); - terminal.writer().print(answer + System.lineSeparator()); - terminal.writer().flush(); - return answer; - } catch (IOException e) { + String answer = take(typed.take()); + return answer == null ? null : answer.trim().toLowerCase(Locale.ROOT); + } catch (InterruptedException e) { + Thread.currentThread().interrupt(); return null; - } finally { - terminal.setAttributes(saved); } } @@ -213,6 +282,11 @@ public Ansi ansi() { @Override public void close() { + closed = true; + Thread reading = input; + if (reading != null) { + reading.interrupt(); + } try { status.update(List.of()); status.close(); diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java index 852bab229..df5f3a99a 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java @@ -419,29 +419,25 @@ static ConsoleSession turn( TurnActivity activity) throws InterruptedException { ConsoleSession session = new ConsoleSession(terminal, fileSystem); - // Atmosphere's execute() is synchronous: it returns only once the whole turn, tool rounds - // included, has finished. Running it on this thread would leave nobody to drive the activity - // line -- which is exactly what happened: the block sat on "… waiting for input …" for the - // entire turn, so the spinner and its word were never seen. - Thread worker = new Thread( - () -> { - try { - runner.run(message, history, session); - } catch (RuntimeException e) { - session.error(e); - } - }, - "agent-turn"); - worker.setDaemon(true); - worker.start(); - boolean finished = awaitWithActivity(session, terminal, () -> stateLine.apply(session), activity); - worker.join(Duration.ofSeconds(5).toMillis()); + // The turn runs on a thread of Atmosphere's, not this one: execute() blocks until the whole + // turn including every tool round is done, which would leave nobody to drive the activity + // line -- the block sat on "… waiting for input …" for entire turns until this changed. + // start() adds the half that makes the always-present prompt worth having: the handle can + // close the stream the model is answering on, so a request typed mid-turn takes effect now. + org.atmosphere.ai.ExecutionHandle handle; + try { + handle = runner.start(message, history, session); + } catch (RuntimeException e) { + session.error(e); + handle = org.atmosphere.ai.ExecutionHandle.completed(); + } + TurnEnd end = awaitWithActivity(session, terminal, () -> stateLine.apply(session), activity, handle); history.add(ChatMessage.user(message)); callLog.add(turnNumber, session.rounds()); if (!session.text().isEmpty()) { history.add(ChatMessage.assistant(session.text())); } - if (!finished) { + if (end == TurnEnd.TIMED_OUT) { session.error(new IllegalStateException("turn did not finish within " + TURN_TIMEOUT)); } return session; @@ -585,14 +581,26 @@ private static long handleCommand( * @param stateLine the second row of the block, asked again on every redraw so the context figure * moves while the turn runs rather than standing still until the next prompt * @param activity paused while an approval question is open - * @return {@code true} when the turn finished within {@link #TURN_TIMEOUT} + * @param handle stops the running turn when the user types instead of waiting + * @return how the turn ended * @throws InterruptedException if interrupted while waiting */ - private static boolean awaitWithActivity( + /** How a turn ended: on its own, because the user typed something, or because it ran too long. */ + enum TurnEnd { + /** The model produced its final answer (or errored). */ + FINISHED, + /** The user typed while it was working; the turn was cut short and that line is next. */ + INTERRUPTED, + /** Nothing arrived within {@link #TURN_TIMEOUT}. */ + TIMED_OUT + } + + static TurnEnd awaitWithActivity( ConsoleSession session, AgentTerminal terminal, java.util.function.Supplier stateLine, - TurnActivity activity) + TurnActivity activity, + org.atmosphere.ai.ExecutionHandle handle) throws InterruptedException { long start = System.nanoTime(); int frame = 0; @@ -601,10 +609,20 @@ private static boolean awaitWithActivity( while (!session.await(ACTIVITY_INTERVAL)) { long seconds = (System.nanoTime() - start) / 1_000_000_000L; if (seconds > TURN_TIMEOUT.toSeconds()) { - return false; + return TurnEnd.TIMED_OUT; } if (activity.isPaused()) { - continue; // an approval question owns the terminal until it is answered + continue; // an approval question is waiting for its answer + } + if (terminal.hasPendingInput()) { + // Something was typed while the agent was working. Stop the turn rather than finish a + // request that has been overtaken: the handle closes the stream the model is answering + // on. The line itself stays queued and becomes the next message, so what the model + // produced so far is kept and the new instruction follows it. + handle.cancel(); + terminal.line(terminal.ansi().yellow("(interrupted — taking your message)")); + terminal.status(List.of(IDLE_LINE, stateLine.get())); + return TurnEnd.INTERRUPTED; } terminal.status(List.of( activityLine( @@ -617,7 +635,7 @@ private static boolean awaitWithActivity( stateLine.get())); } terminal.status(List.of(IDLE_LINE, stateLine.get())); - return true; + return TurnEnd.FINISHED; } /** @@ -783,7 +801,13 @@ private static long compact( "agent-compact"); worker.setDaemon(true); worker.start(); - if (!awaitWithActivity(session, terminal, () -> "/compact", new TurnActivity()) + if (awaitWithActivity( + session, + terminal, + () -> "/compact", + new TurnActivity(), + org.atmosphere.ai.ExecutionHandle.completed()) + != TurnEnd.FINISHED || session.text().isBlank()) { terminal.line("(compact failed; history kept)"); return 0; diff --git a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt index 434009336..52adf35ea 100644 --- a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt +++ b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt @@ -15,7 +15,8 @@ Commands (everything else is sent to the model): /clear drop the history (/reset, /new) /exit leave (/quit) -Shift+Tab at the prompt switches the approval mode without typing a command. +The prompt stays at the bottom while the agent works: you can type at any time. A line typed +during a turn stops it and is sent as the next message. Shift+Tab switches the approval mode. In manual mode every tool that writes or runs a command asks first: [y]es runs it once, [n]o tells the model the user cancelled it, [a]uto stops asking for the rest of the session. diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleSessionTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleSessionTest.java index a58f6f99b..c687dc5e2 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleSessionTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleSessionTest.java @@ -35,6 +35,12 @@ class ConsoleSessionTest { /** A terminal that records what it was told to print. */ private static final class RecordingTerminal implements AgentTerminal { private final List lines = new ArrayList<>(); + private volatile boolean pending; + + @Override + public boolean hasPendingInput() { + return pending; + } @Override public void line(String text) { @@ -107,6 +113,61 @@ void aMultiLineToolResultStaysOnOneLineToo() { assertEveryLineIsOneLine(); } + /** A cancellable turn that only records whether it was stopped. */ + private static final class RecordingHandle implements org.atmosphere.ai.ExecutionHandle { + private final java.util.concurrent.CompletableFuture done = + new java.util.concurrent.CompletableFuture<>(); + private volatile boolean cancelled; + + @Override + public void cancel() { + cancelled = true; + done.complete(null); + } + + @Override + public boolean isDone() { + return done.isDone(); + } + + @Override + public java.util.concurrent.CompletableFuture whenDone() { + return done; + } + } + + @Test + void typingWhileTheAgentWorksStopsTheTurnAndTheLineIsNotConsumed() throws InterruptedException { + // the whole point of a prompt that is there during a turn: a request that has been overtaken + // must not keep running, and what was typed stays queued to become the next message + ConsoleSession session = session(); // never completed: only the typing can end this wait + RecordingHandle handle = new RecordingHandle(); + terminal.pending = true; + + LocalAgent.TurnEnd end = + LocalAgent.awaitWithActivity(session, terminal, () -> "state", new TurnActivity(), handle); + + assertThat(end, is(LocalAgent.TurnEnd.INTERRUPTED)); + assertThat("the stream the model is answering on is closed", handle.cancelled, is(true)); + assertThat( + "and it says so rather than looking like a finished answer", + terminal.lines.stream().anyMatch(l -> l.contains("interrupted")), + is(true)); + } + + @Test + void aTurnThatFinishesOnItsOwnIsNotCancelled() throws InterruptedException { + ConsoleSession session = session(); + RecordingHandle handle = new RecordingHandle(); + session.complete(); + + LocalAgent.TurnEnd end = + LocalAgent.awaitWithActivity(session, terminal, () -> "state", new TurnActivity(), handle); + + assertThat(end, is(LocalAgent.TurnEnd.FINISHED)); + assertThat(handle.cancelled, is(false)); + } + @Test void theContextFigureGrowsWhileTheTurnRuns() { // it used to be rendered once before the turn and handed over as a fixed string, so it stood From cc91731e869b86001ebe2c8eaf0efc477f9dba1d Mon Sep 17 00:00:00 2001 From: Bernard Ladenthin Date: Wed, 23 Sep 2026 18:26:27 +0200 Subject: [PATCH 20/36] llama-atmosphere-agent: frame the input into the bottom block, and classify every tool Two things, both from using it. The input box. The prompt was an ordinary "you>" line above the pinned block, so every Enter left one behind and a few empty ones printed a column of them. The input now sits inside the frame: the top rule is the first line of the reader's prompt, the bottom rule is the first line of the status block -- JLine draws the status below the prompt, so that is the only way to get the input inside a frame at all. Both edges call the same rule(), or they drift apart on a resize. ERASE_LINE_ON_FINISH clears the box on Enter and the typed line is echoed above it, so the transcript keeps what was asked without the leftovers. The reader's "!" history expansion is switched off in the same place; it silently rewrites a request like: git commit -m "fixed!" The tool classification. A read-only "ls" ran without asking in manual mode, which is correct -- reading tools are deliberately never gated -- but nothing checked that the gate list still matches the tools actually offered. It is a list of names, so a tool added or renamed upstream would drop out of it and then run unasked, silently. READ_ONLY_TOOLS now names the other half explicitly and a test asserts every offered tool is in exactly one of the two sets, and that neither set names a tool nobody offers. Same class as a stale spotbugs exclusion. Also documented what cannot be done: the block cannot stay put while the user scrolls the terminal's own scrollback. That needs the alternate screen buffer, which would give up the scrollback and the append-only property this console rests on. 121 tests green. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01E2h8gXyyE5UeimkL9vQv9G --- CLAUDE.md | 26 ++++++++++- llama-atmosphere-agent/README.md | 34 ++++++++++----- .../atmosphere/ConsoleApprovalStrategy.java | 11 +++++ .../llama/atmosphere/JLineTerminal.java | 43 +++++++++++++++---- .../net/ladenthin/llama/atmosphere/help.txt | 6 ++- .../ConsoleApprovalStrategyTest.java | 27 ++++++++++++ 6 files changed, 125 insertions(+), 22 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 230841be6..4c5d2708f 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -2246,7 +2246,14 @@ are decisions, not details: `rename` — reading tools never ask) and a `ConsoleApprovalStrategy`; `ToolExecutionHelper` then blocks the tool loop before the executor runs and turns a denial into the tool result `{"status":"cancelled","message":"Action cancelled by user"}` for the model. Do not reimplement - that message. `ApprovalWireTest` pins both halves over the real server. + that message. `ApprovalWireTest` pins both halves over the real server. **The reading tools + (`ls`, `read_file`, `glob`, `grep`) never ask, and that is a decision, not an omission**: gating + them would make the question so frequent it stops being read. Because the gate is a list of + *names*, a tool upstream adds or renames would drop out of it and then run unasked — so + `READ_ONLY_TOOLS` names the other half explicitly and + `ConsoleApprovalStrategyTest.everyOfferedToolIsEitherGatedOrDeclaredReadOnly` asserts every offered + tool is in exactly one of the two sets, and that neither set names a tool nobody offers. Same class + as the stale `spotbugs-exclude.xml` entries: an allowlist that silently stops matching. 2. **One-shot (`--prompt`) denies a gated call** instead of auto-approving it — `--auto` is the deliberate opt-in. Atmosphere itself fails closed when no strategy is wired, and this keeps that direction: an unattended run must not be the most permissive one. @@ -2366,7 +2373,22 @@ are decisions, not details: status numbers in an `AtomicLong`/`AtomicBoolean` rather than locals, so the widget (which runs inside the reader) can re-render the pinned row with what the last turn left behind. -11. **The prompt stays at the bottom during a turn, and typing stops the turn.** One thread inside +11. **The input is framed into the pinned block, and the prompt stays there during a turn; typing + stops the turn.** The frame is two halves that must be read together: the **top** rule is the first + line of the reader's *prompt* (`rule() + newline + "> "`, rebuilt on every read because the window + can be resized), the **bottom** rule is the first line of the *status* block — JLine draws the + status below the prompt, so that is the only way to get the input inside a frame at all. Both call + the same `rule()`, or the two edges drift apart on a resize. `AgentTerminal.readLine`'s `prompt` + argument is consequently **ignored** here. `ERASE_LINE_ON_FINISH` removes the box on Enter — + without it every submitted line leaves a rule pair in the scrollback, and a few empty Enters print a + wall of them — and the reader thread echoes the line above as `› text` so the transcript keeps it. + `DISABLE_EVENT_EXPANSION` is set in the same builder because the reader's default treats `!` as a + shell history expansion, which silently rewrites a request like `git commit -m "fixed!"`. + **What cannot be done, asked and answered:** keep the block visible while the *user* scrolls the + terminal's scrollback. That needs the alternate screen buffer, i.e. a full-screen application, which + would give up the scrollback and the append-only property the whole console design rests on. + + **The prompt stays at the bottom during a turn, and typing stops the turn.** One thread inside `JLineTerminal` (`startReading`) sits in `readLine` for the whole session and fills a queue; **every** read in that class is served from it, because a terminal has one keyboard and two threads reading it take turns at random. That is why `readKey` no longer reads a single key in raw mode: the diff --git a/llama-atmosphere-agent/README.md b/llama-atmosphere-agent/README.md index b7d71be17..2b844c322 100644 --- a/llama-atmosphere-agent/README.md +++ b/llama-atmosphere-agent/README.md @@ -294,23 +294,37 @@ output **line by line while it runs** (dimmed, `│ `-prefixed) instead of dumpi also keeps the pipe drained; a process whose output nobody reads blocks once the buffer is full, and on Windows that buffer is about 4 KB. -**The block at the bottom has two rows**, below a rule, and both are always present: +**The bottom of the window is one framed block**: the input line between two rules, then what the agent +is doing and the session state. ``` ──────────────────────────────────────────────────────────────────── +> add a test for the parser +──────────────────────────────────────────────────────────────────── ⠙ Fettling… (run_command 47s of 61s · 2 tool calls) [/path/to/project · ⏸ manual · ctx ~3.1k/16k · 9 tools · local-model] ``` -The first row is what the agent is doing — `… waiting for input …` when it is your turn, the spinner -with the running tool and its elapsed time while it works. The second row is the session's state. Two -fixed rows rather than one changing one: a block that changes height makes the output above it jump on -every refresh. - -Note where "the bottom" is: the bottom of the *window*, not the line under the cursor. In a tall -terminal with only a few lines of output there is a gap between your prompt and the block. On a plain -stream (piped input, a one-shot run) nothing can be pinned, so the state line is printed above the -prompt instead and the activity row is dropped rather than repeated into the log. +The top rule is the first line of the reader's prompt, the three below it are the pinned block — which +is how the input ends up inside the frame at all. There is no `you>`: the box already says where the +input is. On Enter the box is erased and the line is echoed above it as `› your text`, so the transcript +keeps what was asked instead of accumulating leftover rules. (`!` history expansion is switched off in +the same place, or a request like `git commit -m "fixed!"` would be rewritten silently.) + +The two lower rows are always both present: the activity row — `… waiting for input …` when it is your +turn, the spinner with the running tool and its elapsed time while it works — and the session's state. +Two fixed rows rather than one changing one: a block that changes height makes the output above it jump +on every refresh. + +**On scrolling.** The block stays put while the agent writes: JLine keeps those lines out of the +terminal's scroll region. It cannot stay while *you* scroll the terminal's own scrollback with the +mouse — then the whole viewport moves and no program on this side of the terminal has a say. Staying +visible through that needs the alternate screen buffer, i.e. a full-screen application, which would +trade away the scrollback and the "written once, never redrawn" property this console is built on. The +established terminal agents behave the same way. + +On a plain stream (piped input, a one-shot run) nothing can be pinned, so the state line is printed +above the prompt instead and the activity row is dropped rather than repeated into the log. **The status line** above the prompt reads `[/path/to/project · ⏸ manual · ctx ~3.1k/16k · 9 tools · local-model]`: the workspace, the approval diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleApprovalStrategy.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleApprovalStrategy.java index e7f806b3b..2955d33bd 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleApprovalStrategy.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleApprovalStrategy.java @@ -42,6 +42,17 @@ public final class ConsoleApprovalStrategy implements ApprovalStrategy { public static final Set GATED_TOOLS = Set.of(ShellTool.TOOL_NAME, "write_file", "edit_file", "delete", "rename"); + /** + * The tools that deliberately never ask, because they only read. + * + *

It exists so that "does not ask" is a decision rather than the absence of one. + * {@link #GATED_TOOLS} is a list of names, so a tool that is added or renamed upstream falls out + * of it silently and then runs unasked — the same class of quiet breakage as a stale exclusion + * file. {@code ConsoleApprovalStrategyTest} asserts every registered tool is in exactly one of the + * two sets, so such a change fails the build instead of the session. + */ + public static final Set READ_ONLY_TOOLS = Set.of("ls", "read_file", "glob", "grep"); + private static final int ARGUMENT_PREVIEW_CHARS = 300; private final AtomicReference mode; diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java index a120a6865..f338c3dc6 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java @@ -85,6 +85,13 @@ private JLineTerminal(Terminal terminal, LineReader reader, Status status, Ansi LineReader reader = LineReaderBuilder.builder() .terminal(terminal) .completer(new StringsCompleter(completions)) + // The input sits in a framed box at the bottom. Without this the box would be + // left behind in the scrollback on every Enter, so a few empty lines would print + // a wall of rules; the line the user typed is echoed above it instead. + .option(LineReader.Option.ERASE_LINE_ON_FINISH, true) + // "!" is a shell history expansion in the reader's default configuration, which + // silently rewrites a request like: git commit -m "fixed!" + .option(LineReader.Option.DISABLE_EVENT_EXPANSION, true) .build(); return new JLineTerminal(terminal, reader, Status.getStatus(terminal), Ansi.detect()); } catch (IOException | RuntimeException e) { @@ -125,9 +132,8 @@ public void line(String text) { * random. So every read in this class — a request, an approval answer — is served from the one * queue this thread fills, and nothing else ever reads the terminal. * - * @param prompt the prompt, kept for the whole session (the first caller decides it) */ - private synchronized void startReading(String prompt) { + private synchronized void startReading() { if (input != null) { return; } @@ -136,7 +142,15 @@ private synchronized void startReading(String prompt) { while (!closed) { reading = true; try { - typed.put(reader.readLine(prompt)); + // Built fresh each time: the rule has to match the window, which can be + // resized between two requests. + String line = reader.readLine(rule() + System.lineSeparator() + "> "); + if (!line.isBlank()) { + // The box is erased on Enter, so the conversation would lose what was + // asked. Echoing it above keeps the transcript readable. + line(ansi.bold("› " + line.strip())); + } + typed.put(line); } catch (UserInterruptException e) { // Ctrl-C: drop what was typed and ask again, as before. } catch (EndOfFileException e) { @@ -158,9 +172,21 @@ private synchronized void startReading(String prompt) { input.start(); } + /** + * The rule that frames the input box, as wide as the window. + * + * @return a line of {@code ─} + */ + private String rule() { + return "─".repeat(Math.max(10, terminal.getSize().getColumns() - 1)); + } + @Override public @Nullable String readLine(String prompt) { - startReading(prompt); + // The prompt is the box this terminal draws, so the caller's is ignored: a "you> " in front of + // an input line that already sits in a frame is noise, and the frame cannot be handed in as a + // string because it is rebuilt on every window size. + startReading(); try { return take(typed.take()); } catch (InterruptedException e) { @@ -196,7 +222,7 @@ public boolean hasPendingInput() { // there while the agent works. A single-key read here would need a second reader on the same // terminal, and the two would take turns at random. line(prompt); - startReading("you> "); + startReading(); try { String answer = take(typed.take()); return answer == null ? null : answer.trim().toLowerCase(Locale.ROOT); @@ -212,14 +238,15 @@ public void status(List lines) { status.update(List.of()); return; } - // A rule above the block separates it from the scrollback, the way the established terminal - // agents frame their input. + // The rule on top of this block is the bottom edge of the input box: the reader draws the top + // edge as the first line of its prompt, so the two together frame the input the way the + // established terminal agents do. // One row must never wrap: a wrapped row occupies two screen lines, the reserved region is // sized in lines, and everything below it is then drawn in the wrong place -- which is how a // long summary tore the block apart. int width = Math.max(10, terminal.getSize().getColumns() - 1); List block = new java.util.ArrayList<>(); - block.add(new AttributedString("─".repeat(width), AttributedStyle.DEFAULT.foreground(AttributedStyle.BRIGHT))); + block.add(new AttributedString(rule(), AttributedStyle.DEFAULT.foreground(AttributedStyle.BRIGHT))); for (String line : lines) { block.add( new AttributedString(fit(line, width), AttributedStyle.DEFAULT.foreground(AttributedStyle.BRIGHT))); diff --git a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt index 52adf35ea..ab5fbc0e5 100644 --- a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt +++ b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt @@ -15,9 +15,11 @@ Commands (everything else is sent to the model): /clear drop the history (/reset, /new) /exit leave (/quit) -The prompt stays at the bottom while the agent works: you can type at any time. A line typed +The input line sits in the box at the bottom and is there while the agent works: you can type at +any time. A line typed during a turn stops it and is sent as the next message. Shift+Tab switches the approval mode. In manual mode every tool that writes or runs a command asks first: [y]es runs it once, [n]o tells the model the user cancelled it, [a]uto stops asking for the rest of the session. -Reading tools never ask. An unknown /command is sent to the model, not rejected. +Reading tools (ls, read_file, glob, grep) never ask — they cannot change anything. +An unknown /command is sent to the model, not rejected. diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleApprovalStrategyTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleApprovalStrategyTest.java index 4b99f0b78..e3264a4ab 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleApprovalStrategyTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleApprovalStrategyTest.java @@ -25,6 +25,33 @@ class ConsoleApprovalStrategyTest { + @org.junit.jupiter.api.Test + void everyOfferedToolIsEitherGatedOrDeclaredReadOnly() { + // The gate is a list of names, so a tool that upstream adds or renames drops out of it and + // then runs without asking -- in manual mode, silently. This is the check that turns that into + // a red build: every tool the model is offered must be classified, one way or the other. + java.util.List offered = + new java.util.ArrayList<>(WorkspaceTools.all(new WorkspaceTools.ReadTracker()).stream() + .map(ToolDefinition::name) + .toList()); + offered.add(ShellTool.TOOL_NAME); + + for (String tool : offered) { + boolean gated = ConsoleApprovalStrategy.GATED_TOOLS.contains(tool); + boolean readOnly = ConsoleApprovalStrategy.READ_ONLY_TOOLS.contains(tool); + assertThat( + tool + " is in neither set: decide whether it has to ask before it runs", + gated || readOnly, + is(true)); + assertThat(tool + " cannot be both", gated && readOnly, is(false)); + } + assertThat( + "a set that names tools nobody offers is stale", + offered.containsAll(ConsoleApprovalStrategy.GATED_TOOLS), + is(true)); + assertThat(offered.containsAll(ConsoleApprovalStrategy.READ_ONLY_TOOLS), is(true)); + } + private final ByteArrayOutputStream console = new ByteArrayOutputStream(); private final TurnActivity activity = new TurnActivity(); From 335b42d2a9112febe011e8f5f1c96844dd6416f6 Mon Sep 17 00:00:00 2001 From: Bernard Ladenthin Date: Wed, 23 Sep 2026 19:27:49 +0200 Subject: [PATCH 21/36] llama-atmosphere-agent: the input box left a rule behind on every Enter Reported as "weird at the bottom when I press Enter a few times", and it was two defects, both mine from the previous commit. The rule above the input was the first line of a two-line prompt, while ERASE_LINE_ON_FINISH erases exactly ONE line -- so every Enter left the rule behind and holding Enter drew a column of them. The rule is now ordinary output (AgentTerminal.separator(), drawn before each read and suppressed when nothing has been printed since, so repeated Enters draw nothing); the prompt is one line again. The look is unchanged: top rule as output, input, then the status block whose own first line is the bottom rule. The stray "[?1h" above the prompt was separate: line() wrote straight to the terminal whenever the reader was not inside readLine, and that instant is exactly when the next read emits its init sequence. The two interleaved and half an escape sequence landed in the scrollback as text. Once the reader thread exists, everything goes through printAbove. JLineTerminalTest is the new part worth keeping: JLineTerminal.over() takes a terminal built over two streams, which renders exactly like a TTY, so the screen can be asserted on the bytes. Two findings that are not guessable and are written down: the test terminal needs stdoutEncoding as well as encoding or every rule arrives as "?", and a box character goes out as UTF-8 from printAbove but as the DEC line-drawing set inside a prompt -- so the first version of the counter passed against the exact bug it was written for. Verified by restoring the two-line prompt and watching it report 5 rules where it now reports 1. 125 tests green. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01E2h8gXyyE5UeimkL9vQv9G --- CLAUDE.md | 29 +++- llama-atmosphere-agent/README.md | 18 ++- .../llama/atmosphere/AgentTerminal.java | 15 ++ .../llama/atmosphere/JLineTerminal.java | 83 +++++++---- .../llama/atmosphere/LocalAgent.java | 1 + .../llama/atmosphere/JLineTerminalTest.java | 132 ++++++++++++++++++ 6 files changed, 242 insertions(+), 36 deletions(-) create mode 100644 llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/JLineTerminalTest.java diff --git a/CLAUDE.md b/CLAUDE.md index 4c5d2708f..4fe77b4ea 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -2376,12 +2376,29 @@ are decisions, not details: 11. **The input is framed into the pinned block, and the prompt stays there during a turn; typing stops the turn.** The frame is two halves that must be read together: the **top** rule is the first line of the reader's *prompt* (`rule() + newline + "> "`, rebuilt on every read because the window - can be resized), the **bottom** rule is the first line of the *status* block — JLine draws the - status below the prompt, so that is the only way to get the input inside a frame at all. Both call - the same `rule()`, or the two edges drift apart on a resize. `AgentTerminal.readLine`'s `prompt` - argument is consequently **ignored** here. `ERASE_LINE_ON_FINISH` removes the box on Enter — - without it every submitted line leaves a rule pair in the scrollback, and a few empty Enters print a - wall of them — and the reader thread echoes the line above as `› text` so the transcript keeps it. + can be resized) — **no.** That is how it was built and it was wrong: `ERASE_LINE_ON_FINISH` erases + exactly **one** line, so a two-line prompt leaves its rule behind on every Enter, and holding Enter + draws a column of them. The **top** rule is therefore ordinary output (`AgentTerminal.separator()`, + printed before each read, suppressed when nothing has been printed since so repeated Enters draw + nothing), the **bottom** rule is the first line of the *status* block — JLine draws the status below + the prompt, so that is the only way to get the input inside a frame at all. Both call the same + `rule()`, or the two edges drift apart on a resize. The prompt is `"> "`, one line, and + `AgentTerminal.readLine`'s `prompt` argument is consequently **ignored** here. + `ERASE_LINE_ON_FINISH` removes the input line on Enter and the reader thread echoes it above as + `› text`, so the transcript keeps what was asked. + + **`JLineTerminalTest` is how any of this is checkable**: `JLineTerminal.over(Terminal, …)` takes a + terminal built over two streams, which renders exactly like a TTY, so the screen can be asserted on + the emitted bytes. Two things that cost an hour each and are not guessable: the test terminal needs + **`stdoutEncoding`** as well as `encoding`, or every `─` arrives as `?`; and a box character is + written as UTF-8 from `printAbove` but as the **DEC line-drawing set** (`ESC(0` + `q`s + `ESC(B`) + inside a *prompt*, so a counter that looks only for `─` passes against the exact bug it was + written for — verified by putting the two-line prompt back and watching the test go from 1 rule to 5. + + **The `[?1h` that appeared as text above the prompt** was a second, separate defect: `line()` wrote + straight to the terminal whenever the reader was not inside `readLine`, and that instant is exactly + when the next read is emitting its init sequence, so the two interleaved and half an escape sequence + landed in the scrollback. Once the reader thread exists, **everything** goes through `printAbove`. `DISABLE_EVENT_EXPANSION` is set in the same builder because the reader's default treats `!` as a shell history expansion, which silently rewrites a request like `git commit -m "fixed!"`. **What cannot be done, asked and answered:** keep the block visible while the *user* scrolls the diff --git a/llama-atmosphere-agent/README.md b/llama-atmosphere-agent/README.md index 2b844c322..dd27edc5a 100644 --- a/llama-atmosphere-agent/README.md +++ b/llama-atmosphere-agent/README.md @@ -305,11 +305,19 @@ is doing and the session state. [/path/to/project · ⏸ manual · ctx ~3.1k/16k · 9 tools · local-model] ``` -The top rule is the first line of the reader's prompt, the three below it are the pinned block — which -is how the input ends up inside the frame at all. There is no `you>`: the box already says where the -input is. On Enter the box is erased and the line is echoed above it as `› your text`, so the transcript -keeps what was asked instead of accumulating leftover rules. (`!` history expansion is switched off in -the same place, or a request like `git commit -m "fixed!"` would be rewritten silently.) +The top rule is ordinary output, printed once before each read; the three rows below are the pinned +block. There is no `you>`: the box already says where the input is. On Enter the input line is erased +and echoed above as `› your text`, so the transcript keeps what was asked. (`!` history expansion is +switched off in the same place, or a request like `git commit -m "fixed!"` would be rewritten silently.) + +**Why the top rule is not part of the prompt**, although that is the obvious way to draw it: the line +reader erases exactly **one** line when the input is submitted, so a two-line prompt leaves its rule +behind on every Enter — hold Enter and you get a column of them. It was built that way first and the +symptom was reported within the hour. `JLineTerminalTest` now drives a real line reader over a pair of +streams and counts the rules in the bytes it emits, which is the only way to see this without a +console. Note what that test had to learn: a box character goes out as UTF-8 `─` from ordinary output +but as the DEC line-drawing set (`ESC(0` + a row of `q`) inside a prompt, so counting only the first +form passed against the very bug it was written for. The two lower rows are always both present: the activity row — `… waiting for input …` when it is your turn, the spinner with the running tool and its elapsed time while it works — and the session's state. diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentTerminal.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentTerminal.java index 03a914150..89bb53f23 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentTerminal.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentTerminal.java @@ -71,6 +71,21 @@ public interface AgentTerminal extends AutoCloseable { */ boolean pinsStatus(); + /** + * Draw the rule that separates the conversation from the input line, unless it is already there. + * + *

It is ordinary output rather than part of the prompt, and that is the whole point: a prompt + * of two lines is erased as one when the line is submitted, so a rule carried in the prompt + * survives every Enter and stacks up. Written once and never touched again, like everything else + * on this console. + * + *

Callers ask for it before every read; asking again with nothing printed in between draws + * nothing, so holding Enter does not produce a column of rules. + */ + default void separator() { + // nothing to frame on a plain stream + } + /** * Whether the user has already typed a line that nobody has read yet. * diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java index f338c3dc6..ea5ad9cdf 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java @@ -58,7 +58,7 @@ public final class JLineTerminal implements AgentTerminal { private final Status status; private final Ansi ansi; private final BlockingQueue typed = new LinkedBlockingQueue<>(); - private volatile boolean reading; + private volatile boolean ruleOnScreen; private volatile boolean closed; private @Nullable Thread input; @@ -82,26 +82,43 @@ private JLineTerminal(Terminal terminal, LineReader reader, Status status, Ansi terminal.close(); return null; } - LineReader reader = LineReaderBuilder.builder() - .terminal(terminal) - .completer(new StringsCompleter(completions)) - // The input sits in a framed box at the bottom. Without this the box would be - // left behind in the scrollback on every Enter, so a few empty lines would print - // a wall of rules; the line the user typed is echoed above it instead. - .option(LineReader.Option.ERASE_LINE_ON_FINISH, true) - // "!" is a shell history expansion in the reader's default configuration, which - // silently rewrites a request like: git commit -m "fixed!" - .option(LineReader.Option.DISABLE_EVENT_EXPANSION, true) - .build(); - return new JLineTerminal(terminal, reader, Status.getStatus(terminal), Ansi.detect()); + return over(terminal, completions); } catch (IOException | RuntimeException e) { // No terminal, no native provider, a restricted environment: the plain console still works. return null; } } + /** + * Wrap a terminal that has already been built. + * + *

The seam the tests use: a terminal over a pair of streams renders exactly like a real one — + * same escape sequences, same line reader — so what the screen would look like can be asserted on + * the emitted bytes, without a TTY. That is the only way to catch a drawing bug like a prompt whose + * height does not match what the reader erases when the line is submitted. + * + * @param terminal the terminal to drive + * @param completions the words tab completes + * @return the wrapper + */ + static JLineTerminal over(Terminal terminal, List completions) { + LineReader reader = LineReaderBuilder.builder() + .terminal(terminal) + .completer(new StringsCompleter(completions)) + // The input line is erased when submitted and echoed above instead, so it does + // not pile up in the scrollback. It erases exactly ONE line, which is why the + // prompt has to stay one line -- see startReading. + .option(LineReader.Option.ERASE_LINE_ON_FINISH, true) + // "!" is a shell history expansion in the reader's default configuration, which + // silently rewrites a request like: git commit -m "fixed!" + .option(LineReader.Option.DISABLE_EVENT_EXPANSION, true) + .build(); + return new JLineTerminal(terminal, reader, Status.getStatus(terminal), Ansi.detect()); + } + @Override public void line(String text) { + ruleOnScreen = false; if (text.indexOf('\n') >= 0 || text.indexOf('\r') >= 0) { // One call must be one screen line: the status block is sized in lines, so a multi-line // string handed over as "a line" desynchronises the reserved region. Callers fold their @@ -109,10 +126,11 @@ public void line(String text) { text.lines().forEach(this::line); return; } - if (reading) { - // Only while the line reader owns the screen: printAbove scrolls the text in above the - // prompt and redraws that prompt afterwards. Calling it when nobody is reading redraws a - // prompt that is not there, which is where the repeated "you>" lines came from. + if (input != null) { + // Once the reader thread exists it owns the screen, and nothing may write around it -- + // not even in the instant between two reads. A direct write there cuts into the escape + // sequence the next read is emitting and half of it lands in the scrollback as text: + // a stray "[?1h" above the prompt was exactly that. reader.printAbove(text); } else { terminal.writer().println(text); @@ -120,6 +138,20 @@ public void line(String text) { } } + @Override + public void separator() { + if (ruleOnScreen) { + // Already drawn, with nothing printed since: it is still directly above the input line. + return; + } + // Starts the reader if this is the first call: the rule belongs above a prompt, so drawing it + // is also the moment the prompt has to exist. Without this the very first box had no top edge + // -- the reader only started on the first read, which comes after. + startReading(); + reader.printAbove(rule()); + ruleOnScreen = true; + } + /** * Start the one thread that owns the keyboard, if it is not running yet. * @@ -140,14 +172,17 @@ private synchronized void startReading() { input = new Thread( () -> { while (!closed) { - reading = true; try { - // Built fresh each time: the rule has to match the window, which can be - // resized between two requests. - String line = reader.readLine(rule() + System.lineSeparator() + "> "); + // One line, and it has to stay one line: the reader erases a single line + // when the input is submitted, so a two-line prompt -- the rule above the + // input, which is where this started -- leaves the rule behind on every + // Enter, a column of them after a few. separator() draws the rule as + // ordinary output instead, which cannot be left behind because it was + // never part of the prompt. + String line = reader.readLine("> "); if (!line.isBlank()) { - // The box is erased on Enter, so the conversation would lose what was - // asked. Echoing it above keeps the transcript readable. + // The input line is erased on Enter, so the conversation would lose + // what was asked. Echoing it above keeps the transcript readable. line(ansi.bold("› " + line.strip())); } typed.put(line); @@ -162,8 +197,6 @@ private synchronized void startReading() { } catch (RuntimeException e) { typed.offer(END_OF_INPUT); // the terminal is gone; stop reading it return; - } finally { - reading = false; } } }, diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java index df5f3a99a..13558a8ce 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java @@ -271,6 +271,7 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream } else { terminal.line(terminal.ansi().dim(status)); } + terminal.separator(); String line = terminal.readLine("you> "); if (line == null) { return 0; diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/JLineTerminalTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/JLineTerminalTest.java new file mode 100644 index 000000000..11d1aaaef --- /dev/null +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/JLineTerminalTest.java @@ -0,0 +1,132 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import static org.hamcrest.MatcherAssert.assertThat; +import static org.hamcrest.Matchers.containsString; +import static org.hamcrest.Matchers.is; + +import java.io.ByteArrayInputStream; +import java.io.ByteArrayOutputStream; +import java.nio.charset.StandardCharsets; +import java.util.List; +import org.jline.terminal.Size; +import org.jline.terminal.Terminal; +import org.jline.terminal.TerminalBuilder; +import org.junit.jupiter.api.Test; + +/** + * What the screen would look like, asserted on the bytes the terminal emits. + * + *

A JLine terminal built over two streams renders exactly like one on a TTY — same escape + * sequences, same line reader — so a drawing bug is visible in the output without a console. That is + * worth having, because this class is the one place where a mistake is invisible to every other test + * and obvious to whoever is using the agent. + * + *

The bug these tests were written for: the rule above the input used to be the first line of a + * two-line prompt, while {@code ERASE_LINE_ON_FINISH} erases exactly one line — so every Enter + * left a rule behind, and holding Enter drew a column of them. + */ +class JLineTerminalTest { + + /** As many columns as a narrow window, so a wrapped line would be obvious. */ + private static final Size SIZE = new Size(60, 10); + + private final ByteArrayOutputStream emitted = new ByteArrayOutputStream(); + + private Terminal terminal(String keystrokes) throws Exception { + return TerminalBuilder.builder() + .streams(new ByteArrayInputStream(keystrokes.getBytes(StandardCharsets.UTF_8)), emitted) + .type("xterm-256color") + // The rule is drawn with U+2500. Both are needed: the writer encodes through the + // stdout charset, which is not the one .encoding() sets, and without it every rule + // arrives as a row of "?" and the assertions compare against bytes nobody wrote. + .encoding(StandardCharsets.UTF_8) + .stdoutEncoding(StandardCharsets.UTF_8) + .size(SIZE) + .provider("exec") + .build(); + } + + private String screen() { + return emitted.toString(StandardCharsets.UTF_8); + } + + /** + * How many times a full-width rule was written, in either of the two forms JLine uses. + * + *

Counting only one of them is how the first version of this test passed against the very bug + * it was written for: a box character goes out as UTF-8 {@code U+2500} from {@code printAbove}, + * but inside a prompt JLine switches to the DEC line-drawing character set and sends + * {@code ESC(0} + a row of {@code q} + {@code ESC(B}. A rule carried in the prompt is therefore + * invisible to a search for {@code ─}. + * + * @return how many rules were written + */ + private int rules() { + int width = SIZE.getColumns() - 1; + return occurrences("─".repeat(width)) + occurrences("(0" + "q".repeat(width)); + } + + private int occurrences(String needle) { + int count = 0; + for (int at = screen().indexOf(needle); at >= 0; at = screen().indexOf(needle, at + 1)) { + count++; + } + return count; + } + + @Test + void pressingEnterOnAnEmptyLineSeveralTimesDrawsTheRuleOnlyOnce() throws Exception { + try (Terminal terminal = terminal("\n\n\n\n"); + JLineTerminal console = JLineTerminal.over(terminal, List.of())) { + for (int i = 0; i < 4; i++) { + console.separator(); + assertThat("an empty line is still a line", console.readLine("ignored"), is("")); + } + + assertThat("nothing was printed in between, so nothing needs separating again", rules(), is(1)); + } + } + + @Test + void whatWasTypedSurvivesAboveTheInputAndTheNextTurnIsSeparatedAgain() throws Exception { + try (Terminal terminal = terminal("hello\nworld\n"); + JLineTerminal console = JLineTerminal.over(terminal, List.of())) { + console.separator(); + assertThat(console.readLine("ignored"), is("hello")); + console.separator(); + assertThat(console.readLine("ignored"), is("world")); + + // the input line itself is erased on Enter, so the echo is what keeps the transcript + assertThat(screen(), containsString("hello")); + assertThat(screen(), containsString("world")); + assertThat("the echo printed something, so the next read is separated again", rules(), is(2)); + } + } + + @Test + void outputPrintedWhileTheAgentWorksSeparatesTheNextInputAgain() throws Exception { + try (Terminal terminal = terminal("\n"); + JLineTerminal console = JLineTerminal.over(terminal, List.of())) { + console.separator(); + console.line("some output from a tool"); + console.separator(); + + assertThat("the rule is no longer next to the input, so it is drawn again", rules(), is(2)); + } + } + + @Test + void aMultiLineStringIsStillPrintedAsSeveralLines() throws Exception { + try (Terminal terminal = terminal("\n"); + JLineTerminal console = JLineTerminal.over(terminal, List.of())) { + console.line("first" + System.lineSeparator() + "second"); + + assertThat(screen(), containsString("first")); + assertThat(screen(), containsString("second")); + } + } +} From beddb163301d79be3e48a6636be3c0b7be86db3b Mon Sep 17 00:00:00 2001 From: Bernard Ladenthin Date: Wed, 23 Sep 2026 19:47:57 +0200 Subject: [PATCH 22/36] llama-atmosphere-agent: stop three writers from fighting over the terminal The bar filled with escape fragments and then disappeared. Three separate writers, only one of which I had accounted for. llama.cpp logs to stderr, which with an in-process model is the same console, and those writes go around JLine entirely -- the terminal scrolls lines it never sees and its reserved region ends up somewhere else than it believes. That is what "> d1.19.029.542 W srv stop: cancel task" is, and after a few of them the block is gone. captureNativeLog routes the log through LlamaModel.setLogger -- the callback sink patches/0014 added, for exactly this -- into terminal.line. The turn thread (Atmosphere's) prints tool lines while the console thread refreshes the block four times a second; neither printAbove nor Status.update knows about the other. A "writing" lock now spans line() and status(). That is the "1H" that was drawn into the middle of the rule. The rule above the input is gone. Carrying it in the prompt leaves one behind on every Enter; drawing it as output leaves one per turn in the scrollback, which is "the line keeps travelling along". JLine's pinned region is below the prompt and never above it, so those were the only two options and both were wrong. One rule, below the input, has neither problem. Holding Enter also cancelled one turn per keystroke: hasPendingInput() counted blank lines. It ignores them now but leaves them queued, because an empty answer to an approval question means yes. Honest limit: the lock is reasoned, not test-covered. Two attempts to pin it both passed with the lock removed -- including one that sliced every OutputStream write in half -- because PrintWriter already makes a single call atomic and the interleaving happens between calls, inside JLine. Deleted rather than kept: a test that is green either way is worse than none. 125 tests green. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01E2h8gXyyE5UeimkL9vQv9G --- CLAUDE.md | 43 +++++++---- llama-atmosphere-agent/README.md | 40 +++++------ .../llama/atmosphere/AgentTerminal.java | 15 ---- .../llama/atmosphere/JLineTerminal.java | 72 ++++++++++--------- .../llama/atmosphere/LocalAgent.java | 33 ++++++++- .../llama/atmosphere/JLineTerminalTest.java | 50 +++++++++---- 6 files changed, 157 insertions(+), 96 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 4fe77b4ea..2f672622d 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -2376,14 +2376,16 @@ are decisions, not details: 11. **The input is framed into the pinned block, and the prompt stays there during a turn; typing stops the turn.** The frame is two halves that must be read together: the **top** rule is the first line of the reader's *prompt* (`rule() + newline + "> "`, rebuilt on every read because the window - can be resized) — **no.** That is how it was built and it was wrong: `ERASE_LINE_ON_FINISH` erases - exactly **one** line, so a two-line prompt leaves its rule behind on every Enter, and holding Enter - draws a column of them. The **top** rule is therefore ordinary output (`AgentTerminal.separator()`, - printed before each read, suppressed when nothing has been printed since so repeated Enters draw - nothing), the **bottom** rule is the first line of the *status* block — JLine draws the status below - the prompt, so that is the only way to get the input inside a frame at all. Both call the same - `rule()`, or the two edges drift apart on a resize. The prompt is `"> "`, one line, and - `AgentTerminal.readLine`'s `prompt` argument is consequently **ignored** here. + can be resized) — **no, and the second attempt was wrong too.** Both are recorded because the + obvious fix is the one that fails. (1) Rule as the first line of a **two-line prompt**: + `ERASE_LINE_ON_FINISH` erases exactly **one** line, so every Enter leaves the rule behind and + holding Enter draws a column of them. (2) Rule as **ordinary output before each read**: nothing is + left behind on Enter any more, but one rule now stays in the scrollback per turn and travels up + with it. **There is no third option** — JLine's status region is below the prompt and never above + it, so a rule above the input can only be part of the prompt (1) or part of the scrollback (2). + The settled shape is therefore **one** rule, the first line of the status block, directly under the + input line; the prompt is `"> "`, one line, and `AgentTerminal.readLine`'s `prompt` argument is + consequently **ignored** here. `ERASE_LINE_ON_FINISH` removes the input line on Enter and the reader thread echoes it above as `› text`, so the transcript keeps what was asked. @@ -2395,10 +2397,27 @@ are decisions, not details: inside a *prompt*, so a counter that looks only for `─` passes against the exact bug it was written for — verified by putting the two-line prompt back and watching the test go from 1 rule to 5. - **The `[?1h` that appeared as text above the prompt** was a second, separate defect: `line()` wrote - straight to the terminal whenever the reader was not inside `readLine`, and that instant is exactly - when the next read is emitting its init sequence, so the two interleaved and half an escape sequence - landed in the scrollback. Once the reader thread exists, **everything** goes through `printAbove`. + **Escape sequences drawn as text** (`[?1h` above the prompt, then a `1H` inside the rule) were two + further defects of the same family, and the second is the one that eventually **destroyed the + block**. First: `line()` wrote straight to the terminal whenever the reader was not inside + `readLine`, which is exactly when the next read emits its init sequence — once the reader thread + exists, **everything** now goes through `printAbove`. Second: three threads write to this terminal + as a matter of course — the turn (Atmosphere's thread) prints tool lines, the console thread + refreshes the block four times a second, and with an in-process model **llama.cpp logs to stderr**, + which is the same console and goes around JLine entirely. The first two are serialised by a + `writing` lock held across `line()` and `status()`; the third is fixed by + `LocalAgent.captureNativeLog`, which routes the native log through `LlamaModel.setLogger` + (the callback sink `patches/0014` added) into `terminal.line`, so it scrolls in above the prompt + like any other output instead of scrolling lines JLine never sees. **Honest limit:** the lock is + reasoned, not test-covered. Two attempts to pin it are recorded in the history of + `JLineTerminalTest` and both passed with the lock removed — even one that sliced every + `OutputStream.write` in half — because `PrintWriter` already makes a single call atomic and the + interleaving happens *between* calls, inside JLine. A test that is green either way is worse than + none, so it was deleted rather than kept. + + **A blank line must not count as pending input.** `hasPendingInput()` ignores blank lines but leaves + them queued: counting them meant that holding Enter cancelled one turn per keystroke and produced + nothing, while dropping them would break the approval prompt, where an empty answer means yes. `DISABLE_EVENT_EXPANSION` is set in the same builder because the reader's default treats `!` as a shell history expansion, which silently rewrites a request like `git commit -m "fixed!"`. **What cannot be done, asked and answered:** keep the block visible while the *user* scrolls the diff --git a/llama-atmosphere-agent/README.md b/llama-atmosphere-agent/README.md index dd27edc5a..7d994f1e1 100644 --- a/llama-atmosphere-agent/README.md +++ b/llama-atmosphere-agent/README.md @@ -294,35 +294,35 @@ output **line by line while it runs** (dimmed, `│ `-prefixed) instead of dumpi also keeps the pipe drained; a process whose output nobody reads blocks once the buffer is full, and on Windows that buffer is about 4 KB. -**The bottom of the window is one framed block**: the input line between two rules, then what the agent -is doing and the session state. +**The bottom of the window is a pinned block**: the input line, then a rule, then what the agent is +doing and the session state. ``` -──────────────────────────────────────────────────────────────────── -> add a test for the parser +› add a test for the parser +… the answer … +> ──────────────────────────────────────────────────────────────────── ⠙ Fettling… (run_command 47s of 61s · 2 tool calls) [/path/to/project · ⏸ manual · ctx ~3.1k/16k · 9 tools · local-model] ``` -The top rule is ordinary output, printed once before each read; the three rows below are the pinned -block. There is no `you>`: the box already says where the input is. On Enter the input line is erased -and echoed above as `› your text`, so the transcript keeps what was asked. (`!` history expansion is +There is no `you>`: the block already says where the input is. On Enter the input line is erased and +echoed above as `› your text`, so the transcript keeps what was asked. (`!` history expansion is switched off in the same place, or a request like `git commit -m "fixed!"` would be rewritten silently.) -**Why the top rule is not part of the prompt**, although that is the obvious way to draw it: the line -reader erases exactly **one** line when the input is submitted, so a two-line prompt leaves its rule -behind on every Enter — hold Enter and you get a column of them. It was built that way first and the -symptom was reported within the hour. `JLineTerminalTest` now drives a real line reader over a pair of -streams and counts the rules in the bytes it emits, which is the only way to see this without a -console. Note what that test had to learn: a box character goes out as UTF-8 `─` from ordinary output -but as the DEC line-drawing set (`ESC(0` + a row of `q`) inside a prompt, so counting only the first -form passed against the very bug it was written for. - -The two lower rows are always both present: the activity row — `… waiting for input …` when it is your -turn, the spinner with the running tool and its elapsed time while it works — and the session's state. -Two fixed rows rather than one changing one: a block that changes height makes the output above it jump -on every refresh. +**Why there is no second rule above the input**, although that is what this looked like at first: JLine's +pinned region sits below the prompt and never above it, so a rule above the input can only be part of the +prompt or part of the scrollback — and both were tried and both were wrong. In the prompt it survives +every Enter, because the reader erases exactly one line (hold Enter, get a column of rules). As output it +leaves one rule per turn behind, travelling up the scrollback. One rule, below the input, is the shape +that has neither problem. + +**Three threads write to this console** and all of them had to be brought into line, because a write that +goes around the line reader scrolls the screen without JLine noticing and the pinned block ends up +somewhere else than it believes — first as a stray `1H` drawn into the rule, then as no block at all. +The turn thread and the console thread share a lock; llama.cpp's own log, which with `--model` goes to +stderr and therefore straight past everything, is routed through the console with `LlamaModel.setLogger`. + **On scrolling.** The block stays put while the agent writes: JLine keeps those lines out of the terminal's scroll region. It cannot stay while *you* scroll the terminal's own scrollback with the diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentTerminal.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentTerminal.java index 89bb53f23..03a914150 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentTerminal.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentTerminal.java @@ -71,21 +71,6 @@ public interface AgentTerminal extends AutoCloseable { */ boolean pinsStatus(); - /** - * Draw the rule that separates the conversation from the input line, unless it is already there. - * - *

It is ordinary output rather than part of the prompt, and that is the whole point: a prompt - * of two lines is erased as one when the line is submitted, so a rule carried in the prompt - * survives every Enter and stacks up. Written once and never touched again, like everything else - * on this console. - * - *

Callers ask for it before every read; asking again with nothing printed in between draws - * nothing, so holding Enter does not produce a column of rules. - */ - default void separator() { - // nothing to frame on a plain stream - } - /** * Whether the user has already typed a line that nobody has read yet. * diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java index ea5ad9cdf..c60687059 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java @@ -58,7 +58,18 @@ public final class JLineTerminal implements AgentTerminal { private final Status status; private final Ansi ansi; private final BlockingQueue typed = new LinkedBlockingQueue<>(); - private volatile boolean ruleOnScreen; + /** + * Held for the length of every write this class makes. + * + *

Two threads write here as a matter of course: the turn runs on a thread of Atmosphere's and + * prints its tool lines and streamed text, while the console thread refreshes the pinned block + * four times a second. Neither JLine's {@code printAbove} nor {@code Status.update} knows about + * the other, so without this their escape sequences interleave and a fragment lands on screen as + * text — a stray {@code 1H}, the tail of a cursor-position sequence, drawn into the middle of the + * rule. + */ + private final Object writing = new Object(); + private volatile boolean closed; private @Nullable Thread input; @@ -118,7 +129,6 @@ static JLineTerminal over(Terminal terminal, List completions) { @Override public void line(String text) { - ruleOnScreen = false; if (text.indexOf('\n') >= 0 || text.indexOf('\r') >= 0) { // One call must be one screen line: the status block is sized in lines, so a multi-line // string handed over as "a line" desynchronises the reserved region. Callers fold their @@ -126,30 +136,18 @@ public void line(String text) { text.lines().forEach(this::line); return; } - if (input != null) { - // Once the reader thread exists it owns the screen, and nothing may write around it -- - // not even in the instant between two reads. A direct write there cuts into the escape - // sequence the next read is emitting and half of it lands in the scrollback as text: - // a stray "[?1h" above the prompt was exactly that. - reader.printAbove(text); - } else { - terminal.writer().println(text); - terminal.writer().flush(); - } - } - - @Override - public void separator() { - if (ruleOnScreen) { - // Already drawn, with nothing printed since: it is still directly above the input line. - return; + synchronized (writing) { + if (input != null) { + // Once the reader thread exists it owns the screen, and nothing may write around it -- + // not even in the instant between two reads. A direct write there cuts into the escape + // sequence the next read is emitting and half of it lands in the scrollback as text: + // a stray "[?1h" above the prompt was exactly that. + reader.printAbove(text); + } else { + terminal.writer().println(text); + terminal.writer().flush(); + } } - // Starts the reader if this is the first call: the rule belongs above a prompt, so drawing it - // is also the moment the prompt has to exist. Without this the very first box had no top edge - // -- the reader only started on the first read, which comes after. - startReading(); - reader.printAbove(rule()); - ruleOnScreen = true; } /** @@ -175,10 +173,8 @@ private synchronized void startReading() { try { // One line, and it has to stay one line: the reader erases a single line // when the input is submitted, so a two-line prompt -- the rule above the - // input, which is where this started -- leaves the rule behind on every - // Enter, a column of them after a few. separator() draws the rule as - // ordinary output instead, which cannot be left behind because it was - // never part of the prompt. + // input, which is what was tried first -- leaves that rule behind on + // every Enter, a column of them after a few. String line = reader.readLine("> "); if (!line.isBlank()) { // The input line is erased on Enter, so the conversation would lose @@ -244,7 +240,10 @@ private String rule() { @Override public boolean hasPendingInput() { - return !typed.isEmpty(); + // A blank line is not a request, so it must not count: the REPL skips it, and counting it + // meant that holding Enter cancelled one turn per keystroke and produced nothing. It stays in + // the queue, because an empty answer to an approval question means yes. + return typed.stream().anyMatch(line -> !line.isBlank()); } @Override @@ -267,13 +266,20 @@ public boolean hasPendingInput() { @Override public void status(List lines) { + synchronized (writing) { + updateStatus(lines); + } + } + + private void updateStatus(List lines) { if (lines.isEmpty()) { status.update(List.of()); return; } - // The rule on top of this block is the bottom edge of the input box: the reader draws the top - // edge as the first line of its prompt, so the two together frame the input the way the - // established terminal agents do. + // The rule on top of this block is the one that separates the conversation from the input. + // It is the ONLY one drawn: a rule above the input line cannot be pinned (JLine's status + // region is below the prompt, never above it) and drawing it as output leaves one behind in + // the scrollback per turn, which is what "the line keeps travelling along" was. // One row must never wrap: a wrapped row occupies two screen lines, the reserved region is // sized in lines, and everything below it is then drawn in the wrong place -- which is how a // long summary tore the block apart. diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java index 13558a8ce..68b6c0118 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java @@ -17,6 +17,7 @@ import java.util.Optional; import java.util.concurrent.atomic.AtomicReference; import net.ladenthin.llama.LlamaModel; +import net.ladenthin.llama.args.LogFormat; import net.ladenthin.llama.parameters.ModelParameters; import net.ladenthin.llama.server.OpenAiCompatServer; import net.ladenthin.llama.server.OpenAiServerConfig; @@ -199,6 +200,7 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream if (terminal == null) { terminal = new PlainTerminal(out, reader, Ansi.detect()); } + captureNativeLog(terminal, options); // One-shot runs have nobody at the keyboard, so the strategy gets no console and denies // gated calls unless --auto was passed (see ConsoleApprovalStrategy). TurnActivity activity = new TurnActivity(); @@ -271,7 +273,6 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream } else { terminal.line(terminal.ansi().dim(status)); } - terminal.separator(); String line = terminal.readLine("you> "); if (line == null) { return 0; @@ -408,6 +409,36 @@ private static void loop( : terminal.ansi().yellow("loop: " + outcome.reason())); } + /** + * Route llama.cpp's own log through the console instead of letting it write to stderr. + * + *

**This is what destroys a pinned block, and nothing on the Java side can defend against it.** + * With an in-process model the server logs to stderr, which is the same console; those writes go + * around the line reader, so the terminal scrolls lines JLine never sees and its reserved region + * ends up somewhere else than it believes. What that looks like: a warning printed into the middle + * of the input line (`> d1.19.029.542 W srv stop: cancel task`), and after a few of them the + * block is gone. Routing the log through {@link AgentTerminal#line} puts it under the same lock as + * everything else, so it scrolls in above the prompt like any other output. + * + *

Only for an in-process model: with {@code --base-url} the server is another process and its + * log is its own business, and calling this would load the native library for nothing. + * + * @param terminal where the log lines go + * @param options the parsed command line + */ + private static void captureNativeLog(AgentTerminal terminal, AgentOptions options) { + if (options.getModelPath() == null) { + return; + } + Ansi ansi = terminal.ansi(); + LlamaModel.setLogger(LogFormat.TEXT, (level, message) -> { + String text = message == null ? "" : message.strip(); + if (!text.isEmpty()) { + terminal.line(ansi.dim(text)); + } + }); + } + static ConsoleSession turn( AgentRunner runner, AgentFileSystem fileSystem, diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/JLineTerminalTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/JLineTerminalTest.java index 11d1aaaef..b6ccb1f74 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/JLineTerminalTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/JLineTerminalTest.java @@ -7,6 +7,7 @@ import static org.hamcrest.MatcherAssert.assertThat; import static org.hamcrest.Matchers.containsString; import static org.hamcrest.Matchers.is; +import static org.hamcrest.Matchers.nullValue; import java.io.ByteArrayInputStream; import java.io.ByteArrayOutputStream; @@ -27,7 +28,10 @@ * *

The bug these tests were written for: the rule above the input used to be the first line of a * two-line prompt, while {@code ERASE_LINE_ON_FINISH} erases exactly one line — so every Enter - * left a rule behind, and holding Enter drew a column of them. + * left a rule behind, and holding Enter drew a column of them. Drawing it as ordinary output instead + * only moved the problem: then one rule stayed in the scrollback per turn and travelled up with it. + * There is now exactly one rule, the first line of the pinned block, and these tests hold the console + * to writing none at all. */ class JLineTerminalTest { @@ -67,7 +71,7 @@ private String screen() { */ private int rules() { int width = SIZE.getColumns() - 1; - return occurrences("─".repeat(width)) + occurrences("(0" + "q".repeat(width)); + return occurrences("─".repeat(width)) + occurrences("\u001b(0" + "q".repeat(width)); } private int occurrences(String needle) { @@ -79,43 +83,59 @@ private int occurrences(String needle) { } @Test - void pressingEnterOnAnEmptyLineSeveralTimesDrawsTheRuleOnlyOnce() throws Exception { + void pressingEnterSeveralTimesLeavesNoRuleInTheScrollback() throws Exception { try (Terminal terminal = terminal("\n\n\n\n"); JLineTerminal console = JLineTerminal.over(terminal, List.of())) { for (int i = 0; i < 4; i++) { - console.separator(); assertThat("an empty line is still a line", console.readLine("ignored"), is("")); } - assertThat("nothing was printed in between, so nothing needs separating again", rules(), is(1)); + assertThat("the only rule is the pinned one, and it is never written as output", rules(), is(0)); } } @Test - void whatWasTypedSurvivesAboveTheInputAndTheNextTurnIsSeparatedAgain() throws Exception { + void whatWasTypedSurvivesAboveTheInputLine() throws Exception { try (Terminal terminal = terminal("hello\nworld\n"); JLineTerminal console = JLineTerminal.over(terminal, List.of())) { - console.separator(); assertThat(console.readLine("ignored"), is("hello")); - console.separator(); assertThat(console.readLine("ignored"), is("world")); // the input line itself is erased on Enter, so the echo is what keeps the transcript assertThat(screen(), containsString("hello")); assertThat(screen(), containsString("world")); - assertThat("the echo printed something, so the next read is separated again", rules(), is(2)); + assertThat("and still no rule travels along with it", rules(), is(0)); } } @Test - void outputPrintedWhileTheAgentWorksSeparatesTheNextInputAgain() throws Exception { - try (Terminal terminal = terminal("\n"); + void anEmptyLineIsNotAPendingRequest() throws Exception { + try (Terminal terminal = terminal("\n \nreal\n"); JLineTerminal console = JLineTerminal.over(terminal, List.of())) { - console.separator(); - console.line("some output from a tool"); - console.separator(); + console.readLine("ignored"); // starts the reader; the rest queues up behind it + + waitFor(console::hasPendingInput); + // Two blank lines are queued in front of it, and it is still the real one that counts: + // otherwise holding Enter cancels one turn per keystroke and produces nothing. + assertThat(console.hasPendingInput(), is(true)); + assertThat(console.readLine("ignored").isBlank(), is(true)); + assertThat(console.readLine("ignored"), is("real")); + // End of input is pending too, and has to be: a session whose console has closed must + // stop waiting rather than keep a turn running for nobody. + assertThat(console.hasPendingInput(), is(true)); + assertThat("and it reads as no line at all", console.readLine("ignored"), is(nullValue())); + } + } - assertThat("the rule is no longer next to the input, so it is drawn again", rules(), is(2)); + /** + * Wait for the reader thread to have caught up. + * + * @param condition what to wait for + * @throws InterruptedException if interrupted while waiting + */ + private void waitFor(java.util.function.BooleanSupplier condition) throws InterruptedException { + for (int attempt = 0; attempt < 200 && !condition.getAsBoolean(); attempt++) { + Thread.sleep(10); } } From 1ad2965fa8841a40de80bdc7baeb281468a838c0 Mon Sep 17 00:00:00 2001 From: Bernard Ladenthin Date: Wed, 23 Sep 2026 19:54:11 +0200 Subject: [PATCH 23/36] llama-atmosphere-agent: scroll to the bottom once, so the input starts there The input looked like an ordinary prompt again after the last change, and only settled into the block after a few turns. That is geometry, not chance: the line reader draws its prompt where the cursor is -- directly after the last thing printed -- while only the status block is pinned to the window. On a half-empty screen the two are far apart, and they meet once enough output has scrolled the cursor down by itself. Pushing the cursor to the last row before the first prompt makes that the state from the beginning. The cost is a screenful of blank lines above the session, which is what a program that wants its input at the bottom without taking over the whole screen has to pay; the alternative is the alternate screen buffer, and that costs the scrollback. 126 tests green. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01E2h8gXyyE5UeimkL9vQv9G --- CLAUDE.md | 10 +++++++ llama-atmosphere-agent/README.md | 8 ++++++ .../llama/atmosphere/JLineTerminal.java | 28 +++++++++++++++++++ .../llama/atmosphere/JLineTerminalTest.java | 18 ++++++++++++ 4 files changed, 64 insertions(+) diff --git a/CLAUDE.md b/CLAUDE.md index 2f672622d..b237b07dd 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -2415,6 +2415,16 @@ are decisions, not details: interleaving happens *between* calls, inside JLine. A test that is green either way is worse than none, so it was deleted rather than kept. + **The screen is scrolled to the bottom once, before the first prompt** (`scrollToBottom`). The + reader draws its prompt at the cursor, i.e. after the last line printed, while only the status + block is pinned to the window — so on a half-empty screen the input floats in the middle with the + block far below it, and they only meet once output has scrolled the cursor down by itself. That is + why it looked right after a few turns and like an ordinary prompt at the start. Emitting + `rows - 1` newlines once makes it the state from the first prompt on; from then on every printed + line scrolls and the cursor stays on the last row. The cost is a screenful of blank lines above the + session, which is what a program that wants its input at the bottom *without* taking over the + screen has to pay. + **A blank line must not count as pending input.** `hasPendingInput()` ignores blank lines but leaves them queued: counting them meant that holding Enter cancelled one turn per keystroke and produced nothing, while dropping them would break the approval prompt, where an empty answer means yes. diff --git a/llama-atmosphere-agent/README.md b/llama-atmosphere-agent/README.md index 7d994f1e1..3175d7c50 100644 --- a/llama-atmosphere-agent/README.md +++ b/llama-atmosphere-agent/README.md @@ -317,6 +317,14 @@ every Enter, because the reader erases exactly one line (hold Enter, get a colum leaves one rule per turn behind, travelling up the scrollback. One rule, below the input, is the shape that has neither problem. +**Why the screen is scrolled once at startup.** The line reader draws its prompt where the cursor is, +which is directly after the last thing printed; only the block below it is pinned to the window. On a +half-empty screen that leaves the input floating in the middle with the block far below, and the two +only meet once enough output has scrolled the cursor down by itself — which is why it looks right after +a few turns and wrong at the start. Pushing the cursor to the last row before the first prompt makes +that the state from the beginning. The cost is a screen of blank lines above the session: the +alternative is taking over the whole screen (alternate buffer), which costs the scrollback. + **Three threads write to this console** and all of them had to be brought into line, because a write that goes around the line reader scrolls the screen without JLine noticing and the pinned block ends up somewhere else than it believes — first as a stray `1H` drawn into the rule, then as no block at all. diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java index c60687059..f9733ef59 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java @@ -167,6 +167,7 @@ private synchronized void startReading() { if (input != null) { return; } + scrollToBottom(); input = new Thread( () -> { while (!closed) { @@ -201,6 +202,33 @@ private synchronized void startReading() { input.start(); } + /** + * Push the cursor to the last usable row, once, before the first prompt is drawn. + * + *

The line reader draws its prompt wherever the cursor happens to be, which is directly after + * the last thing printed; only the status block is pinned to the bottom of the window. On a + * half-empty screen that leaves the input floating in the middle with the block far below it, and + * it only looks like one piece once enough output has scrolled the cursor down by itself — which + * is why it looked right after a few turns and wrong at the start. + * + *

Scrolling the screen once at startup makes that the state from the first prompt on: from + * then on every line printed scrolls, so the cursor stays on the last row for the rest of the + * session. The cost is a screen of blank lines above the session, which is what any program that + * wants its input at the bottom without taking over the whole screen has to pay. + */ + private void scrollToBottom() { + int rows = terminal.getSize().getRows(); + if (rows <= 1) { + return; // no size to speak of (a pipe, a terminal that will not say): nothing to scroll + } + synchronized (writing) { + for (int row = 0; row < rows - 1; row++) { + terminal.writer().println(); + } + terminal.writer().flush(); + } + } + /** * The rule that frames the input box, as wide as the window. * diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/JLineTerminalTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/JLineTerminalTest.java index b6ccb1f74..a9b05710a 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/JLineTerminalTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/JLineTerminalTest.java @@ -139,6 +139,24 @@ private void waitFor(java.util.function.BooleanSupplier condition) throws Interr } } + @Test + void theScreenIsScrolledSoTheInputStartsAtTheBottom() throws Exception { + // The reader draws its prompt where the cursor is, and only the block below is pinned to the + // window. Without this the input floats after the output with the block far below it, and the + // two only meet once enough output has scrolled the cursor down on its own. + try (Terminal terminal = terminal("\n"); + JLineTerminal console = JLineTerminal.over(terminal, List.of())) { + console.readLine("ignored"); + + long blankLines = + screen().chars().filter(character -> character == '\n').count(); + assertThat( + "the cursor is pushed to the last row before the first prompt", + blankLines >= SIZE.getRows() - 1, + is(true)); + } + } + @Test void aMultiLineStringIsStillPrintedAsSeveralLines() throws Exception { try (Terminal terminal = terminal("\n"); From 10deecddf05eca7d49be3b5842efce161282cce1 Mon Sep 17 00:00:00 2001 From: Bernard Ladenthin Date: Wed, 23 Sep 2026 20:49:31 +0200 Subject: [PATCH 24/36] llama-atmosphere-agent: /cls wipes the screen, /clear wipes it and the history AgentTerminal.clearScreen() defaults to nothing (a stream has no screen); JLineTerminal sends the terminal's own clear_screen capability through printAbove, like every other write, then refills the blank rows and redraws the block. /clear wipes the screen too: what is still on it afterwards is a conversation the model no longer has, which reads as if it were still in play. Neither touches the emulator's scrollback. Two things the tests found rather than the reading. getStringCapability returns terminfo SOURCE -- "\E[H\E[2J", with the escape spelled out -- so writing it as it comes prints that text on the screen; Curses.tputs expands it. And Ctrl-L already did all of this before the command existed, bound by JLine's own keymap, so /cls is the second way to do it and not the only one. Both are pinned. 128 tests green. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01E2h8gXyyE5UeimkL9vQv9G --- CLAUDE.md | 9 ++++ llama-atmosphere-agent/README.md | 10 ++++- .../llama/atmosphere/AgentTerminal.java | 12 ++++++ .../llama/atmosphere/JLineTerminal.java | 43 ++++++++++++++++--- .../llama/atmosphere/LocalAgent.java | 4 ++ .../llama/atmosphere/SlashCommands.java | 4 +- .../net/ladenthin/llama/atmosphere/help.txt | 3 +- .../llama/atmosphere/JLineTerminalTest.java | 34 +++++++++++++++ 8 files changed, 111 insertions(+), 8 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index b237b07dd..72199760a 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -2415,6 +2415,15 @@ are decisions, not details: interleaving happens *between* calls, inside JLine. A test that is green either way is worse than none, so it was deleted rather than kept. + **`/cls` wipes the screen, `/clear` wipes it and the history.** `AgentTerminal.clearScreen()` + defaults to doing nothing (a stream has no screen); `JLineTerminal` expands the terminal's + `clear_screen` capability and sends it **through `printAbove`**, like every other write, then + refills the blank rows and redraws the block. One trap, caught by the test rather than by reading: + `getStringCapability` returns **terminfo source** (`\E[H\E[2J`, with the escape spelled out), so + writing it as it comes prints that text on the screen — `Curses.tputs` expands it. Ctrl-L already + did this before the command existed, bound by JLine's own keymap; a test pins that too, so a keymap + option cannot quietly remove it. + **The screen is scrolled to the bottom once, before the first prompt** (`scrollToBottom`). The reader draws its prompt at the cursor, i.e. after the last line printed, while only the status block is pinned to the window — so on a half-empty screen the input floats in the middle with the diff --git a/llama-atmosphere-agent/README.md b/llama-atmosphere-agent/README.md index 3175d7c50..dd6fd992a 100644 --- a/llama-atmosphere-agent/README.md +++ b/llama-atmosphere-agent/README.md @@ -189,7 +189,8 @@ unknown `/command` included — goes to the model: | `/mode [manual\|auto]` (`/approve`) | show or set the approval mode (`⏸ manual` / `⏵⏵ auto`) | | `/compact [focus]` | summarize the conversation and continue from the summary | | `/loop [--every 5m] [--max 20] [--check ''] ` | keep working on one task until it is done | -| `/clear` (`/reset`, `/new`) | drop the history | +| `/clear` (`/reset`, `/new`) | drop the history, and wipe the screen with it | +| `/cls` (`/clear-screen`) | wipe the screen, keep the conversation — Ctrl-L does the same | | `/exit` (`/quit`) | leave | **The prompt.** On a real terminal the agent uses [JLine](https://github.com/jline/jline3): arrow keys @@ -317,6 +318,13 @@ every Enter, because the reader erases exactly one line (hold Enter, get a colum leaves one rule per turn behind, travelling up the scrollback. One rule, below the input, is the shape that has neither problem. +**Wiping the screen.** `/cls` clears the window and leaves the input and the block at the bottom — +Ctrl-L does the same, bound by the line reader itself rather than by this project (a test pins that, so +a keymap change cannot quietly take it away). `/clear` wipes the screen *and* drops the history: what is +still on screen after a `/clear` is a conversation the model no longer has, which reads as if it were +still in play. Neither touches the terminal emulator's own scrollback — what was written stays where +the scrollbar can reach it. + **Why the screen is scrolled once at startup.** The line reader draws its prompt where the cursor is, which is directly after the last thing printed; only the block below it is pinned to the window. On a half-empty screen that leaves the input floating in the middle with the block far below, and the two diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentTerminal.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentTerminal.java index 03a914150..6358a420e 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentTerminal.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentTerminal.java @@ -71,6 +71,18 @@ public interface AgentTerminal extends AutoCloseable { */ boolean pinsStatus(); + /** + * Wipe the window, leaving the input line and the block at the bottom. + * + *

The scrollback of the terminal emulator is not touched — what was written stays where the + * scrollbar can reach it. This only clears what is on screen, the way {@code clear} or Ctrl-L does. + * + *

A plain stream has no screen to clear, so the default does nothing. + */ + default void clearScreen() { + // nothing to wipe on a stream + } + /** * Whether the user has already typed a line that nobody has read yet. * diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java index f9733ef59..3560579dd 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java @@ -202,6 +202,43 @@ private synchronized void startReading() { input.start(); } + @Override + public void clearScreen() { + String capability = terminal.getStringCapability(InfoCmp.Capability.clear_screen); + if (capability == null) { + return; // a terminal that cannot clear: better nothing than a guessed escape sequence + } + // The capability is terminfo source, not the sequence itself: it reads "\E[H\E[2J", with the + // escape spelled out. Writing it as it comes prints that text on the screen, which is what a + // test caught. Curses expands it the way terminal.puts would, but into a string this class can + // hand to the reader instead of writing behind its back. + StringBuilder expanded = new StringBuilder(); + org.jline.utils.Curses.tputs(expanded, capability); + String clear = expanded.toString(); + synchronized (writing) { + if (input == null) { + terminal.writer().print(clear); + terminal.writer().flush(); + return; + } + // Through the reader, like every other write once it exists: printAbove leaves the prompt + // redrawn and the reader's idea of the cursor intact, which writing the escape sequence + // around it would not. The blank rows put the input back on the last row, where clearing + // to the top-left corner has just moved it away from. + reader.printAbove(clear + System.lineSeparator().repeat(blankRows())); + status.redraw(); + } + } + + /** + * How many rows to fill so the cursor ends up on the last usable one. + * + * @return the count, never negative + */ + private int blankRows() { + return Math.max(0, terminal.getSize().getRows() - 1); + } + /** * Push the cursor to the last usable row, once, before the first prompt is drawn. * @@ -217,12 +254,8 @@ private synchronized void startReading() { * wants its input at the bottom without taking over the whole screen has to pay. */ private void scrollToBottom() { - int rows = terminal.getSize().getRows(); - if (rows <= 1) { - return; // no size to speak of (a pipe, a terminal that will not say): nothing to scroll - } synchronized (writing) { - for (int row = 0; row < rows - 1; row++) { + for (int row = 0; row < blankRows(); row++) { terminal.writer().println(); } terminal.writer().flush(); diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java index 68b6c0118..81edd307b 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java @@ -559,8 +559,12 @@ private static long handleCommand( case HELP -> prompt(HELP_TEXT).lines().forEach(terminal::line); case CLEAR -> { history.clear(); + // The screen goes with it: what is still on it is a conversation the model no longer + // has, which reads as if it were still in play. + terminal.clearScreen(); terminal.line("(history cleared)"); } + case CLS -> terminal.clearScreen(); case CALLS -> callLog.render().lines().forEach(terminal::line); case TOOLS -> { terminal.line("tools: " + String.join(", ", runner.toolNames())); diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/SlashCommands.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/SlashCommands.java index a58504638..18f653705 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/SlashCommands.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/SlashCommands.java @@ -28,8 +28,10 @@ public record SlashCommands(Command command, String arguments) { public enum Command { /** Print the command overview. */ HELP("/help", "/?", "/commands"), - /** Drop the conversation history. */ + /** Drop the conversation history, and wipe the screen with it. */ CLEAR("/clear", "/reset", "/new"), + /** Wipe the screen, keeping the conversation. */ + CLS("/cls", "/clear-screen"), /** Summarize the history and continue with the summary; the argument steers the summary. */ COMPACT("/compact"), /** Keep working on one task until it is done; see {@link TaskLoop}. */ diff --git a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt index ab5fbc0e5..12c28c98d 100644 --- a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt +++ b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt @@ -12,7 +12,8 @@ Commands (everything else is sent to the model): /loop [--every 5m] [--max 20] [--check ''] keep working on until the model answers <>; the state lives in AGENT-LOOP.md, not in the conversation - /clear drop the history (/reset, /new) + /clear drop the history, and wipe the screen with it (/reset, /new) + /cls wipe the screen, keep the conversation (/clear-screen, or Ctrl-L) /exit leave (/quit) The input line sits in the box at the bottom and is there while the agent works: you can type at diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/JLineTerminalTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/JLineTerminalTest.java index a9b05710a..6243b2a85 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/JLineTerminalTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/JLineTerminalTest.java @@ -7,6 +7,7 @@ import static org.hamcrest.MatcherAssert.assertThat; import static org.hamcrest.Matchers.containsString; import static org.hamcrest.Matchers.is; +import static org.hamcrest.Matchers.notNullValue; import static org.hamcrest.Matchers.nullValue; import java.io.ByteArrayInputStream; @@ -38,6 +39,9 @@ class JLineTerminalTest { /** As many columns as a narrow window, so a wrapped line would be obvious. */ private static final Size SIZE = new Size(60, 10); + /** What a cleared screen looks like on the wire: erase the whole display. */ + private static final String ERASE_DISPLAY = "\u001b[2J"; + private final ByteArrayOutputStream emitted = new ByteArrayOutputStream(); private Terminal terminal(String keystrokes) throws Exception { @@ -157,6 +161,36 @@ void theScreenIsScrolledSoTheInputStartsAtTheBottom() throws Exception { } } + @Test + void clearingTheScreenWipesItAndLeavesTheReaderWorking() throws Exception { + try (Terminal terminal = terminal("first\nsecond\n"); + JLineTerminal console = JLineTerminal.over(terminal, List.of())) { + assertThat(console.readLine("ignored"), is("first")); + int before = screen().length(); + + console.clearScreen(); + + // The capability is terminfo source ("\E[H\E[2J"), so what must reach the screen is the + // expanded form. Writing the capability as it comes prints it as text, which is what this + // assertion caught the first time it ran. + assertThat( + terminal.getStringCapability(org.jline.utils.InfoCmp.Capability.clear_screen), is(notNullValue())); + assertThat("erase display reached the screen", screen().substring(before), containsString(ERASE_DISPLAY)); + assertThat("and the prompt still reads afterwards", console.readLine("ignored"), is("second")); + } + } + + @Test + void controlLIsBoundToTheReadersOwnClearScreen() throws Exception { + // 0x0C is Ctrl-L. It is bound by JLine itself, so /cls is the second way to do this rather + // than the only one -- worth pinning, because a keymap option could silently take it away. + try (Terminal terminal = terminal("\u000cstill here\n"); + JLineTerminal console = JLineTerminal.over(terminal, List.of())) { + assertThat(console.readLine("ignored"), is("still here")); + assertThat("Ctrl-L cleared rather than being typed into the line", screen(), containsString(ERASE_DISPLAY)); + } + } + @Test void aMultiLineStringIsStillPrintedAsSeveralLines() throws Exception { try (Terminal terminal = terminal("\n"); From 38d8e08771e3a08c4bedb5ad870d7d0eca0df5d4 Mon Sep 17 00:00:00 2001 From: Bernard Ladenthin Date: Wed, 23 Sep 2026 21:35:20 +0200 Subject: [PATCH 25/36] llama-atmosphere-agent: put the block back after a screen wipe After /cls the bar at the bottom was gone. Status.redraw() writes nothing in that situation: it draws what has changed, and a wipe changes nothing about the block's content -- it only takes it off the screen, which the object has no way of knowing. The rendered block is now kept in a field and restored with reset() (forget what is believed to be on screen) plus update(block). Found by writing the test first: it reproduced the empty bottom, and putting redraw() back makes it red again, so it is a regression test rather than a description. The probe also settled what actually happens on the wire -- the clear sequence was written correctly all along, only the block never came back. 129 tests green. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01E2h8gXyyE5UeimkL9vQv9G --- CLAUDE.md | 8 ++++++- llama-atmosphere-agent/README.md | 4 +++- .../llama/atmosphere/JLineTerminal.java | 24 +++++++++++++++---- .../llama/atmosphere/JLineTerminalTest.java | 21 ++++++++++++++++ 4 files changed, 50 insertions(+), 7 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 72199760a..6295d36f9 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -2420,7 +2420,13 @@ are decisions, not details: `clear_screen` capability and sends it **through `printAbove`**, like every other write, then refills the blank rows and redraws the block. One trap, caught by the test rather than by reading: `getStringCapability` returns **terminfo source** (`\E[H\E[2J`, with the escape spelled out), so - writing it as it comes prints that text on the screen — `Curses.tputs` expands it. Ctrl-L already + writing it as it comes prints that text on the screen — `Curses.tputs` expands it. A second one, and + the reason the bar went missing after a `/cls`: **`Status.redraw()` writes nothing after a wipe.** + It draws what has *changed*, and a wipe changes nothing about its content — it only removes it from + the screen, which the object has no way of knowing. So the block is kept in a field as it was last + rendered and put back with `status.reset()` (forget what is believed to be on screen) followed by + `status.update(block)`; `redraw()` alone is a no-op, verified by putting it back and watching the + test go red. Ctrl-L already did this before the command existed, bound by JLine's own keymap; a test pins that too, so a keymap option cannot quietly remove it. diff --git a/llama-atmosphere-agent/README.md b/llama-atmosphere-agent/README.md index dd6fd992a..30a425b67 100644 --- a/llama-atmosphere-agent/README.md +++ b/llama-atmosphere-agent/README.md @@ -323,7 +323,9 @@ Ctrl-L does the same, bound by the line reader itself rather than by this projec a keymap change cannot quietly take it away). `/clear` wipes the screen *and* drops the history: what is still on screen after a `/clear` is a conversation the model no longer has, which reads as if it were still in play. Neither touches the terminal emulator's own scrollback — what was written stays where -the scrollbar can reach it. +the scrollbar can reach it. The block at the bottom is redrawn from a kept copy afterwards: JLine draws +the pinned region only when its *content* changes, and a wipe does not change the content, it only takes +it off the screen — so asking it to redraw does nothing and the bottom of the window stays empty. **Why the screen is scrolled once at startup.** The line reader draws its prompt where the cursor is, which is directly after the last thing printed; only the block below it is pinned to the window. On a diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java index 3560579dd..7b5d9fc38 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java @@ -70,6 +70,15 @@ public final class JLineTerminal implements AgentTerminal { */ private final Object writing = new Object(); + /** + * The block as it was last handed over, so it can be put back after the screen is wiped. + * + *

{@code Status} draws only what has changed, and a wipe does not change its content — it just + * removes it from the screen. Without keeping a copy there is nothing to redraw it from, and the + * bottom of the window stays empty until the next refresh happens to differ. + */ + private volatile List block = List.of(); + private volatile boolean closed; private @Nullable Thread input; @@ -226,7 +235,10 @@ public void clearScreen() { // around it would not. The blank rows put the input back on the last row, where clearing // to the top-left corner has just moved it away from. reader.printAbove(clear + System.lineSeparator().repeat(blankRows())); - status.redraw(); + // reset() makes it forget what it believes is on screen; without that the update below is + // a no-op, because the content it would draw is the content it thinks is already there. + status.reset(); + status.update(block); } } @@ -334,6 +346,7 @@ public void status(List lines) { private void updateStatus(List lines) { if (lines.isEmpty()) { + block = List.of(); status.update(List.of()); return; } @@ -345,13 +358,14 @@ private void updateStatus(List lines) { // sized in lines, and everything below it is then drawn in the wrong place -- which is how a // long summary tore the block apart. int width = Math.max(10, terminal.getSize().getColumns() - 1); - List block = new java.util.ArrayList<>(); - block.add(new AttributedString(rule(), AttributedStyle.DEFAULT.foreground(AttributedStyle.BRIGHT))); + List rows = new java.util.ArrayList<>(); + rows.add(new AttributedString(rule(), AttributedStyle.DEFAULT.foreground(AttributedStyle.BRIGHT))); for (String line : lines) { - block.add( + rows.add( new AttributedString(fit(line, width), AttributedStyle.DEFAULT.foreground(AttributedStyle.BRIGHT))); } - status.update(block); + block = List.copyOf(rows); + status.update(rows); } /** diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/JLineTerminalTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/JLineTerminalTest.java index 6243b2a85..3bc29feff 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/JLineTerminalTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/JLineTerminalTest.java @@ -180,6 +180,27 @@ void clearingTheScreenWipesItAndLeavesTheReaderWorking() throws Exception { } } + @Test + void theBlockIsBackOnScreenAfterAClear() throws Exception { + // Clearing erases the block along with everything else, and the pinned region is redrawn only + // when its content changes -- so after a clear it believes it is still on screen and draws + // nothing, leaving the bottom of the window empty. + try (Terminal terminal = terminal("go\n"); + JLineTerminal console = JLineTerminal.over(terminal, List.of())) { + // As in a session: the reader owns the screen before anything is drawn into the block. + console.readLine("ignored"); + console.status(List.of("a distinctive state row")); + int before = screen().length(); + + console.clearScreen(); + + assertThat( + "the block is drawn again after the wipe", + screen().substring(before), + containsString("a distinctive state row")); + } + } + @Test void controlLIsBoundToTheReadersOwnClearScreen() throws Exception { // 0x0C is Ctrl-L. It is bound by JLine itself, so /cls is the second way to do this rather From dc0689c45e20d4871651b0fd8a7572c2eeea3e4b Mon Sep 17 00:00:00 2001 From: Bernard Ladenthin Date: Wed, 23 Sep 2026 21:50:15 +0200 Subject: [PATCH 26/36] llama-atmosphere-agent: handle window resizes, and icons in the status line Dragging the window narrower drew a row of "> > > > >" across the screen. Status keeps the size it was created with, so its reserved rows stop matching the window, everything below them is drawn in the wrong place, and a prompt redraw lands beside the previous one instead of over it. A WINCH handler now resizes and resets it and re-renders the block from the text it was built from -- never from the rendered rows, which were cut to a width that no longer exists. Both halves are pinned by tests that go red when the handler is reduced to redraw(). The status line is icons and values now: folder, mode glyph, context, tools, and a robot for a model loaded in this process or a globe for one reached over the network. Each is an icon, a space, its value. fit() therefore measures screen COLUMNS rather than characters -- an icon is one character and two columns, and counting characters lets a row come out wider than the window, wrap, and push the pinned block out of place, which is the tearing this console has had twice already through other doors. Removed again: skipping a status write when the block is unchanged. It looked like the fix for the stray "?1h" -- fewer writes, fewer chances to collide with the setup sequence the reader emits inside JLine where no lock of ours reaches. Taking the guard out again left the emitted bytes identical, because JLine already skips an unchanged block. Deleted rather than kept with a comment claiming a benefit it does not have. The "?1h" therefore still has no established cause, and this commit does not claim one. 133 tests green. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01E2h8gXyyE5UeimkL9vQv9G --- CLAUDE.md | 18 ++++++ llama-atmosphere-agent/README.md | 11 ++++ .../llama/atmosphere/JLineTerminal.java | 54 +++++++++++++++++- .../llama/atmosphere/LocalAgent.java | 12 ++-- .../llama/atmosphere/StatusLine.java | 33 +++++++++-- .../atmosphere/ConsoleFormattingTest.java | 25 ++++++--- .../llama/atmosphere/JLineTerminalTest.java | 56 +++++++++++++++++++ 7 files changed, 189 insertions(+), 20 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 6295d36f9..8925b6217 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -2440,6 +2440,24 @@ are decisions, not details: session, which is what a program that wants its input at the bottom *without* taking over the screen has to pay. + **A window resize has to be handled, and it is the one signal nothing else covers.** `Status` keeps + the size it was built with, so after the window is made narrower its reserved rows no longer match + it, everything below them is drawn in the wrong place, and a prompt redraw lands next to the + previous one instead of over it — a row of `> > > > >` across the screen. A `Signal.WINCH` handler + calls `status.resize()` + `reset()` and re-renders the block **from the text it was built from** + (`requested`), never from the rendered rows: those were cut to a width that no longer exists, and a + row too wide wraps onto a second screen line, which is precisely what the reserved region cannot + survive. Both halves are pinned by tests that go red when the handler is reduced to `redraw()`. + `fit()` measures in **screen columns** (`AttributedString.columnLength`), not characters, for the + same reason — an icon is one character and two columns. + + **What was tried and removed: skipping an unchanged block.** It looked like the fix for the `?1h` + fragment (fewer writes, fewer chances to collide with the reader's own setup sequence). Removing + the guard again left the emitted bytes identical, because **JLine already skips a block whose + content has not changed**. It was deleted rather than kept with a comment claiming a benefit it + does not have — and the `?1h` therefore still has no established cause; the lock covers our writes, + the reader's own are inside JLine. + **A blank line must not count as pending input.** `hasPendingInput()` ignores blank lines but leaves them queued: counting them meant that holding Enter cancelled one turn per keystroke and produced nothing, while dropping them would break the approval prompt, where an empty answer means yes. diff --git a/llama-atmosphere-agent/README.md b/llama-atmosphere-agent/README.md index 30a425b67..73efd97ce 100644 --- a/llama-atmosphere-agent/README.md +++ b/llama-atmosphere-agent/README.md @@ -335,6 +335,17 @@ a few turns and wrong at the start. Pushing the cursor to the last row before th that the state from the beginning. The cost is a screen of blank lines above the session: the alternative is taking over the whole screen (alternate buffer), which costs the scrollback. +**The status line is icons and values**: `📁` workspace, the mode glyph, `📊` context, +`🔧` tools, and `🤖` for a model this process loaded or `🌐` for one reached over the +network. Each is an icon, a space, its value. Rows are cut by **screen columns** rather than characters, +because an icon is one character and two columns — counting characters lets a row come out wider than +the window, wrap, and push the pinned block out of place. + +**Resizing the window** is handled explicitly: the pinned region keeps the size it was created with, so +without a `WINCH` handler its reserved rows stop matching the window and a prompt redraw lands beside +the previous one (a row of `> > > > >` across the screen). The block is then re-rendered from the text +it was built from, not from the rows that were cut for the old width. + **Three threads write to this console** and all of them had to be brought into line, because a write that goes around the line reader scrolls the screen without JLine noticing and the pinned block ends up somewhere else than it believes — first as a stray `1H` drawn into the rule, then as no block at all. diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java index 7b5d9fc38..80bc4e31f 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java @@ -79,6 +79,15 @@ public final class JLineTerminal implements AgentTerminal { */ private volatile List block = List.of(); + /** + * The block as it was asked for, before being cut to the window. + * + *

A resize changes what "cut to the window" means, so the rendered rows cannot be reused — they + * were shortened for a width that no longer exists. What is kept is the text the caller handed + * over, which is re-rendered at the new size. + */ + private volatile List requested = List.of(); + private volatile boolean closed; private @Nullable Thread input; @@ -133,7 +142,32 @@ static JLineTerminal over(Terminal terminal, List completions) { // silently rewrites a request like: git commit -m "fixed!" .option(LineReader.Option.DISABLE_EVENT_EXPANSION, true) .build(); - return new JLineTerminal(terminal, reader, Status.getStatus(terminal), Ansi.detect()); + JLineTerminal console = new JLineTerminal(terminal, reader, Status.getStatus(terminal), Ansi.detect()); + // Without this the pinned region keeps the size it was built with: its reserved rows no longer + // match the window, everything below is drawn in the wrong place, and a redraw of the prompt + // lands next to the previous one instead of over it -- a row of "> > > > >" across the screen + // was the window being made narrower. + terminal.handle(Terminal.Signal.WINCH, ignored -> console.resized()); + return console; + } + + /** + * Redraw everything that was sized for the old window. + * + *

The rows are re-rendered from the text they were built from rather than reused: they were cut + * to a width that no longer exists, and a row that is too wide wraps onto a second screen line, + * which is exactly what the reserved region cannot survive. + */ + void resized() { + synchronized (writing) { + status.resize(); + status.reset(); + List lines = requested; + block = List.of(); // whatever is on screen was drawn for another size + if (!lines.isEmpty()) { + updateStatus(lines); + } + } } @Override @@ -238,7 +272,9 @@ public void clearScreen() { // reset() makes it forget what it believes is on screen; without that the update below is // a no-op, because the content it would draw is the content it thinks is already there. status.reset(); - status.update(block); + // A fresh list every time: JLine keeps the one it is given and works on it, so handing it + // the kept copy makes that copy its own and the next update trips over it. + status.update(new java.util.ArrayList<>(block)); } } @@ -345,7 +381,11 @@ public void status(List lines) { } private void updateStatus(List lines) { + requested = List.copyOf(lines); if (lines.isEmpty()) { + if (block.isEmpty()) { + return; + } block = List.of(); status.update(List.of()); return; @@ -376,7 +416,15 @@ private void updateStatus(List lines) { * @return the row, ending in {@code …} when it had to be cut */ static String fit(String text, int width) { - return text.length() <= width ? text : text.substring(0, Math.max(1, width - 1)) + "…"; + // Counted in screen columns, not characters. An icon or an emoji occupies two columns and one + // character, so cutting by character length lets a row come out wider than the window, wrap + // onto a second screen line, and push everything below the reserved region out of place -- + // the same tearing a long summary caused, arriving through a different door. + AttributedString measured = new AttributedString(text); + if (measured.columnLength() <= width) { + return text; + } + return measured.columnSubSequence(0, Math.max(1, width - 1)).toString() + "…"; } /** diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java index 81edd307b..da8e886b7 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java @@ -243,7 +243,8 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream estimated.get(), contextSize, tools.size(), - options.getModelId()); + options.getModelId(), + options.getModelPath() == null); boolean shortcut = console.onCycleMode(() -> { mode.set(mode.get().next()); if (console.pinsStatus()) { @@ -267,7 +268,8 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream estimated.get(), contextSize, tools.size(), - options.getModelId()); + options.getModelId(), + options.getModelPath() == null); if (terminal.pinsStatus()) { terminal.status(List.of(IDLE_LINE, status)); } else { @@ -327,7 +329,8 @@ options, contextSize, estimateTokens(systemPrompt(options), history) + line.leng running.inputTokens() == 0, contextSize, tools.size(), - options.getModelId()), + options.getModelId(), + options.getModelPath() == null), activity); pendingNote = toolNote(completed.rounds()); estimated.set(completed.inputTokens() == 0); @@ -590,7 +593,8 @@ private static long handleCommand( estimated, contextSize, runner.toolNames().size(), - options.getModelId())); + options.getModelId(), + options.getModelPath() == null)); terminal.line("workspace: " + options.getWorkspace()); terminal.line("history: " + history.size() + " messages"); } diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/StatusLine.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/StatusLine.java index f4ba96a4f..0f18ba717 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/StatusLine.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/StatusLine.java @@ -26,6 +26,21 @@ */ public final class StatusLine { + /** In front of the workspace path. */ + private static final String WORKSPACE_ICON = "📁"; + + /** In front of the context figure. */ + private static final String CONTEXT_ICON = "📊"; + + /** In front of the tool count. */ + private static final String TOOLS_ICON = "🔧"; + + /** In front of a model this process loaded itself. */ + private static final String LOCAL_MODEL_ICON = "🤖"; + + /** In front of a model served by something else over the network. */ + private static final String REMOTE_MODEL_ICON = "🌐"; + /** Above this many characters the workspace path is shortened to its last two segments. */ private static final int MAX_PATH_CHARS = 40; @@ -44,6 +59,7 @@ private StatusLine() {} * @param contextSize the context window in tokens, or {@link #UNKNOWN_CONTEXT} * @param tools how many tools are offered to the model * @param modelId the model id sent in every request + * @param remote whether the model is served by another process rather than loaded here * @return one line, without a trailing newline */ public static String render( @@ -53,9 +69,16 @@ public static String render( boolean estimated, int contextSize, int tools, - String modelId) { - return "[" + shorten(workspace) + " · " + mode.badge() + " · " + context(inputTokens, estimated, contextSize) - + " · " + tools + " tools · " + modelId + "]"; + String modelId, + boolean remote) { + // Every part is an icon, a space, and its value. The icons carry what the words used to, so + // the line stays short enough to survive a narrow window; the space is what keeps an icon from + // running into its value, which is easy to lose when the glyph is wide. + return "[" + WORKSPACE_ICON + " " + shorten(workspace) + + " · " + mode.badge() + + " · " + CONTEXT_ICON + " " + context(inputTokens, estimated, contextSize) + + " · " + TOOLS_ICON + " " + tools + + " · " + (remote ? REMOTE_MODEL_ICON : LOCAL_MODEL_ICON) + " " + modelId + "]"; } /** @@ -83,13 +106,13 @@ static String shorten(java.nio.file.Path workspace) { * @param inputTokens the input tokens of the last turn * @param estimated whether that number is an estimate (rendered with a leading {@code ~}) * @param contextSize the context window in tokens, or {@link #UNKNOWN_CONTEXT} - * @return e.g. {@code "ctx 1.2k/16k"}, or {@code "ctx 1.2k"} when the size is unknown + * @return e.g. {@code "1.2k/16k"}, or {@code "1.2k"} when the size is unknown */ static String context(long inputTokens, boolean estimated, int contextSize) { // Two numbers in k, no percentage: everyone reads 12k/16k at a glance, and a percentage of a // number that is itself an estimate suggests a precision this does not have. String used = (estimated ? "~" : "") + abbreviate(inputTokens); - return contextSize <= UNKNOWN_CONTEXT ? "ctx " + used : "ctx " + used + "/" + abbreviate(contextSize); + return contextSize <= UNKNOWN_CONTEXT ? used : used + "/" + abbreviate(contextSize); } /** diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleFormattingTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleFormattingTest.java index b06d06a5e..c91f08bb3 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleFormattingTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleFormattingTest.java @@ -51,9 +51,11 @@ void aPlainInstanceReturnsTheTextUnchanged() { @Test void theStatusLineShowsWorkspaceModeContextToolsAndModel() { String line = StatusLine.render( - java.nio.file.Path.of("/tmp/ws"), ApprovalMode.MANUAL, 1234, false, 16384, 9, "local-model"); + java.nio.file.Path.of("/tmp/ws"), ApprovalMode.MANUAL, 1234, false, 16384, 9, "local-model", false); - assertThat(line, containsString("ws · ⏸ manual · ctx 1.2k/16k · 9 tools · local-model]")); + // an icon, a space, its value -- the same shape for every part of the line + assertThat(line, containsString("📁 ")); + assertThat(line, containsString("ws · ⏸ manual · 📊 1.2k/16k · 🔧 9 · 🤖 local-model]")); } @Test @@ -81,10 +83,17 @@ void eachModeCarriesItsOwnSymbol() { assertThat(ApprovalMode.MANUAL.badge(), is("⏸ manual")); assertThat(ApprovalMode.AUTO.badge(), is("⏵⏵ auto")); assertThat( - StatusLine.render(java.nio.file.Path.of("/tmp/ws"), ApprovalMode.AUTO, 0, true, 0, 1, "m"), + StatusLine.render(java.nio.file.Path.of("/tmp/ws"), ApprovalMode.AUTO, 0, true, 0, 1, "m", false), containsString("⏵⏵ auto")); } + @Test + void aRemoteEndpointIsMarkedDifferentlyFromAModelLoadedHere() { + java.nio.file.Path ws = java.nio.file.Path.of("/tmp/ws"); + assertThat(StatusLine.render(ws, ApprovalMode.MANUAL, 0, true, 0, 1, "m", false), containsString("🤖 m")); + assertThat(StatusLine.render(ws, ApprovalMode.MANUAL, 0, true, 0, 1, "m", true), containsString("🌐 m")); + } + @Test void aLongWorkspacePathIsShortenedToItsLastTwoSegments() { // the path is on every line of the session, so it must not push the rest off the screen @@ -99,17 +108,17 @@ void aLongWorkspacePathIsShortenedToItsLastTwoSegments() { @Test void contextIsShownInThousandsAndWithoutASizeWhenItIsUnknown() { - assertThat(StatusLine.context(812, false, 32768), is("ctx 812/33k")); - assertThat(StatusLine.context(16000, false, 32768), is("ctx 16k/33k")); - assertThat(StatusLine.context(0, false, StatusLine.UNKNOWN_CONTEXT), is("ctx 0")); - assertThat(StatusLine.context(2500, false, StatusLine.UNKNOWN_CONTEXT), is("ctx 2.5k")); + assertThat(StatusLine.context(812, false, 32768), is("812/33k")); + assertThat(StatusLine.context(16000, false, 32768), is("16k/33k")); + assertThat(StatusLine.context(0, false, StatusLine.UNKNOWN_CONTEXT), is("0")); + assertThat(StatusLine.context(2500, false, StatusLine.UNKNOWN_CONTEXT), is("2.5k")); } @Test void anEstimatedCountIsMarkedWithATilde() { // llama.cpp reports usage only to clients that ask for it, and Atmosphere does not, so the // number normally comes from LocalAgent.estimateTokens -- the tilde says so. - assertThat(StatusLine.context(2500, true, 16384), is("ctx ~2.5k/16k")); + assertThat(StatusLine.context(2500, true, 16384), is("~2.5k/16k")); assertThat( LocalAgent.estimateTokens( "0123456789", java.util.List.of(org.atmosphere.ai.llm.ChatMessage.user("0123456789"))), diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/JLineTerminalTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/JLineTerminalTest.java index 3bc29feff..4ee10fc42 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/JLineTerminalTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/JLineTerminalTest.java @@ -180,6 +180,43 @@ void clearingTheScreenWipesItAndLeavesTheReaderWorking() throws Exception { } } + @Test + void makingTheWindowNarrowerRedrawsTheBlockAtTheNewWidth() throws Exception { + // Reported as a row of "> > > > >" across the screen after dragging the window smaller. The + // pinned region keeps the size it was built with unless it is told, so its reserved rows stop + // matching the window and everything below them is drawn in the wrong place. + try (Terminal terminal = terminal("\n"); + JLineTerminal console = JLineTerminal.over(terminal, List.of())) { + console.status(List.of("state")); + int before = screen().length(); + + terminal.setSize(new Size(30, 10)); + console.resized(); + + String afterResize = screen().substring(before); + assertThat("the block was drawn again", afterResize.isEmpty(), is(false)); + assertThat( + "and never again at the width of the window that is gone", + afterResize.contains("─".repeat(SIZE.getColumns() - 1)), + is(false)); + } + } + + @Test + void aRowTooWideForTheNewWindowIsCutRatherThanWrapped() throws Exception { + try (Terminal terminal = terminal("\n"); + JLineTerminal console = JLineTerminal.over(terminal, List.of())) { + console.status(List.of("x".repeat(50))); + + terminal.setSize(new Size(20, 10)); + console.resized(); + + assertThat( + "the row that was rendered for the wide window is not reused", occurrences("x".repeat(50)), is(1)); + assertThat("it is cut for the narrow one", screen(), containsString("…")); + } + } + @Test void theBlockIsBackOnScreenAfterAClear() throws Exception { // Clearing erases the block along with everything else, and the pinned region is redrawn only @@ -212,6 +249,25 @@ void controlLIsBoundToTheReadersOwnClearScreen() throws Exception { } } + @Test + void aRowIsCutByScreenColumnsNotByCharacters() { + // An icon takes two columns and one character. Cutting by character length lets the row come + // out wider than the window, wrap onto a second screen line, and push everything below the + // reserved region out of place -- the tearing that a long summary caused, through another door. + String icons = "📁".repeat(20); + + String cut = JLineTerminal.fit(icons, 10); + + // What matters is the width on screen, not how many characters that took. + assertThat("the row fits the window", new org.jline.utils.AttributedString(cut).columnLength() <= 10, is(true)); + assertThat( + "a cut by characters would have kept nine icons, which is eighteen columns", + cut.codePointCount(0, cut.length()) < 9, + is(true)); + assertThat(cut.endsWith("…"), is(true)); + assertThat("plain text is untouched when it fits", JLineTerminal.fit("short", 10), is("short")); + } + @Test void aMultiLineStringIsStillPrintedAsSeveralLines() throws Exception { try (Terminal terminal = terminal("\n"); From 19f2765c47ad101ea9f8e87111b5aedc50d28e94 Mon Sep 17 00:00:00 2001 From: Bernard Ladenthin Date: Wed, 23 Sep 2026 21:59:09 +0200 Subject: [PATCH 27/36] llama-atmosphere-agent: drop the resize handler that made resizing worse I added a WINCH handler for the "> > > > >" row after a resize; it was reported worse, not better. LineReaderImpl.handleSignal already calls Status.resize, and the reader installs its own handler for as long as it is reading -- the whole session -- so mine either never ran or ran in addition, putting a second writer on the terminal from the signal thread while the reader was redrawing. Removed. A probe raising a real resize signal against a pipe-backed terminal shows JLine doing it correctly alone: scroll region reset, rule re-cut to the new width, one prompt. So the remaining report is not reproducible here and this commit does not claim to fix it. Also removed, both because measurement said they did nothing: skipping a status write when the block is unchanged (JLine already does), and re-cutting rows on a width change (Status.resize does). Neither is kept with a comment claiming a benefit it does not have. What is kept and does matter: fit() measures screen columns rather than characters, because an icon is one character and two columns; and restoring the block after a wipe takes reset, then an empty update, then a render from the text the caller gave -- with only the first two, Status emitted a single character where a whole block was missing. 133 tests green. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01E2h8gXyyE5UeimkL9vQv9G --- CLAUDE.md | 43 +++++++++------ llama-atmosphere-agent/README.md | 10 ++-- .../llama/atmosphere/JLineTerminal.java | 55 ++++++++----------- .../llama/atmosphere/JLineTerminalTest.java | 12 ++-- 4 files changed, 61 insertions(+), 59 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 8925b6217..c65b747fc 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -2440,23 +2440,32 @@ are decisions, not details: session, which is what a program that wants its input at the bottom *without* taking over the screen has to pay. - **A window resize has to be handled, and it is the one signal nothing else covers.** `Status` keeps - the size it was built with, so after the window is made narrower its reserved rows no longer match - it, everything below them is drawn in the wrong place, and a prompt redraw lands next to the - previous one instead of over it — a row of `> > > > >` across the screen. A `Signal.WINCH` handler - calls `status.resize()` + `reset()` and re-renders the block **from the text it was built from** - (`requested`), never from the rendered rows: those were cut to a width that no longer exists, and a - row too wide wraps onto a second screen line, which is precisely what the reserved region cannot - survive. Both halves are pinned by tests that go red when the handler is reduced to `redraw()`. - `fit()` measures in **screen columns** (`AttributedString.columnLength`), not characters, for the - same reason — an icon is one character and two columns. - - **What was tried and removed: skipping an unchanged block.** It looked like the fix for the `?1h` - fragment (fewer writes, fewer chances to collide with the reader's own setup sequence). Removing - the guard again left the emitted bytes identical, because **JLine already skips a block whose - content has not changed**. It was deleted rather than kept with a comment claiming a benefit it - does not have — and the `?1h` therefore still has no established cause; the lock covers our writes, - the reader's own are inside JLine. + **Do not add a `WINCH` handler, and the reason is measured.** A resize drawing a row of + `> > > > >` across the screen looks like the pinned region not being told about the new size, so a + `Signal.WINCH` handler that resized and re-rendered it was added — and the user reported it + **worse**, not better. `LineReaderImpl.handleSignal(WINCH)` already calls `Status.resize(Size)`, + and the reader installs its own handler for as long as it is reading, which is the whole session; + ours therefore either never ran or ran *in addition*, putting a second writer on the terminal from + the signal thread at the exact moment the reader was redrawing. A probe driving a real + `terminal.raise(WINCH)` against a pipe-backed terminal shows JLine doing it correctly on its own: + scroll region reset, the rule re-cut to the new width, **one** prompt. So the remaining report is + not reproducible in the harness and has no fix here yet — stated rather than papered over. + `fit()` measuring in **screen columns** (`AttributedString.columnLength`) rather than characters is + ours and does matter: an icon is one character and two columns, and a row wider than the window + wraps onto a second screen line, which the reserved region cannot survive. + + **Two things tried and removed, both because measurement said they did nothing.** (1) Skipping a + status write when the block is unchanged — removing the guard again left the emitted bytes + identical, because JLine already skips an unchanged block. (2) Re-cutting the rows when the window + width changed — `Status.resize()` re-cuts the rows it holds itself. Neither was kept with a comment + claiming a benefit it does not have. The stray `?1h` consequently still has **no established + cause**: the lock covers our writes, the reader's own are inside JLine. + + **Restoring the block after a wipe takes three steps, found by measurement not by reading**: + `status.reset()`, then `status.update(List.of())`, then render it again from `requested` (the text + the caller gave, not the rendered rows). With only the first two, `Status` draws the difference it + computes against a belief the wipe invalidated — observed as a single character emitted where a + whole block was missing. `JLineTerminalTest.theBlockIsBackOnScreenAfterAClear` is what says so. **A blank line must not count as pending input.** `hasPendingInput()` ignores blank lines but leaves them queued: counting them meant that holding Enter cancelled one turn per keystroke and produced diff --git a/llama-atmosphere-agent/README.md b/llama-atmosphere-agent/README.md index 73efd97ce..f77f43a16 100644 --- a/llama-atmosphere-agent/README.md +++ b/llama-atmosphere-agent/README.md @@ -341,10 +341,12 @@ network. Each is an icon, a space, its value. Rows are cut by **screen columns** because an icon is one character and two columns — counting characters lets a row come out wider than the window, wrap, and push the pinned block out of place. -**Resizing the window** is handled explicitly: the pinned region keeps the size it was created with, so -without a `WINCH` handler its reserved rows stop matching the window and a prompt redraw lands beside -the previous one (a row of `> > > > >` across the screen). The block is then re-rendered from the text -it was built from, not from the rows that were cut for the old width. +**Resizing the window** is left to JLine, which resizes the pinned region and re-cuts its rows itself. +A handler of our own was tried for a reported row of `> > > > >` after dragging the window smaller and +made it worse: the line reader installs its own handler for as long as it is reading, so ours only +added a second writer on the terminal while the reader was redrawing. That report is **not currently +reproducible** here — a probe raising a real resize signal shows a clean redraw with one prompt — so it +is listed as open rather than claimed as fixed. **Three threads write to this console** and all of them had to be brought into line, because a write that goes around the line reader scrolls the screen without JLine noticing and the pinned block ends up diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java index 80bc4e31f..3eb1e935b 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java @@ -80,11 +80,12 @@ public final class JLineTerminal implements AgentTerminal { private volatile List block = List.of(); /** - * The block as it was asked for, before being cut to the window. + * The block as the caller asked for it, before being cut to the window. * - *

A resize changes what "cut to the window" means, so the rendered rows cannot be reused — they - * were shortened for a width that no longer exists. What is kept is the text the caller handed - * over, which is re-rendered at the new size. + *

Kept so it can be drawn from scratch after the screen is wiped. Handing the rendered rows + * back is not enough: {@code Status} draws the difference between them and what it believes is on + * screen, and after a wipe that belief is wrong in a way it cannot detect — measured, it emitted a + * single character where a whole block was missing. */ private volatile List requested = List.of(); @@ -142,32 +143,12 @@ static JLineTerminal over(Terminal terminal, List completions) { // silently rewrites a request like: git commit -m "fixed!" .option(LineReader.Option.DISABLE_EVENT_EXPANSION, true) .build(); - JLineTerminal console = new JLineTerminal(terminal, reader, Status.getStatus(terminal), Ansi.detect()); - // Without this the pinned region keeps the size it was built with: its reserved rows no longer - // match the window, everything below is drawn in the wrong place, and a redraw of the prompt - // lands next to the previous one instead of over it -- a row of "> > > > >" across the screen - // was the window being made narrower. - terminal.handle(Terminal.Signal.WINCH, ignored -> console.resized()); - return console; - } - - /** - * Redraw everything that was sized for the old window. - * - *

The rows are re-rendered from the text they were built from rather than reused: they were cut - * to a width that no longer exists, and a row that is too wide wraps onto a second screen line, - * which is exactly what the reserved region cannot survive. - */ - void resized() { - synchronized (writing) { - status.resize(); - status.reset(); - List lines = requested; - block = List.of(); // whatever is on screen was drawn for another size - if (!lines.isEmpty()) { - updateStatus(lines); - } - } + // No WINCH handler here on purpose. The line reader installs its own for as long as it is + // reading -- which is the whole session -- and it already resizes the pinned region itself + // (LineReaderImpl.handleSignal calls Status.resize). Adding one of ours only put a second + // writer on the terminal, on the signal thread, at the exact moment the reader was redrawing: + // the row of "> > > > >" after a resize got worse, not better, when it was tried. + return new JLineTerminal(terminal, reader, Status.getStatus(terminal), Ansi.detect()); } @Override @@ -271,10 +252,18 @@ public void clearScreen() { reader.printAbove(clear + System.lineSeparator().repeat(blankRows())); // reset() makes it forget what it believes is on screen; without that the update below is // a no-op, because the content it would draw is the content it thinks is already there. + List lines = requested; + // Three steps, and all three were needed to make the block come back after a wipe: + // forget the drawing state, hand over an empty block so nothing is believed to be on + // screen, then render it again. With only the first two, Status drew the difference it + // computed against a belief the wipe had invalidated -- measured as a single character + // where a whole block was missing. status.reset(); - // A fresh list every time: JLine keeps the one it is given and works on it, so handing it - // the kept copy makes that copy its own and the next update trips over it. - status.update(new java.util.ArrayList<>(block)); + block = List.of(); + status.update(List.of()); + if (!lines.isEmpty()) { + updateStatus(lines); + } } } diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/JLineTerminalTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/JLineTerminalTest.java index 4ee10fc42..0aeb8438c 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/JLineTerminalTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/JLineTerminalTest.java @@ -182,16 +182,18 @@ void clearingTheScreenWipesItAndLeavesTheReaderWorking() throws Exception { @Test void makingTheWindowNarrowerRedrawsTheBlockAtTheNewWidth() throws Exception { - // Reported as a row of "> > > > >" across the screen after dragging the window smaller. The - // pinned region keeps the size it was built with unless it is told, so its reserved rows stop - // matching the window and everything below them is drawn in the wrong place. + // This pins an assumption about JLine rather than logic of ours: it re-cuts the rows it holds + // when the window shrinks, so nothing here has to. That is worth a test because the whole + // bottom block depends on it -- a row wider than the window wraps onto a second screen line, + // and the reserved region cannot survive that. An upgrade that changed it would show up here + // instead of on somebody's screen. try (Terminal terminal = terminal("\n"); JLineTerminal console = JLineTerminal.over(terminal, List.of())) { console.status(List.of("state")); int before = screen().length(); terminal.setSize(new Size(30, 10)); - console.resized(); + console.status(List.of("state")); String afterResize = screen().substring(before); assertThat("the block was drawn again", afterResize.isEmpty(), is(false)); @@ -209,7 +211,7 @@ void aRowTooWideForTheNewWindowIsCutRatherThanWrapped() throws Exception { console.status(List.of("x".repeat(50))); terminal.setSize(new Size(20, 10)); - console.resized(); + console.status(List.of("x".repeat(50))); assertThat( "the row that was rendered for the wide window is not reused", occurrences("x".repeat(50)), is(1)); From 3dd4581114c6d58c82e2435d058771729ad34815 Mon Sep 17 00:00:00 2001 From: Bernard Ladenthin Date: Wed, 23 Sep 2026 22:06:47 +0200 Subject: [PATCH 28/36] llama-atmosphere-agent: pin the resize path with a test that drives the real signal everyResizeDrawsExactlyOnePrompt drives what actually happens on a resize -- a size change plus terminal.raise(WINCH), with the reader sitting in readLine as it does all session -- across shrinking, growing and a changed row count. Each one draws exactly one prompt, run repeatedly to check it is not flaky. That is the finding worth keeping: the leftover "> " row still being reported is not produced by this path. The remaining suspect is the console reflowing its own screen buffer on a resize, which moves lines the program never wrote again and which nothing on this side can reproduce. Written down as open, with /cls and Ctrl-L as the tested way out. A second test asserting "a clear leaves exactly one prompt" was written and dropped: it depends on where the reader thread happens to be, and the two existing clear tests already cover the erase reaching the screen and the block coming back. 134 tests green. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01E2h8gXyyE5UeimkL9vQv9G --- .../llama/atmosphere/JLineTerminalTest.java | 38 ++++++++++++++++++- 1 file changed, 37 insertions(+), 1 deletion(-) diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/JLineTerminalTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/JLineTerminalTest.java index 0aeb8438c..1f7332f53 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/JLineTerminalTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/JLineTerminalTest.java @@ -45,8 +45,12 @@ class JLineTerminalTest { private final ByteArrayOutputStream emitted = new ByteArrayOutputStream(); private Terminal terminal(String keystrokes) throws Exception { + return terminal(new ByteArrayInputStream(keystrokes.getBytes(StandardCharsets.UTF_8))); + } + + private Terminal terminal(java.io.InputStream keystrokes) throws Exception { return TerminalBuilder.builder() - .streams(new ByteArrayInputStream(keystrokes.getBytes(StandardCharsets.UTF_8)), emitted) + .streams(keystrokes, emitted) .type("xterm-256color") // The rule is drawn with U+2500. Both are needed: the writer encodes through the // stdout charset, which is not the one .encoding() sets, and without it every rule @@ -219,6 +223,38 @@ void aRowTooWideForTheNewWindowIsCutRatherThanWrapped() throws Exception { } } + @Test + void everyResizeDrawsExactlyOnePrompt() throws Exception { + // The reported artefact is a second, stale "> " left on screen after dragging the window. + // This drives the path that redraws it -- a real size change plus the signal, with the reader + // sitting in readLine as it does all session -- and pins that shrinking, growing and changing + // the row count each produce one prompt and not two. It holds for every size tried, which is + // what says the remaining artefact is not in this path. + java.io.PipedOutputStream keys = new java.io.PipedOutputStream(); + try (Terminal terminal = terminal(new java.io.PipedInputStream(keys)); + JLineTerminal console = JLineTerminal.over(terminal, List.of())) { + Thread reader = new Thread(() -> console.readLine("ignored")); + reader.setDaemon(true); + reader.start(); + Thread.sleep(200); + console.status(List.of("state row")); + + int[][] sizes = {{30, 10}, {90, 10}, {45, 10}, {120, 24}}; + for (int[] size : sizes) { + int before = screen().length(); + + terminal.setSize(new Size(size[0], size[1])); + terminal.raise(Terminal.Signal.WINCH); + Thread.sleep(150); + + String drawn = screen().substring(before); + long prompts = + drawn.chars().filter(character -> character == '>').count(); + assertThat("one prompt after resizing to " + size[0] + "x" + size[1], prompts, is(1L)); + } + } + } + @Test void theBlockIsBackOnScreenAfterAClear() throws Exception { // Clearing erases the block along with everything else, and the pinned region is redrawn only From 55aaaf050f10263e485bec0690a5b667896a17b1 Mon Sep 17 00:00:00 2001 From: Bernard Ladenthin Date: Wed, 23 Sep 2026 22:07:06 +0200 Subject: [PATCH 29/36] =?UTF-8?q?llama-atmosphere-agent:=20README=20?= =?UTF-8?q?=E2=80=94=20say=20what=20the=20resize=20test=20proves=20and=20w?= =?UTF-8?q?hat=20stays=20open?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The previous commit updated CLAUDE.md but the README edit did not apply, so the two said different things about the same finding. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01E2h8gXyyE5UeimkL9vQv9G --- llama-atmosphere-agent/README.md | 9 ++++++--- 1 file changed, 6 insertions(+), 3 deletions(-) diff --git a/llama-atmosphere-agent/README.md b/llama-atmosphere-agent/README.md index f77f43a16..5217d899f 100644 --- a/llama-atmosphere-agent/README.md +++ b/llama-atmosphere-agent/README.md @@ -344,9 +344,12 @@ the window, wrap, and push the pinned block out of place. **Resizing the window** is left to JLine, which resizes the pinned region and re-cuts its rows itself. A handler of our own was tried for a reported row of `> > > > >` after dragging the window smaller and made it worse: the line reader installs its own handler for as long as it is reading, so ours only -added a second writer on the terminal while the reader was redrawing. That report is **not currently -reproducible** here — a probe raising a real resize signal shows a clean redraw with one prompt — so it -is listed as open rather than claimed as fixed. +added a second writer on the terminal while the reader was redrawing. A test drives the real path — a +size change plus the resize signal, with the reader reading as it does all session — across shrinking, +growing and a changed row count, and every one draws exactly one prompt. The leftover `> ` row that is +still reported is therefore not produced there; the remaining suspect is the console reflowing its own +screen buffer on a resize, which moves lines the program never wrote again and which nothing on this +side can reproduce. `/cls` or Ctrl-L cleans it up. **Three threads write to this console** and all of them had to be brought into line, because a write that goes around the line reader scrolls the screen without JLine noticing and the pinned block ends up From 10304049bc53e736989c70125dca19578f5f1831 Mon Sep 17 00:00:00 2001 From: Bernard Ladenthin Date: Wed, 23 Sep 2026 22:20:26 +0200 Subject: [PATCH 30/36] llama-atmosphere-agent: drive interrupt-then-continue end to end in a test Reported as "after an interrupt it does not carry on by itself". The sequence is now driven through the real OpenAiCompatServer with a scripted backend: a turn is cut short by pending input, and the next turn is then run. It reaches the server, carries the interrupted question in its history, and answers. So the mechanism works, and this commit does not claim to fix the report -- the cause is elsewhere and I would rather say that than ship a guess. Two things the writing of the test settled, both worth keeping. A turn that finishes within one activity tick is never looked at for interruption, which is correct (there is nothing to cut short) but means the scripted backend has to be slowed down or the test proves nothing -- the first version passed while interrupting nothing at all. And a second line typed during the replacement turn stops that one too: that is the design, not a defect, but from the outside it looks exactly like a turn that never started, which is the likeliest way to misread the symptom. 136 tests green. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01E2h8gXyyE5UeimkL9vQv9G --- CLAUDE.md | 10 + .../llama/atmosphere/InterruptedTurnTest.java | 211 ++++++++++++++++++ 2 files changed, 221 insertions(+) create mode 100644 llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/InterruptedTurnTest.java diff --git a/CLAUDE.md b/CLAUDE.md index c65b747fc..21db4860f 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -2467,6 +2467,16 @@ are decisions, not details: computes against a belief the wipe invalidated — observed as a single character emitted where a whole block was missing. `JLineTerminalTest.theBlockIsBackOnScreenAfterAClear` is what says so. + **The turn after an interrupted one is pinned end to end** (`InterruptedTurnTest`): with a scripted + backend behind the real `OpenAiCompatServer`, a turn is cut short by pending input and the next one + is then driven through — it reaches the server, carries the interrupted question in its history, + and answers. Reported as "it does not carry on by itself"; the mechanism works, so the cause of + that report is elsewhere and is **not** claimed to be fixed. Two things the writing of it settled: + a turn that finishes inside one activity tick is never even looked at for interruption (correct — + there is nothing to cut short), which is why the scripted backend has to be made slow or the test + proves nothing; and a second line typed during the replacement turn stops that one too, which is + the design and not a defect, but looks from the outside exactly like a turn that never started. + **A blank line must not count as pending input.** `hasPendingInput()` ignores blank lines but leaves them queued: counting them meant that holding Enter cancelled one turn per keystroke and produced nothing, while dropping them would break the approval prompt, where an empty answer means yes. diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/InterruptedTurnTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/InterruptedTurnTest.java new file mode 100644 index 000000000..98fb4b0d3 --- /dev/null +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/InterruptedTurnTest.java @@ -0,0 +1,211 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import static org.hamcrest.MatcherAssert.assertThat; +import static org.hamcrest.Matchers.containsString; +import static org.hamcrest.Matchers.hasSize; +import static org.hamcrest.Matchers.is; + +import com.fasterxml.jackson.databind.JsonNode; +import java.nio.file.Path; +import java.util.ArrayList; +import java.util.List; +import net.ladenthin.llama.server.OpenAiCompatServer; +import net.ladenthin.llama.server.OpenAiServerConfig; +import org.atmosphere.ai.RetryPolicy; +import org.atmosphere.ai.fs.AgentFileSystem; +import org.atmosphere.ai.fs.WorkspaceAgentFileSystem; +import org.atmosphere.ai.llm.ChatMessage; +import org.junit.jupiter.api.Test; +import org.junit.jupiter.api.io.TempDir; + +/** + * Typing while the agent works stops that turn — and the turn after it has to run. + * + *

Reported as "it does not carry on by itself": the interruption printed its line and then nothing + * followed. The sequence is driven here end to end against the real {@link OpenAiCompatServer} with + * scripted llama.cpp chunks, because every part of it is a different piece of machinery — the + * cancellable entry point of the runtime, the queue the console keeps, and the history the next turn + * is sent with — and a defect in any of them looks identical from the outside. + */ +class InterruptedTurnTest { + + private static final String MODEL_ID = "local-model"; + + /** Longer than one activity tick, so the wait actually looks for typed input. */ + private static final java.time.Duration SLOW_ENOUGH_TO_INTERRUPT = java.time.Duration.ofMillis(700); + + @TempDir + Path workspace; + + /** A console that answers "yes, something was typed" whenever the test says so. */ + private final class Typing implements AgentTerminal { + private volatile boolean pending; + private final List lines = new ArrayList<>(); + + @Override + public void line(String text) { + lines.add(text); + } + + @Override + public String readLine(String prompt) { + return null; + } + + @Override + public String readKey(String prompt) { + return null; + } + + @Override + public boolean hasPendingInput() { + return pending; + } + + @Override + public void status(List statusLines) {} + + @Override + public boolean pinsStatus() { + return false; + } + + @Override + public Ansi ansi() { + return Ansi.PLAIN; + } + + @Override + public void close() {} + } + + private final Typing terminal = new Typing(); + + private OpenAiCompatServer server(ScriptedBackend backend) throws Exception { + return new OpenAiCompatServer( + backend, + OpenAiServerConfig.builder() + .host("127.0.0.1") + .port(0) + .modelId(MODEL_ID) + .build()) + .start(); + } + + private AgentRunner runner(OpenAiCompatServer server) { + return new AgentRunner( + "http://127.0.0.1:" + server.getPort() + "/v1", + "k", + MODEL_ID, + List.of(), + "You are a test agent.", + 0.0, + 64, + 4) + .retryPolicy(RetryPolicy.NONE); + } + + /** + * A turn that takes long enough to be interrupted. + * + *

Without this the test proves nothing: the interruption is only looked for while waiting for + * the turn, and a scripted answer arrives before the first look. That is also the behaviour in a + * real session — a turn that is already finished is not cut short — so the delay is what makes + * this the reported situation rather than a different one. + * + * @param call which model call this is + * @return the scripted answer, after a pause on the first call + */ + private static List slowFirstTurn(int call) { + if (call == 1) { + try { + Thread.sleep(SLOW_ENOUGH_TO_INTERRUPT.toMillis()); + } catch (InterruptedException e) { + Thread.currentThread().interrupt(); + } + } + return ScriptedBackend.textTurn("answer " + call); + } + + @Test + void theTurnAfterAnInterruptedOneRunsAndCarriesTheHistory() throws Exception { + ScriptedBackend backend = new ScriptedBackend((call, request) -> slowFirstTurn(call)); + try (OpenAiCompatServer server = server(backend)) { + AgentRunner runner = runner(server); + AgentFileSystem files = new WorkspaceAgentFileSystem(workspace, AgentFileSystem.Limits.defaults()); + List history = new ArrayList<>(); + ToolCallLog log = new ToolCallLog(); + + // The user types while the first turn is still running. + terminal.pending = true; + ConsoleSession first = LocalAgent.turn( + runner, files, "a poem please", history, terminal, log, 1, ignored -> "", new TurnActivity()); + + assertThat( + "the interruption is announced", + terminal.lines.stream().anyMatch(line -> line.contains("interrupted")), + is(true)); + + // What was typed is now the next message, and nothing is pending any more. + terminal.pending = false; + ConsoleSession second = LocalAgent.turn( + runner, files, "make it longer", history, terminal, log, 2, ignored -> "", new TurnActivity()); + + assertThat("the turn after the interruption produced an answer", second.text(), containsString("answer")); + assertThat( + "and it was not itself reported as failed", + second.failure(), + is(org.hamcrest.Matchers.nullValue())); + assertThat( + "the interrupted turn is in the history as asked", + history.get(0).content(), + is("a poem please")); + + List requests = backend.requests(); + assertThat("both turns reached the server", requests.size() >= 2, is(true)); + JsonNode last = requests.get(requests.size() - 1).path("messages"); + assertThat( + "the second request carries the typed line", + last.get(last.size() - 1).path("content").asText(), + is("make it longer")); + assertThat("and the interrupted question before it", last.toString(), containsString("a poem please")); + assertThat(first.rounds(), hasSize(0)); + } + } + + @Test + void aSecondLineTypedDuringTheReplacementTurnStopsThatOneToo() throws Exception { + // Not a defect: each typed line overtakes the turn it arrived in. It is pinned because the + // symptom -- an interruption that seems to lead nowhere -- looks the same as a turn that never + // starts, and telling the two apart afterwards is what took the longest. + ScriptedBackend backend = new ScriptedBackend((call, request) -> { + try { + Thread.sleep(SLOW_ENOUGH_TO_INTERRUPT.toMillis()); + } catch (InterruptedException e) { + Thread.currentThread().interrupt(); + } + return ScriptedBackend.textTurn("answer " + call); + }); + try (OpenAiCompatServer server = server(backend)) { + AgentRunner runner = runner(server); + AgentFileSystem files = new WorkspaceAgentFileSystem(workspace, AgentFileSystem.Limits.defaults()); + List history = new ArrayList<>(); + ToolCallLog log = new ToolCallLog(); + + terminal.pending = true; + LocalAgent.turn(runner, files, "one", history, terminal, log, 1, ignored -> "", new TurnActivity()); + LocalAgent.turn(runner, files, "two", history, terminal, log, 2, ignored -> "", new TurnActivity()); + + assertThat( + "both were cut short, and both said so", + terminal.lines.stream() + .filter(line -> line.contains("interrupted")) + .count(), + is(2L)); + } + } +} From 877ec547b4ec76dd66da5d8555f4a20c1ddb8769 Mon Sep 17 00:00:00 2001 From: Bernard Ladenthin Date: Thu, 24 Sep 2026 18:48:14 +0200 Subject: [PATCH 31/36] llama-atmosphere-agent: --plain, a second console that only appends lines An additional interface, next to the existing one rather than instead of it. The line-oriented console was already there -- it is what a piped or one-shot run has always used -- but nothing could ask for it on a real terminal. --plain does. It only ever appends: the status is printed as an ordinary line before the prompt and scrolls away with everything else, there is no pinned block and no spinner, and nothing on screen is rewritten. That is what a session needs when it is piped, logged, recorded, or carried by anything that forwards lines rather than a screen. LocalAgent.usesFullTerminal(options, interactive) is now the single place that decides between the two, which is also what made it testable: the rich console needs someone typing AND permission to move the cursor, and --plain withholds the second even when the first is true. Worth saying plainly in the docs, because it is the likely reason to reach for this flag and it is not the right one: a normal SSH session does not need it. A remote terminal reports its size and handles cursor control like a local one. --plain is for the cases where that is not true. The trade is documented as a table rather than left to be discovered: no input while the agent works (PlainTerminal reports nothing pending, so a typed line cannot stop a turn), no history, no Tab completion, no Ctrl-L, no Shift+Tab. 137 tests green. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01E2h8gXyyE5UeimkL9vQv9G --- CLAUDE.md | 10 +++++- llama-atmosphere-agent/README.md | 32 +++++++++++++++++++ .../llama/atmosphere/AgentOptions.java | 24 +++++++++++++- .../llama/atmosphere/LocalAgent.java | 19 ++++++++++- .../llama/atmosphere/AgentOptionsTest.java | 18 +++++++++++ 5 files changed, 100 insertions(+), 3 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 21db4860f..4c7645ee5 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -2257,7 +2257,15 @@ are decisions, not details: 2. **One-shot (`--prompt`) denies a gated call** instead of auto-approving it — `--auto` is the deliberate opt-in. Atmosphere itself fails closed when no strategy is wired, and this keeps that direction: an unattended run must not be the most permissive one. -3. **`AgentTerminal` has exactly two implementations, chosen once at startup.** `JLineTerminal` (a real +3. **`AgentTerminal` has exactly two implementations, chosen once at startup, and `--plain` picks the + line-oriented one on purpose.** `LocalAgent.usesFullTerminal(options, interactive)` is the single + place that decides: the cursor-controlling console needs someone typing **and** permission to move + the cursor, and `--plain` withholds the second even on a real terminal. That is not a fallback but + a supported mode — for a session that is piped, logged, recorded, or carried by something that + forwards lines rather than a screen. It gives up the pinned block, the spinner, history/completion + and typing-during-a-turn (`PlainTerminal.hasPendingInput()` is always false), and gains being + correct when the output is a file. A normal SSH session needs none of this: a remote terminal + reports its size and handles cursor control like a local one. `JLineTerminal` (a real terminal: line editing, history, Tab completion of the command names, a status line pinned to the bottom via JLine's `Status`, single-key answers through `enterRawMode`, streamed output via `LineReader.printAbove` so the bottom block stays put) and `PlainTerminal` (a `PrintStream` plus a diff --git a/llama-atmosphere-agent/README.md b/llama-atmosphere-agent/README.md index 5217d899f..3f98569f4 100644 --- a/llama-atmosphere-agent/README.md +++ b/llama-atmosphere-agent/README.md @@ -175,6 +175,38 @@ code page it saw at startup, so umlauts and emoji in the answer would turn into project's `.mvn/jvm.config` pins `-Dstdout.encoding=UTF-8 -Dstderr.encoding=UTF-8` for the `mvn` JVM so both sides agree. +### Two consoles: the full one, and `--plain` + +There are two, and both stay. The default is the **full console** described below: a block pinned to +the bottom of the window, an input line that is there while the agent works, a spinner, single-key +navigation. It positions the cursor, so it needs a terminal that reports its size and understands the +sequences. + +`--plain` chooses the **line-oriented console** instead. It only ever appends lines: the status is +printed as an ordinary line before the prompt and scrolls away with everything else, there is no +pinned block and no spinner, and nothing on screen is ever rewritten. That makes a session readable +when it is piped, logged, recorded, or carried by anything that forwards lines rather than a screen: + +```bash +mvn -q compile exec:java -Dexec.args="--base-url http://127.0.0.1:8080/v1 --plain" | tee session.log +``` + +The same console is what the agent falls back to on its own when there is no usable terminal — a pipe, +a `dumb` terminal, an editor's run window — so `--plain` only *forces* what would otherwise be +detected. Note that a normal SSH session does **not** need it: a remote terminal reports its size and +handles cursor control like a local one. It is for the cases where that is not true. + +What the line-oriented console gives up, so the choice is an informed one: + +| | full (default) | `--plain` | +|---|---|---| +| input while the agent works | yes, and a typed line stops the turn | no, the prompt appears between turns | +| status | pinned at the bottom | printed once before each prompt | +| activity / spinner | yes | dropped rather than repeated into the log | +| approvals | `y` + Enter | `y` + Enter | +| history, Tab completion, Ctrl-L, Shift+Tab | yes | no | +| correct when the output is a file | — | yes | + ### Commands, approval and the status line A line that starts with `/` and names a command is answered by the agent itself; anything else — an diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentOptions.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentOptions.java index 769adeb88..35028c5cd 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentOptions.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentOptions.java @@ -68,6 +68,7 @@ public final class AgentOptions { private final String modelId; private final Path workspace; private final boolean allowShell; + private final boolean plain; private final boolean auto; private final boolean autoCompact; private final int compactAt; @@ -89,6 +90,7 @@ private AgentOptions(Builder b) { this.modelId = b.modelId; this.workspace = b.workspace; this.allowShell = b.allowShell; + this.plain = b.plain; this.auto = b.auto; this.autoCompact = b.autoCompact; this.compactAt = b.compactAt; @@ -115,6 +117,7 @@ public static AgentOptions parse(String[] args) { switch (a) { case "-h", "--help" -> b.help = true; case "--allow-shell" -> b.allowShell = true; + case "--plain" -> b.plain = true; case "--auto" -> b.auto = true; case "--auto-compact" -> b.autoCompact = booleanValue(args, ++i, a); case "--compact-at" -> b.compactAt = percentValue(args, ++i, a); @@ -208,6 +211,7 @@ public static String usage() { "Agent:", " --workspace

directory the file tools are confined to (default: cwd)", " --allow-shell add the run_command tool (runs any command line, starting in the workspace)", + " --plain line-oriented console: no pinned block, no cursor control", " --auto run tools without asking (default: ask before writes and commands)", " --auto-compact summarize the history before it overflows the context (default " + DEFAULT_AUTO_COMPACT + ")", @@ -340,6 +344,22 @@ public boolean isAllowShell() { return allowShell; } + /** + * Whether to use the line-oriented console even when a full terminal is available. + * + *

The rich console positions the cursor: it pins a block to the bottom of the window and keeps + * the input line there while output scrolls above it. That needs a terminal that reports its size + * and understands the sequences, which is the normal case over SSH as well — but not in a plain + * pipe, a CI log, a `dumb` terminal, an editor's run window or a serial console, and not when the + * session is being recorded as text. This flag chooses the console that only ever appends lines, + * which is also what the agent falls back to on its own when there is no usable terminal. + * + * @return {@code true} when {@code --plain} was passed + */ + public boolean isPlain() { + return plain; + } + /** * Sampling temperature. * @@ -399,7 +419,8 @@ public String toString() { return "AgentOptions{baseUrl=" + baseUrl + ", modelPath=" + modelPath + ", gpuLayers=" + gpuLayers + ", ctxSize=" + ctxSize + ", logVerbosity=" + (verbose ? "verbose" : logVerbosity) + ", modelId=" + modelId + ", workspace=" + workspace - + ", allowShell=" + allowShell + ", auto=" + auto + ", autoCompact=" + autoCompact + ", temperature=" + + ", allowShell=" + allowShell + ", plain=" + plain + ", auto=" + auto + ", autoCompact=" + autoCompact + + ", temperature=" + temperature + ", maxTokens=" + maxTokens + ", maxToolRounds=" + maxToolRounds + ", prompt=" + (prompt == null ? "" : "") @@ -421,6 +442,7 @@ private static final class Builder { String modelId = DEFAULT_MODEL_ID; Path workspace = Paths.get("").toAbsolutePath().normalize(); boolean allowShell; + boolean plain; boolean auto; boolean autoCompact = DEFAULT_AUTO_COMPACT; int compactAt = DEFAULT_COMPACT_AT; diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java index da8e886b7..a48c8b7a9 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java @@ -196,7 +196,7 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream new AtomicReference<>(options.isAuto() ? ApprovalMode.AUTO : ApprovalMode.MANUAL); boolean interactive = options.getPrompt() == null && input != null; BufferedReader reader = input == null ? null : new BufferedReader(input); - terminal = interactive ? JLineTerminal.open(commandNames()) : null; + terminal = usesFullTerminal(options, interactive) ? JLineTerminal.open(commandNames()) : null; if (terminal == null) { terminal = new PlainTerminal(out, reader, Ansi.detect()); } @@ -442,6 +442,23 @@ private static void captureNativeLog(AgentTerminal terminal, AgentOptions option }); } + /** + * Whether to drive the cursor-controlling console rather than the line-oriented one. + * + *

Both consoles are kept, and this is the only place that decides between them. The rich one + * needs someone at a terminal and permission to position the cursor; {@code --plain} + * withholds the second even when the first is true, which is what a session that is piped, + * logged, recorded, or run through something that only forwards lines needs. A run without an + * interactive input has no use for it either way. + * + * @param options the parsed command line + * @param interactive whether there is someone typing + * @return {@code true} to try the full terminal + */ + static boolean usesFullTerminal(AgentOptions options, boolean interactive) { + return interactive && !options.isPlain(); + } + static ConsoleSession turn( AgentRunner runner, AgentFileSystem fileSystem, diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/AgentOptionsTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/AgentOptionsTest.java index 44117a05f..6d039ff4e 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/AgentOptionsTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/AgentOptionsTest.java @@ -17,6 +17,24 @@ class AgentOptionsTest { + @Test + void plainChoosesTheLineOrientedConsoleEvenWithATerminal() { + // Both consoles stay; this is the only switch between them. A run with no one typing never + // uses the rich one anyway, which is why the flag is not the whole answer. + AgentOptions rich = AgentOptions.parse(new String[] {"--base-url", "http://localhost:1/v1"}); + AgentOptions plain = AgentOptions.parse(new String[] {"--base-url", "http://localhost:1/v1", "--plain"}); + + assertThat(rich.isPlain(), is(false)); + assertThat(plain.isPlain(), is(true)); + assertThat(LocalAgent.usesFullTerminal(rich, true), is(true)); + assertThat("asked for plain, so not even with a terminal", LocalAgent.usesFullTerminal(plain, true), is(false)); + assertThat( + "nobody typing, so there is nothing to pin either way", + LocalAgent.usesFullTerminal(rich, false), + is(false)); + assertThat(LocalAgent.usesFullTerminal(plain, false), is(false)); + } + @Test void baseUrlModeWithDefaults() { AgentOptions options = AgentOptions.parse(new String[] {"--base-url", "http://127.0.0.1:8080/v1/"}); From 49b19e55b024d4b76fcfef3fc64cb6d6ea0656ee Mon Sep 17 00:00:00 2001 From: Bernard Ladenthin Date: Thu, 24 Sep 2026 22:10:52 +0200 Subject: [PATCH 32/36] llama-atmosphere-agent: record what was said, with timestamps, and /save it Nothing kept a session log before. The model's history has no timestamps and is rewritten by /compact -- a summary replaces the turns -- and ToolCallLog is a hard-cut receipt for tool calls only. So "what did I ask, what came back, when" had no answer. Transcript records user lines, answers, tool calls and session events, stamped, append-only. /save [name] writes it into the workspace; --transcript appends each entry as it is said so a killed session still leaves what it had, and a failed write is swallowed because a record meant to survive a bad ending must not cause one. /compact keeps it and notes that it happened; /clear empties it, since that command means forget this session. A list, not a map keyed by the timestamp: a tool result and the answer after it regularly land in the same millisecond, and a map would keep one and drop the other without saying so. Insertion order already is time order. Tested at both levels: the class (ordering, same-millisecond entries, blank lines dropped, live append, an unwritable path not ending the session) and end to end through the real server, where /save after a /compact still contains the answer and /save after a /clear does not contain the question. 148 tests green. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01E2h8gXyyE5UeimkL9vQv9G --- CLAUDE.md | 21 +- llama-atmosphere-agent/README.md | 20 ++ .../llama/atmosphere/AgentOptions.java | 21 +- .../llama/atmosphere/LocalAgent.java | 30 ++- .../llama/atmosphere/SlashCommands.java | 2 + .../llama/atmosphere/Transcript.java | 198 ++++++++++++++++++ .../net/ladenthin/llama/atmosphere/help.txt | 1 + .../llama/atmosphere/LocalAgentTest.java | 47 +++++ .../llama/atmosphere/TranscriptTest.java | 143 +++++++++++++ 9 files changed, 475 insertions(+), 8 deletions(-) create mode 100644 llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/Transcript.java create mode 100644 llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/TranscriptTest.java diff --git a/CLAUDE.md b/CLAUDE.md index 4c7645ee5..88ef7b2c9 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -2318,7 +2318,18 @@ are decisions, not details: an embedding index (Cursor's production effect is +0.3 %), and LSP tools (the one isolation study finds them token-negative and *worse* at multi-file rename, because renames touch comments and strings that semantic references exclude). -7. **Tool calls are carried into the conversation as a text note, and logged for `/calls`.** +7. **The session transcript (`Transcript`) is not the conversation the model is sent, and must not be + merged with it.** The model's history is rewritten by `/compact` — a summary replaces the turns — + and has never carried a timestamp; the transcript only grows and stamps every entry. `/compact` + adds a note to it and changes nothing else, `/clear` empties it (the command means "forget this + session"), `/save [name]` writes it into the workspace, and `--transcript ` appends live so a + killed session still leaves what it had — that write failing is swallowed, because a record that + exists to survive a bad ending may not cause one. **A list, not a map keyed by the timestamp**: a + tool result and the answer after it regularly share a millisecond and a map would drop one + silently; insertion order already is time order. `ToolCallLog` stays as the separate, hard-cut + receipt for `/calls` — it answers "did that really run", which prose cannot. + +8. **Tool calls are carried into the conversation as a text note, and logged for `/calls`.** `LocalAgent.withToolNotes` prefixes each turn's answer in the history with `(tools I actually ran this turn: -> )`, and `ToolCallLog` keeps the same data for the `/calls` command. **Why it is a note and not real `tool_calls` @@ -2345,7 +2356,7 @@ are decisions, not details: prompt then reads a key in raw mode on that worker thread while the console thread redraws four times a second, so `TurnActivity` pauses the redraw for as long as the question is open. Reading the pipe incrementally is not only cosmetic — an unread pipe blocks the child once it is full, which on Windows is roughly 4 KB. -8. **The context number in the status line is an estimate, marked `~`, and it moves during the turn.** +9. **The context number in the status line is an estimate, marked `~`, and it moves during the turn.** llama.cpp emits its usage chunk only when the client sets `stream_options.include_usage`, and Atmosphere's client does not; `ConsoleSession.usage()` takes the real count when one arrives, otherwise `LocalAgent.estimateTokens` uses four characters per token. The window size is @@ -2359,7 +2370,7 @@ are decisions, not details: request carried when it was sent, and yields to the server's own count as soon as one arrives. `TaskLoop` passes a constant function, and its step label must be copied into a local first: a lambda may not close over the loop counter. -9. **One call to `AgentTerminal.line` is one screen line**, and `ConsoleSessionTest` is what defends +10. **One call to `AgentTerminal.line` is one screen line**, and `ConsoleSessionTest` is what defends it. The pinned block is reserved in **lines**, so a single "line" carrying twenty newlines moves the screen twenty rows further than the terminal accounted for and the block is then drawn across the output — reported twice, both times from a `write_file` call whose `content` argument was the file. @@ -2368,7 +2379,7 @@ are decisions, not details: *name*; results and errors are folded the same way. `JLineTerminal.line` splits a multi-line string as a backstop for a caller that forgets. Only the console is cut — the model gets everything, and `ConsoleSession.rounds()` keeps the full arguments for the history note and `/calls`. -10. **The approval mode carries a glyph, and shift+tab switches it**: `ApprovalMode.symbol()` / +11. **The approval mode carries a glyph, and shift+tab switches it**: `ApprovalMode.symbol()` / `badge()` render `⏸ manual` and `⏵⏵ auto` on the status line and in `/mode`, the transport symbols the established terminal agents use for the same distinction; `ApprovalMode.next()` is the cycle the key walks. The binding is `AgentTerminal.onCycleMode(Runnable)`, which **defaults to declining** @@ -2381,7 +2392,7 @@ are decisions, not details: status numbers in an `AtomicLong`/`AtomicBoolean` rather than locals, so the widget (which runs inside the reader) can re-render the pinned row with what the last turn left behind. -11. **The input is framed into the pinned block, and the prompt stays there during a turn; typing +12. **The input is framed into the pinned block, and the prompt stays there during a turn; typing stops the turn.** The frame is two halves that must be read together: the **top** rule is the first line of the reader's *prompt* (`rule() + newline + "> "`, rebuilt on every read because the window can be resized) — **no, and the second attempt was wrong too.** Both are recorded because the diff --git a/llama-atmosphere-agent/README.md b/llama-atmosphere-agent/README.md index 3f98569f4..220ec9900 100644 --- a/llama-atmosphere-agent/README.md +++ b/llama-atmosphere-agent/README.md @@ -207,6 +207,26 @@ What the line-oriented console gives up, so the choice is an informed one: | history, Tab completion, Ctrl-L, Shift+Tab | yes | no | | correct when the output is a file | — | yes | +### The session transcript + +What was said is recorded separately from the conversation the model is sent, because those are two +different things: `/compact` **rewrites** the model's conversation (a summary replaces the turns it +summarises) and it never carried a timestamp at all. The transcript only grows. + +- `/save [name]` writes it into the workspace, one stamped line per entry: + `[2026-09-24 18:41:07] you: what does hello.txt say?`. Without a name the file is named after the + time. +- `--transcript ` appends every entry **as it is said**, so a session that is killed still + leaves what it had. A failure to write is swallowed on purpose: a record that exists to survive a + bad ending must not cause one. +- `/compact` keeps the record and notes that it happened. `/clear` empties it, because that command + means "forget this session" and leaving the text behind would make that untrue. +- `/calls` is the shorter, tool-only receipt and is unchanged. + +It is a list, not a map keyed by the timestamp: a tool result and the answer that follows it regularly +land in the same millisecond, and a map would keep one and drop the other silently. Insertion order +already is time order. + ### Commands, approval and the status line A line that starts with `/` and names a command is answered by the agent itself; anything else — an diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentOptions.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentOptions.java index 35028c5cd..e90afdff8 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentOptions.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentOptions.java @@ -69,6 +69,7 @@ public final class AgentOptions { private final Path workspace; private final boolean allowShell; private final boolean plain; + private final java.nio.file.@org.jspecify.annotations.Nullable Path transcript; private final boolean auto; private final boolean autoCompact; private final int compactAt; @@ -91,6 +92,7 @@ private AgentOptions(Builder b) { this.workspace = b.workspace; this.allowShell = b.allowShell; this.plain = b.plain; + this.transcript = b.transcript; this.auto = b.auto; this.autoCompact = b.autoCompact; this.compactAt = b.compactAt; @@ -118,6 +120,7 @@ public static AgentOptions parse(String[] args) { case "-h", "--help" -> b.help = true; case "--allow-shell" -> b.allowShell = true; case "--plain" -> b.plain = true; + case "--transcript" -> b.transcript = java.nio.file.Path.of(value(args, ++i, a)); case "--auto" -> b.auto = true; case "--auto-compact" -> b.autoCompact = booleanValue(args, ++i, a); case "--compact-at" -> b.compactAt = percentValue(args, ++i, a); @@ -212,6 +215,7 @@ public static String usage() { " --workspace

directory the file tools are confined to (default: cwd)", " --allow-shell add the run_command tool (runs any command line, starting in the workspace)", " --plain line-oriented console: no pinned block, no cursor control", + " --transcript append what is said, with timestamps, as it happens", " --auto run tools without asking (default: ask before writes and commands)", " --auto-compact summarize the history before it overflows the context (default " + DEFAULT_AUTO_COMPACT + ")", @@ -360,6 +364,19 @@ public boolean isPlain() { return plain; } + /** + * Where to append the session transcript as it happens, if anywhere. + * + *

{@code /save} writes the whole thing on request; this writes each line as it is said, so a + * session that is killed still leaves what it had. A failure to write is swallowed: a record that + * exists to survive a bad ending must not cause one. + * + * @return the file, or {@code null} when the transcript is kept in memory only + */ + public java.nio.file.@org.jspecify.annotations.Nullable Path getTranscript() { + return transcript; + } + /** * Sampling temperature. * @@ -419,7 +436,8 @@ public String toString() { return "AgentOptions{baseUrl=" + baseUrl + ", modelPath=" + modelPath + ", gpuLayers=" + gpuLayers + ", ctxSize=" + ctxSize + ", logVerbosity=" + (verbose ? "verbose" : logVerbosity) + ", modelId=" + modelId + ", workspace=" + workspace - + ", allowShell=" + allowShell + ", plain=" + plain + ", auto=" + auto + ", autoCompact=" + autoCompact + + ", allowShell=" + allowShell + ", plain=" + plain + ", transcript=" + transcript + ", auto=" + auto + + ", autoCompact=" + autoCompact + ", temperature=" + temperature + ", maxTokens=" + maxTokens @@ -443,6 +461,7 @@ private static final class Builder { Path workspace = Paths.get("").toAbsolutePath().normalize(); boolean allowShell; boolean plain; + java.nio.file.@org.jspecify.annotations.Nullable Path transcript; boolean auto; boolean autoCompact = DEFAULT_AUTO_COMPACT; int compactAt = DEFAULT_COMPACT_AT; diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java index a48c8b7a9..794526537 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java @@ -192,6 +192,9 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream List history = new ArrayList<>(); ToolCallLog callLog = new ToolCallLog(); + // What was said, with the time, kept apart from the conversation the model is sent: that + // one is rewritten by /compact and has no timestamps at all. + Transcript transcript = new Transcript(options.getTranscript()); AtomicReference mode = new AtomicReference<>(options.isAuto() ? ApprovalMode.AUTO : ApprovalMode.MANUAL); boolean interactive = options.getPrompt() == null && input != null; @@ -298,9 +301,11 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream inputTokens.get(), estimated.get(), terminal, - callLog)); + callLog, + transcript)); continue; } + transcript.add(Transcript.Kind.USER, line); turnNumber++; // Compact BEFORE the request that would overflow, not after it: afterwards the // oversized request has already been sent, which is the one thing to avoid. @@ -332,6 +337,11 @@ options, contextSize, estimateTokens(systemPrompt(options), history) + line.leng options.getModelId(), options.getModelPath() == null), activity); + for (ConsoleSession.ToolRound round : completed.rounds()) { + transcript.add( + Transcript.Kind.TOOL, round.name() + " " + round.argumentsJson() + " -> " + round.result()); + } + transcript.add(Transcript.Kind.AGENT, completed.text()); pendingNote = toolNote(completed.rounds()); estimated.set(completed.inputTokens() == 0); inputTokens.set( @@ -559,6 +569,7 @@ static String toolNote(List rounds) { * @param estimated whether that number is an estimate * @param terminal the console * @param callLog every tool call of the session, for {@code /calls} + * @param transcript what was said, with the time, for {@code /save} * @return the input tokens to show from now on (unchanged, or the summary's after {@code /compact}) * @throws InterruptedException if interrupted while a summary is generated */ @@ -573,18 +584,30 @@ private static long handleCommand( long inputTokens, boolean estimated, AgentTerminal terminal, - ToolCallLog callLog) + ToolCallLog callLog, + Transcript transcript) throws InterruptedException { switch (command.command()) { case HELP -> prompt(HELP_TEXT).lines().forEach(terminal::line); case CLEAR -> { history.clear(); + // "Forget this session" has to mean the record too, or the word is not true. + transcript.clear(); // The screen goes with it: what is still on it is a conversation the model no longer // has, which reads as if it were still in play. terminal.clearScreen(); terminal.line("(history cleared)"); } case CLS -> terminal.clearScreen(); + case SAVE -> { + try { + java.nio.file.Path written = transcript.save( + options.getWorkspace(), command.hasArguments() ? command.arguments() : null); + terminal.line("transcript: " + transcript.size() + " entries -> " + written); + } catch (java.io.IOException e) { + terminal.line("could not write the transcript: " + e.getMessage()); + } + } case CALLS -> callLog.render().lines().forEach(terminal::line); case TOOLS -> { terminal.line("tools: " + String.join(", ", runner.toolNames())); @@ -616,6 +639,9 @@ private static long handleCommand( terminal.line("history: " + history.size() + " messages"); } case COMPACT -> { + // The record is deliberately untouched: compacting rewrites what the model is sent, + // not what happened. Only the fact that it happened is worth a line. + transcript.add(Transcript.Kind.NOTE, "compacted the conversation"); return compact(runner, fileSystem, history, command.arguments(), terminal); } case LOOP -> loop(runner, fileSystem, terminal, options, mode, command.arguments(), callLog); diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/SlashCommands.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/SlashCommands.java index 18f653705..7d976f0cf 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/SlashCommands.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/SlashCommands.java @@ -32,6 +32,8 @@ public enum Command { CLEAR("/clear", "/reset", "/new"), /** Wipe the screen, keeping the conversation. */ CLS("/cls", "/clear-screen"), + /** Write what was said, with the time, to a file in the workspace. */ + SAVE("/save", "/transcript"), /** Summarize the history and continue with the summary; the argument steers the summary. */ COMPACT("/compact"), /** Keep working on one task until it is done; see {@link TaskLoop}. */ diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/Transcript.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/Transcript.java new file mode 100644 index 000000000..e27484a3a --- /dev/null +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/Transcript.java @@ -0,0 +1,198 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import java.io.IOException; +import java.nio.charset.StandardCharsets; +import java.nio.file.Files; +import java.nio.file.Path; +import java.nio.file.StandardOpenOption; +import java.time.LocalDateTime; +import java.time.format.DateTimeFormatter; +import java.util.List; +import java.util.concurrent.CopyOnWriteArrayList; +import org.jspecify.annotations.Nullable; + +/** + * What was said in this session, in the order it was said, with the time. + * + *

It is not the conversation the model is sent. That one is rewritten by {@code /compact} — a + * summary replaces the turns it summarises — and it never carried a time at all, so it cannot answer + * "what did I ask before lunch" or "what did it actually reply". This record only ever grows. + * {@code /compact} adds a note to it and changes nothing else; {@code /clear} empties it, because + * that command means "forget this session" and leaving the text behind would make that untrue. + * + *

A list, not a map keyed by the time. Two entries can share a millisecond — a tool result + * and the answer that follows it regularly do — and a map would keep one of them and silently drop + * the other. Insertion order already is time order, which is the only ordering anyone wants here. + * + *

With a file configured, every entry is also appended as it happens, so a session that is killed + * still leaves what it had. Without one, {@code /save} writes the whole thing on request. + */ +public final class Transcript { + + private static final DateTimeFormatter STAMP = DateTimeFormatter.ofPattern("yyyy-MM-dd HH:mm:ss"); + + private static final DateTimeFormatter FILE_STAMP = DateTimeFormatter.ofPattern("yyyy-MM-dd_HH-mm-ss"); + + /** Who said it. */ + public enum Kind { + /** What the user typed. */ + USER("you"), + /** What the model answered. */ + AGENT("agent"), + /** A tool call and what it returned. */ + TOOL("tool"), + /** Something the session did: compacted, interrupted, mode changed. */ + NOTE("note"); + + private final String label; + + Kind(String label) { + this.label = label; + } + + /** + * The word written in the file. + * + * @return the label + */ + public String label() { + return label; + } + } + + /** + * One thing that was said. + * + * @param at when + * @param kind who + * @param text what, verbatim and uncut + */ + public record Entry(LocalDateTime at, Kind kind, String text) {} + + private final List entries = new CopyOnWriteArrayList<>(); + private final @Nullable Path liveFile; + + /** + * Keep a session transcript in memory only. + */ + public Transcript() { + this(null); + } + + /** + * Keep a session transcript, and append every entry to a file as it happens. + * + * @param liveFile the file to append to, or {@code null} to keep it in memory + */ + public Transcript(@Nullable Path liveFile) { + this.liveFile = liveFile; + } + + /** + * Record something. + * + *

Blank text is dropped: an empty answer is not worth a line, and the file would fill with + * them on a session of interrupted turns. + * + * @param kind who said it + * @param text what was said + */ + public void add(Kind kind, String text) { + if (text == null || text.isBlank()) { + return; + } + Entry entry = new Entry(LocalDateTime.now(), kind, text.strip()); + entries.add(entry); + appendLive(entry); + } + + /** + * Everything recorded, oldest first. + * + * @return the entries + */ + public List entries() { + return List.copyOf(entries); + } + + /** + * How many entries were recorded. + * + * @return the count + */ + public int size() { + return entries.size(); + } + + /** Forget the session, as {@code /clear} means it. */ + public void clear() { + entries.clear(); + } + + /** + * The whole transcript as text. + * + * @return one block per entry, oldest first + */ + public String render() { + StringBuilder text = new StringBuilder(); + for (Entry entry : entries) { + text.append(format(entry)); + } + return text.toString(); + } + + /** + * Write the transcript to a file. + * + * @param directory where it goes, normally the workspace + * @param name the file name, or {@code null} for one named after the time it was written + * @return the file that was written + * @throws IOException if it cannot be written + */ + public Path save(Path directory, @Nullable String name) throws IOException { + String fileName = name == null || name.isBlank() + ? "transcript-" + LocalDateTime.now().format(FILE_STAMP) + ".txt" + : name.strip(); + Path file = directory.resolve(fileName); + Files.createDirectories(file.toAbsolutePath().getParent()); + Files.writeString(file, render(), StandardCharsets.UTF_8); + return file; + } + + private void appendLive(Entry entry) { + if (liveFile == null) { + return; + } + try { + Path parent = liveFile.toAbsolutePath().getParent(); + if (parent != null) { + Files.createDirectories(parent); + } + Files.writeString( + liveFile, + format(entry), + StandardCharsets.UTF_8, + StandardOpenOption.CREATE, + StandardOpenOption.APPEND); + } catch (IOException e) { + // A transcript that cannot be written must not end the session: the point of it is to + // survive a bad ending, so it may not cause one. + } + } + + /** + * One entry as it appears in the file. + * + * @param entry the entry + * @return the block, ending in a newline + */ + private static String format(Entry entry) { + String separator = System.lineSeparator(); + return "[" + entry.at().format(STAMP) + "] " + entry.kind().label() + ": " + entry.text() + separator; + } +} diff --git a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt index 12c28c98d..8e308418f 100644 --- a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt +++ b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt @@ -14,6 +14,7 @@ Commands (everything else is sent to the model): the state lives in AGENT-LOOP.md, not in the conversation /clear drop the history, and wipe the screen with it (/reset, /new) /cls wipe the screen, keep the conversation (/clear-screen, or Ctrl-L) + /save [name] write what was said, with timestamps, into the workspace (/transcript) /exit leave (/quit) The input line sits in the box at the bottom and is there while the agent works: you can type at diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java index ee5efdfe6..3344a08a5 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java @@ -128,6 +128,53 @@ void interactiveModeRunsTurnsUntilExitAndKeepsHistory() throws Exception { assertThat(out.toString(StandardCharsets.UTF_8), containsString("answer 2")); } + @Test + void theSessionIsRecordedAndSavedWhereTheToolsWork() throws Exception { + // End to end through the real server: what was typed, what came back, written by /save with a + // timestamp on every line. /compact keeps the record -- it rewrites what the model is sent, + // not what happened -- and only /clear empties it. + ScriptedBackend backend = new ScriptedBackend((call, request) -> ScriptedBackend.textTurn("answer " + call)); + try (OpenAiCompatServer server = server(backend)) { + AgentOptions options = AgentOptions.parse(new String[] { + "--base-url", "http://127.0.0.1:" + server.getPort() + "/v1", "--workspace", workspace.toString() + }); + + int exit = LocalAgent.run( + options, + new StringReader("what is two plus two\n/compact\n/save session.txt\n/exit\n"), + new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8), + new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8)); + + assertThat(exit, is(0)); + java.nio.file.Path written = workspace.resolve("session.txt"); + assertThat("the file lands in the workspace", java.nio.file.Files.exists(written), is(true)); + String text = java.nio.file.Files.readString(written, StandardCharsets.UTF_8); + assertThat(text, containsString("you: what is two plus two")); + assertThat("the answer survived the compaction", text, containsString("agent: answer 1")); + assertThat("and the compaction is noted rather than hidden", text, containsString("compacted")); + assertThat("every line is stamped", text.startsWith("["), is(true)); + } + } + + @Test + void clearingEmptiesTheRecordAsWell() throws Exception { + ScriptedBackend backend = new ScriptedBackend((call, request) -> ScriptedBackend.textTurn("answer " + call)); + try (OpenAiCompatServer server = server(backend)) { + AgentOptions options = AgentOptions.parse(new String[] { + "--base-url", "http://127.0.0.1:" + server.getPort() + "/v1", "--workspace", workspace.toString() + }); + + LocalAgent.run( + options, + new StringReader("remember this\n/clear\n/save after-clear.txt\n/exit\n"), + new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8), + new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8)); + + String text = java.nio.file.Files.readString(workspace.resolve("after-clear.txt"), StandardCharsets.UTF_8); + assertThat("forget the session means the record too", text, not(containsString("remember this"))); + } + } + @Test void failedTurnExitsNonZero() throws Exception { ScriptedBackend backend = new ScriptedBackend((call, request) -> ScriptedBackend.textTurn("never")); diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/TranscriptTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/TranscriptTest.java new file mode 100644 index 000000000..82a8d79ad --- /dev/null +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/TranscriptTest.java @@ -0,0 +1,143 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import static org.hamcrest.MatcherAssert.assertThat; +import static org.hamcrest.Matchers.contains; +import static org.hamcrest.Matchers.containsString; +import static org.hamcrest.Matchers.is; +import static org.hamcrest.Matchers.not; + +import java.nio.charset.StandardCharsets; +import java.nio.file.Files; +import java.nio.file.Path; +import java.util.List; +import org.junit.jupiter.api.Test; +import org.junit.jupiter.api.io.TempDir; + +/** + * The record of what was said, which is a different thing from the conversation the model is sent. + */ +class TranscriptTest { + + @TempDir + Path directory; + + @Test + void entriesComeBackInTheOrderTheyWereSaid() { + Transcript transcript = new Transcript(); + + transcript.add(Transcript.Kind.USER, "first"); + transcript.add(Transcript.Kind.TOOL, "ls {} -> a b"); + transcript.add(Transcript.Kind.AGENT, "second"); + + assertThat( + transcript.entries().stream().map(Transcript.Entry::text).toList(), + contains("first", "ls {} -> a b", "second")); + } + + @Test + void twoEntriesInTheSameMomentAreBothKept() { + // The reason this is a list and not a map keyed by the time: a tool result and the answer that + // follows it regularly land in the same millisecond, and a map would keep one and lose the + // other without saying so. + Transcript transcript = new Transcript(); + + for (int i = 0; i < 50; i++) { + transcript.add(Transcript.Kind.AGENT, "entry " + i); + } + + assertThat(transcript.size(), is(50)); + } + + @Test + void blankTextIsNotWorthALine() { + Transcript transcript = new Transcript(); + + transcript.add(Transcript.Kind.AGENT, ""); + transcript.add(Transcript.Kind.AGENT, " "); + transcript.add(Transcript.Kind.AGENT, null); + + assertThat("a session of interrupted turns would otherwise fill the file", transcript.size(), is(0)); + } + + @Test + void everyLineCarriesTheTimeAndWhoSaidIt() { + Transcript transcript = new Transcript(); + transcript.add(Transcript.Kind.USER, "what time is it"); + + String rendered = transcript.render(); + + assertThat(rendered, containsString("you: what time is it")); + assertThat( + "a date and a clock time", + rendered.matches("(?s)\\[\\d{4}-\\d{2}-\\d{2} \\d{2}:\\d{2}:\\d{2}\\].*"), + is(true)); + } + + @Test + void savingWritesEverythingToTheNamedFile() throws Exception { + Transcript transcript = new Transcript(); + transcript.add(Transcript.Kind.USER, "hello"); + transcript.add(Transcript.Kind.AGENT, "hi"); + + Path written = transcript.save(directory, "session.txt"); + + assertThat(written.getFileName().toString(), is("session.txt")); + String text = Files.readString(written, StandardCharsets.UTF_8); + assertThat(text, containsString("you: hello")); + assertThat(text, containsString("agent: hi")); + } + + @Test + void savingWithoutANameUsesTheTime() throws Exception { + Transcript transcript = new Transcript(); + transcript.add(Transcript.Kind.USER, "hello"); + + Path written = transcript.save(directory, null); + + assertThat(written.getFileName().toString(), containsString("transcript-")); + assertThat(written.getFileName().toString().endsWith(".txt"), is(true)); + } + + @Test + void aConfiguredFileIsAppendedToAsThingsAreSaid() throws Exception { + // The point of it: a session that is killed still leaves what it had. Nothing is written at + // the end, so there is no end to miss. + Path live = directory.resolve("logs").resolve("live.txt"); + Transcript transcript = new Transcript(live); + + transcript.add(Transcript.Kind.USER, "one"); + assertThat("written already, not at the end", Files.readString(live), containsString("one")); + + transcript.add(Transcript.Kind.AGENT, "two"); + + List lines = Files.readAllLines(live, StandardCharsets.UTF_8); + assertThat(lines.size(), is(2)); + assertThat(lines.get(0), containsString("you: one")); + assertThat(lines.get(1), containsString("agent: two")); + } + + @Test + void aFileThatCannotBeWrittenDoesNotEndTheSession() { + // It exists to survive a bad ending, so it may not cause one. + Transcript transcript = new Transcript(directory); + + transcript.add(Transcript.Kind.USER, "the path is a directory, not a file"); + + assertThat("kept in memory regardless", transcript.size(), is(1)); + } + + @Test + void clearingForgetsTheSession() { + Transcript transcript = new Transcript(); + transcript.add(Transcript.Kind.USER, "hello"); + + transcript.clear(); + + assertThat(transcript.size(), is(0)); + assertThat(transcript.render(), not(containsString("hello"))); + } +} From 91210b67dadc659f6e2800ca23931d2f25bfa60e Mon Sep 17 00:00:00 2001 From: Bernard Ladenthin Date: Thu, 24 Sep 2026 22:15:57 +0200 Subject: [PATCH 33/36] llama-atmosphere-agent: /retry and --system-file /retry sends the last question again and removes the answer that came back from the conversation first -- leaving it would show the model what it said last time, and a model that sees its own answer repeats it, which is the opposite of a retry. The question goes with it because the turn adds it back. Only a trailing exchange that really is the one being retried is touched, so a history that was just replaced by a summary, or one that never got an answer, is left alone. Before anything has been asked it says so instead of sending an empty turn. --system-file replaces the system prompt from a file, which is the same as --system without fighting the shell over quoting and newlines. Read at startup, so a wrong path is a usage error immediately rather than a surprise on the first turn -- a prompt long enough to be worth a file is long enough that its absence should not be discovered mid-session. Four tests: the retry really re-asks with the answer gone, an empty retry is refused without sending anything, the file's content arrives as the system message, and a missing file is rejected at parse time. 152 tests green. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01E2h8gXyyE5UeimkL9vQv9G --- llama-atmosphere-agent/README.md | 11 +++ .../llama/atmosphere/AgentOptions.java | 21 +++++ .../llama/atmosphere/LocalAgent.java | 72 ++++++++++++---- .../llama/atmosphere/SlashCommands.java | 2 + .../net/ladenthin/llama/atmosphere/help.txt | 1 + .../llama/atmosphere/LocalAgentTest.java | 85 +++++++++++++++++++ 6 files changed, 177 insertions(+), 15 deletions(-) diff --git a/llama-atmosphere-agent/README.md b/llama-atmosphere-agent/README.md index 220ec9900..dad72e0a9 100644 --- a/llama-atmosphere-agent/README.md +++ b/llama-atmosphere-agent/README.md @@ -207,6 +207,17 @@ What the line-oriented console gives up, so the choice is an informed one: | history, Tab completion, Ctrl-L, Shift+Tab | yes | no | | correct when the output is a file | — | yes | +### Asking again, and your own system prompt + +`/retry` (`/again`) sends the last question once more and **drops the answer that came back** from the +conversation first. That is the whole point: leaving it in place would show the model what it said +last time, and a model that sees its own answer repeats it. Before anything has been asked it says so +rather than sending an empty turn. + +`--system-file ` replaces the system prompt with the content of a file — the same as `--system` +but without fighting the shell over quoting and newlines. It is read at startup, so a wrong path is a +usage error immediately instead of a surprise on the first turn. + ### The session transcript What was said is recorded separately from the conversation the model is sent, because those are two diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentOptions.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentOptions.java index e90afdff8..018401895 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentOptions.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentOptions.java @@ -139,6 +139,7 @@ public static AgentOptions parse(String[] args) { case "--max-tokens" -> b.maxTokens = intValue(args, ++i, a); case "--max-tool-rounds" -> b.maxToolRounds = intValue(args, ++i, a); case "--system" -> b.systemPrompt = value(args, ++i, a); + case "--system-file" -> b.systemPrompt = readSystemPrompt(value(args, ++i, a)); case "--prompt", "-p" -> b.prompt = value(args, ++i, a); default -> throw new IllegalArgumentException("Unknown argument: " + a); } @@ -214,6 +215,7 @@ public static String usage() { "Agent:", " --workspace

directory the file tools are confined to (default: cwd)", " --allow-shell add the run_command tool (runs any command line, starting in the workspace)", + " --system-file replace the system prompt with the content of a file", " --plain line-oriented console: no pinned block, no cursor control", " --transcript append what is said, with timestamps, as it happens", " --auto run tools without asking (default: ask before writes and commands)", @@ -413,6 +415,25 @@ public int getMaxToolRounds() { return systemPrompt; } + /** + * Read a system prompt from a file. + * + *

Read here rather than when it is used, so a path that does not exist is a usage error at + * startup instead of a surprise on the first turn. A prompt long enough to be worth a file is also + * long enough that a typo in the path is easy to miss. + * + * @param path the file + * @return its content + * @throws IllegalArgumentException when it cannot be read + */ + private static String readSystemPrompt(String path) { + try { + return java.nio.file.Files.readString(java.nio.file.Path.of(path), java.nio.charset.StandardCharsets.UTF_8); + } catch (java.io.IOException | RuntimeException e) { + throw new IllegalArgumentException("--system-file cannot be read: " + path + " (" + e.getMessage() + ")"); + } + } + /** * One-shot prompt. * diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java index 794526537..3832b732c 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java @@ -260,6 +260,7 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream + (shortcut ? " shift+tab switches the approval mode." : "")); int turnNumber = 0; String pendingNote = ""; + String lastMessage = ""; while (true) { // Pinned to the bottom of the window on a real terminal; printed above the prompt on a // plain stream, where there is nothing to pin and a repeatedly refreshed line would @@ -290,22 +291,37 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream if (command.get().command() == SlashCommands.Command.EXIT) { return 0; } - inputTokens.set(handleCommand( - command.get(), - runner, - fileSystem, - history, - mode, - options, - contextSize, - inputTokens.get(), - estimated.get(), - terminal, - callLog, - transcript)); - continue; + if (command.get().command() == SlashCommands.Command.RETRY) { + // Handled here rather than with the other commands, because it is not a command + // that answers something: it runs a turn, which only this loop can do. + if (lastMessage.isEmpty()) { + terminal.line("nothing to retry yet"); + continue; + } + dropLastExchange(history, lastMessage); + transcript.add(Transcript.Kind.NOTE, "retrying: " + lastMessage); + line = lastMessage; + } else { + inputTokens.set(handleCommand( + command.get(), + runner, + fileSystem, + history, + mode, + options, + contextSize, + inputTokens.get(), + estimated.get(), + terminal, + callLog, + transcript)); + continue; + } } - transcript.add(Transcript.Kind.USER, line); + if (!line.equals(lastMessage)) { + transcript.add(Transcript.Kind.USER, line); + } + lastMessage = line; turnNumber++; // Compact BEFORE the request that would overflow, not after it: afterwards the // oversized request has already been sent, which is the one thing to avoid. @@ -465,6 +481,32 @@ private static void captureNativeLog(AgentTerminal terminal, AgentOptions option * @param interactive whether there is someone typing * @return {@code true} to try the full terminal */ + /** + * Take the last exchange out of the conversation, so a retry asks again instead of following on. + * + *

Leaving the failed answer in place would be the opposite of a retry: the model would see what + * it said last time and, being a model, would say it again. The question is removed with it, + * because the turn that follows adds it back. + * + *

Only a trailing exchange that really is the one being retried is touched — a history that was + * just replaced by a summary, or one that never got an answer, is left alone. + * + * @param history the conversation, modified in place + * @param message the question being asked again + */ + static void dropLastExchange(List history, String message) { + if (!history.isEmpty() + && "assistant".equals(history.get(history.size() - 1).role())) { + history.remove(history.size() - 1); + } + if (!history.isEmpty()) { + ChatMessage last = history.get(history.size() - 1); + if ("user".equals(last.role()) && message.equals(last.content())) { + history.remove(history.size() - 1); + } + } + } + static boolean usesFullTerminal(AgentOptions options, boolean interactive) { return interactive && !options.isPlain(); } diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/SlashCommands.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/SlashCommands.java index 7d976f0cf..6b0c56531 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/SlashCommands.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/SlashCommands.java @@ -34,6 +34,8 @@ public enum Command { CLS("/cls", "/clear-screen"), /** Write what was said, with the time, to a file in the workspace. */ SAVE("/save", "/transcript"), + /** Ask the last question again, without the answer that came back. */ + RETRY("/retry", "/again"), /** Summarize the history and continue with the summary; the argument steers the summary. */ COMPACT("/compact"), /** Keep working on one task until it is done; see {@link TaskLoop}. */ diff --git a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt index 8e308418f..a0a0103d1 100644 --- a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt +++ b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt @@ -15,6 +15,7 @@ Commands (everything else is sent to the model): /clear drop the history, and wipe the screen with it (/reset, /new) /cls wipe the screen, keep the conversation (/clear-screen, or Ctrl-L) /save [name] write what was said, with timestamps, into the workspace (/transcript) + /retry ask the last question again, dropping the answer (/again) /exit leave (/quit) The input line sits in the box at the bottom and is there while the agent works: you can type at diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java index 3344a08a5..07ce227f2 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java @@ -175,6 +175,91 @@ void clearingEmptiesTheRecordAsWell() throws Exception { } } + @Test + void retryAsksAgainWithoutTheAnswerThatCameBack() throws Exception { + // The point of a retry: the model must not see what it said last time, or it says it again. + ScriptedBackend backend = new ScriptedBackend((call, request) -> ScriptedBackend.textTurn("answer " + call)); + try (OpenAiCompatServer server = server(backend)) { + AgentOptions options = AgentOptions.parse(new String[] { + "--base-url", "http://127.0.0.1:" + server.getPort() + "/v1", "--workspace", workspace.toString() + }); + + LocalAgent.run( + options, + new StringReader("the question\n/retry\n/exit\n"), + new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8), + new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8)); + } + List requests = backend.requests(); + assertThat("it was asked twice", requests, hasSize(2)); + JsonNode second = requests.get(1).path("messages"); + assertThat("system and the question, and nothing else", second.size(), is(2)); + assertThat(second.get(1).path("content").asText(), is("the question")); + assertThat( + "the first answer is gone from the conversation", second.toString(), not(containsString("answer 1"))); + } + + @Test + void retryBeforeAnythingWasAskedSaysSo() throws Exception { + ScriptedBackend backend = new ScriptedBackend((call, request) -> ScriptedBackend.textTurn("never")); + ByteArrayOutputStream out = new ByteArrayOutputStream(); + try (OpenAiCompatServer server = server(backend)) { + AgentOptions options = AgentOptions.parse(new String[] { + "--base-url", "http://127.0.0.1:" + server.getPort() + "/v1", "--workspace", workspace.toString() + }); + + LocalAgent.run( + options, + new StringReader("/retry\n/exit\n"), + new PrintStream(out, true, StandardCharsets.UTF_8), + new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8)); + } + assertThat(out.toString(StandardCharsets.UTF_8), containsString("nothing to retry")); + assertThat("and nothing was sent", backend.requests(), hasSize(0)); + } + + @Test + void aSystemPromptCanComeFromAFile() throws Exception { + java.nio.file.Path promptFile = workspace.resolve("persona.txt"); + java.nio.file.Files.writeString(promptFile, "You answer only in haiku."); + ScriptedBackend backend = new ScriptedBackend((call, request) -> ScriptedBackend.textTurn("ok")); + try (OpenAiCompatServer server = server(backend)) { + AgentOptions options = AgentOptions.parse(new String[] { + "--base-url", + "http://127.0.0.1:" + server.getPort() + "/v1", + "--workspace", + workspace.toString(), + "--system-file", + promptFile.toString() + }); + + LocalAgent.run( + options, + new StringReader("hello\n/exit\n"), + new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8), + new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8)); + } + JsonNode system = backend.requests().get(0).path("messages").get(0); + assertThat(system.path("role").asText(), is("system")); + assertThat(system.path("content").asText(), containsString("only in haiku")); + } + + @Test + void aSystemFileThatIsNotThereIsAUsageError() { + // Read at startup rather than at first use: a typo in a path is easy to miss, and a prompt long + // enough to be worth a file is long enough that its absence should not be a surprise mid-turn. + IllegalArgumentException thrown = org.junit.jupiter.api.Assertions.assertThrows( + IllegalArgumentException.class, + () -> AgentOptions.parse(new String[] { + "--base-url", + "http://localhost:1/v1", + "--system-file", + workspace.resolve("gone.txt").toString() + })); + + assertThat(thrown.getMessage(), containsString("--system-file cannot be read")); + } + @Test void failedTurnExitsNonZero() throws Exception { ScriptedBackend backend = new ScriptedBackend((call, request) -> ScriptedBackend.textTurn("never")); From 2c1b7297ed20d9f2a30e8e2d5ca568e459b04430 Mon Sep 17 00:00:00 2001 From: Bernard Ladenthin Date: Thu, 24 Sep 2026 22:18:43 +0200 Subject: [PATCH 34/36] llama-atmosphere-agent: multi-line input between """ fences A console sends on Enter, so a stack trace or a function pasted into the prompt became several questions. A line that is exactly """ opens a block and the next one closes it; everything between is one message with its newlines. Chosen over a key combination because it works in both consoles, survives a paste (the fences arrive as part of the pasted text) and asks nothing of the terminal. End of input inside an unfinished block ends the session rather than sending half a thought. Also stabilised a test that had started failing intermittently: the clear-screen one cleared while the reader thread was starting the next prompt, which made what JLine emits depend on which got there first. It waits for the reader to settle now -- that test is about the block coming back, not about that race. Run three times to confirm. 154 tests green. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01E2h8gXyyE5UeimkL9vQv9G --- llama-atmosphere-agent/README.md | 18 +++++++ .../llama/atmosphere/LocalAgent.java | 43 +++++++++++++++ .../llama/atmosphere/JLineTerminalTest.java | 6 ++- .../llama/atmosphere/LocalAgentTest.java | 52 +++++++++++++++++++ 4 files changed, 118 insertions(+), 1 deletion(-) diff --git a/llama-atmosphere-agent/README.md b/llama-atmosphere-agent/README.md index dad72e0a9..2ed8012c6 100644 --- a/llama-atmosphere-agent/README.md +++ b/llama-atmosphere-agent/README.md @@ -218,6 +218,24 @@ rather than sending an empty turn. but without fighting the shell over quoting and newlines. It is read at startup, so a wrong path is a usage error immediately instead of a surprise on the first turn. +### Multi-line input + +A console sends on Enter, so pasting a stack trace or a function into the prompt would turn it into +several questions. A line that is exactly `"""` opens a block and the next one closes it; everything +between is **one** message, newlines and all: + +``` +you> """ +... public int add(int a, int b) { +... return a - b; +... } +... """ +``` + +Chosen over a key combination because it works in both consoles, survives a paste — the fences arrive +as part of the pasted text — and asks nothing of the terminal. End of input inside an unfinished block +ends the session without sending it. + ### The session transcript What was said is recorded separately from the conversation the model is sent, because those are two diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java index 3832b732c..fc2f9f67a 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java @@ -98,6 +98,10 @@ public final class LocalAgent { static final String SPINNER_WORDS = "spinner-words.txt"; private static final Duration SHELL_TIMEOUT = Duration.ofSeconds(120); + + /** A line that is exactly this opens and closes a multi-line message. */ + static final String BLOCK_FENCE = "\"\"\""; + private static final int SHELL_MAX_OUTPUT_CHARS = 20_000; private LocalAgent() {} @@ -283,6 +287,12 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream if (line == null) { return 0; } + if (BLOCK_FENCE.equals(line.strip())) { + line = readBlock(terminal); + if (line == null) { + return 0; + } + } if (line.trim().isEmpty()) { continue; } @@ -507,6 +517,39 @@ static void dropLastExchange(List history, String message) { } } + /** + * Read a block of lines, the way a fenced code block is written. + * + *

A console reads a line at a time, and Enter sends it — which makes pasting a stack trace or a + * function into the prompt impossible without it becoming several questions. A line that is + * exactly {@value #BLOCK_FENCE} starts a block and the next one closes it; everything between is + * one message, newlines and all. + * + *

Chosen over a key combination because it works in both consoles, survives a paste (the fence + * arrives as part of the pasted text), and needs nothing from the terminal. + * + * @param terminal where the lines come from + * @return the block, or {@code null} at end of input + */ + static @Nullable String readBlock(AgentTerminal terminal) { + StringBuilder block = new StringBuilder(); + while (true) { + String line = terminal.readLine("... "); + if (line == null) { + // End of input inside a block: what was collected is still a question worth asking, + // but there is nobody left to answer it, so the session ends as it would anyway. + return null; + } + if (BLOCK_FENCE.equals(line.strip())) { + return block.toString(); + } + if (block.length() > 0) { + block.append(System.lineSeparator()); + } + block.append(line); + } + } + static boolean usesFullTerminal(AgentOptions options, boolean interactive) { return interactive && !options.isPlain(); } diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/JLineTerminalTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/JLineTerminalTest.java index 1f7332f53..be58cc79c 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/JLineTerminalTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/JLineTerminalTest.java @@ -262,8 +262,12 @@ void theBlockIsBackOnScreenAfterAClear() throws Exception { // nothing, leaving the bottom of the window empty. try (Terminal terminal = terminal("go\n"); JLineTerminal console = JLineTerminal.over(terminal, List.of())) { - // As in a session: the reader owns the screen before anything is drawn into the block. + // As in a session: the reader owns the screen before anything is drawn into the block. The + // pause is not decoration -- the reader thread starts the next prompt as soon as one + // returns, and clearing while that is in flight makes what JLine emits depend on which of + // the two got there first. This test is about the block coming back, not about that race. console.readLine("ignored"); + Thread.sleep(200); console.status(List.of("a distinctive state row")); int before = screen().length(); diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java index 07ce227f2..e481b3b5b 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java @@ -260,6 +260,58 @@ void aSystemFileThatIsNotThereIsAUsageError() { assertThat(thrown.getMessage(), containsString("--system-file cannot be read")); } + @Test + void aFencedBlockIsOneMessageWithItsNewlinesKept() throws Exception { + // A console sends on Enter, so a stack trace pasted into the prompt would become several + // questions. Between two fences it is one. + ScriptedBackend backend = new ScriptedBackend((call, request) -> ScriptedBackend.textTurn("ok")); + try (OpenAiCompatServer server = server(backend)) { + AgentOptions options = AgentOptions.parse(new String[] { + "--base-url", "http://127.0.0.1:" + server.getPort() + "/v1", "--workspace", workspace.toString() + }); + + LocalAgent.run( + options, + new StringReader( + LocalAgent.BLOCK_FENCE + "\nline one\nline two\n" + LocalAgent.BLOCK_FENCE + "\n/exit\n"), + new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8), + new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8)); + } + assertThat("one question, not two", backend.requests(), hasSize(1)); + String asked = backend.requests() + .get(0) + .path("messages") + .get(1) + .path("content") + .asText(); + assertThat(asked, containsString("line one")); + assertThat(asked, containsString("line two")); + assertThat( + "the newline between them survived", + asked.contains("line one") && asked.indexOf("line two") > asked.indexOf("line one"), + is(true)); + assertThat("and the fences are not part of it", asked.contains("\"\"\""), is(false)); + } + + @Test + void endOfInputInsideABlockEndsTheSession() throws Exception { + ScriptedBackend backend = new ScriptedBackend((call, request) -> ScriptedBackend.textTurn("never")); + try (OpenAiCompatServer server = server(backend)) { + AgentOptions options = AgentOptions.parse(new String[] { + "--base-url", "http://127.0.0.1:" + server.getPort() + "/v1", "--workspace", workspace.toString() + }); + + int exit = LocalAgent.run( + options, + new StringReader(LocalAgent.BLOCK_FENCE + "\nhalf a thought\n"), + new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8), + new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8)); + + assertThat("an unfinished block is not sent", exit, is(0)); + } + assertThat(backend.requests(), hasSize(0)); + } + @Test void failedTurnExitsNonZero() throws Exception { ScriptedBackend backend = new ScriptedBackend((call, request) -> ScriptedBackend.textTurn("never")); From 06980c69c81c3a14d2dde72c17be2fc8744f3e39 Mon Sep 17 00:00:00 2001 From: Bernard Ladenthin Date: Thu, 24 Sep 2026 22:21:52 +0200 Subject: [PATCH 35/36] llama-atmosphere-agent: /load reads a saved transcript back as the conversation Save in one session, load in the next: the questions and the answers become messages again, so the model can be asked to carry on rather than to start over. Two decisions worth stating. Only USER and AGENT entries are replayed -- a tool result outside its round is not something a chat template has a place for, and inventing a shape for it would be worse than letting the model call the tool again; tool lines and notes stay in the record. And the parser treats a line without a stamp as a continuation of the entry above it, because an entry is not a line: an answer keeps its newlines when it is written, so reading line by line would turn one answer into several, each nonsense on its own. A file that is not a transcript yields no entries rather than one wrong one -- anything it yielded would be replayed to the model as if it had been said. A path without a directory resolves in the workspace, so /load session.txt finds what /save session.txt wrote. The round trip is tested both ways: the parser on its own (multi-line entries, junk files, the timestamp surviving to the second the file records it in) and end to end across two sessions, where the second request carries the earlier pair before the new question. 159 tests green. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01E2h8gXyyE5UeimkL9vQv9G --- CLAUDE.md | 7 ++ llama-atmosphere-agent/README.md | 8 +++ .../llama/atmosphere/LocalAgent.java | 70 +++++++++++++++++++ .../llama/atmosphere/SlashCommands.java | 2 + .../llama/atmosphere/Transcript.java | 64 +++++++++++++++++ .../net/ladenthin/llama/atmosphere/help.txt | 1 + .../llama/atmosphere/LocalAgentTest.java | 63 +++++++++++++++++ .../llama/atmosphere/TranscriptTest.java | 45 ++++++++++++ 8 files changed, 260 insertions(+) diff --git a/CLAUDE.md b/CLAUDE.md index 88ef7b2c9..4afed0117 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -2328,6 +2328,13 @@ are decisions, not details: tool result and the answer after it regularly share a millisecond and a map would drop one silently; insertion order already is time order. `ToolCallLog` stays as the separate, hard-cut receipt for `/calls` — it answers "did that really run", which prose cannot. + **`/load` replays only `USER` and `AGENT` entries** as messages: a tool result outside its round is + not something a chat template has a place for, and inventing a shape for it would be worse than + letting the model call the tool again. The parser treats a line without a stamp as a continuation + of the entry above it, because an entry is not a line — an answer keeps its newlines when written, + and reading line by line would turn one answer into several. A file that is not a transcript yields + **no** entries rather than one wrong one, since anything it yielded would be replayed to the model + as if it had been said. 8. **Tool calls are carried into the conversation as a text note, and logged for `/calls`.** `LocalAgent.withToolNotes` prefixes each turn's answer in the history with diff --git a/llama-atmosphere-agent/README.md b/llama-atmosphere-agent/README.md index 2ed8012c6..30c5c8ef1 100644 --- a/llama-atmosphere-agent/README.md +++ b/llama-atmosphere-agent/README.md @@ -252,6 +252,14 @@ summarises) and it never carried a timestamp at all. The transcript only grows. means "forget this session" and leaving the text behind would make that untrue. - `/calls` is the shorter, tool-only receipt and is unchanged. +`/load ` (`/resume`) reads one back **as the conversation**: the questions and the answers +become messages again, so the model can be asked to carry on rather than to start over. Tool calls and +session notes stay in the record but are **not** replayed — a tool result outside its round is not +something a chat template has a place for, and inventing one would be worse than letting the model +call the tool again. A path without a directory is resolved in the workspace, so `/load session.txt` +finds what `/save session.txt` wrote; a file that cannot be read, or that is not a transcript, says so +and changes nothing. + It is a list, not a map keyed by the timestamp: a tool result and the answer that follows it regularly land in the same millisecond, and a map would keep one and drop the other silently. Insertion order already is time order. diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java index fc2f9f67a..76d39d070 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java @@ -550,6 +550,69 @@ static void dropLastExchange(List history, String message) { } } + /** + * Read a saved transcript back in, as the conversation and as the record. + * + *

Only the questions and the answers become messages again. Tool calls and session notes are + * kept in the record but not replayed to the model: a tool result out of its round is not + * something any chat template has a place for, and inventing one would be worse than leaving the + * model to call the tool again if it needs to. + * + *

The file is resolved against the workspace when it is not an absolute path, so {@code /load + * session.txt} finds what {@code /save session.txt} wrote. + * + * @param name the file + * @param options the parsed command line, for the workspace + * @param history the conversation, replaced + * @param transcript the record, replaced + * @param terminal where to report + * @return the estimated input tokens of the loaded conversation + */ + private static long load( + String name, + AgentOptions options, + List history, + Transcript transcript, + AgentTerminal terminal) { + java.nio.file.Path file = java.nio.file.Path.of(name.strip()); + if (!file.isAbsolute()) { + file = options.getWorkspace().resolve(file); + } + List loaded; + try { + loaded = Transcript.parse(java.nio.file.Files.readString(file, java.nio.charset.StandardCharsets.UTF_8)); + } catch (java.io.IOException e) { + terminal.line("cannot read " + file + ": " + e.getMessage()); + return estimateTokens(systemPrompt(options), history); + } + if (loaded.isEmpty()) { + terminal.line(file + " holds no transcript entries"); + return estimateTokens(systemPrompt(options), history); + } + history.clear(); + int messages = 0; + for (Transcript.Entry entry : loaded) { + switch (entry.kind()) { + case USER -> { + history.add(ChatMessage.user(entry.text())); + messages++; + } + case AGENT -> { + history.add(ChatMessage.assistant(entry.text())); + messages++; + } + default -> { + // kept in the record, not replayed as a message + } + } + } + transcript.replaceWith(loaded); + transcript.add(Transcript.Kind.NOTE, "loaded " + file); + terminal.line("loaded " + loaded.size() + " entries from " + file + " (" + messages + + " of them replayed to the model)"); + return estimateTokens(systemPrompt(options), history); + } + static boolean usesFullTerminal(AgentOptions options, boolean interactive) { return interactive && !options.isPlain(); } @@ -684,6 +747,13 @@ private static long handleCommand( terminal.line("(history cleared)"); } case CLS -> terminal.clearScreen(); + case LOAD -> { + if (!command.hasArguments()) { + terminal.line("say which file: /load "); + return inputTokens; + } + return load(command.arguments(), options, history, transcript, terminal); + } case SAVE -> { try { java.nio.file.Path written = transcript.save( diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/SlashCommands.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/SlashCommands.java index 6b0c56531..bc8e2949c 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/SlashCommands.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/SlashCommands.java @@ -36,6 +36,8 @@ public enum Command { SAVE("/save", "/transcript"), /** Ask the last question again, without the answer that came back. */ RETRY("/retry", "/again"), + /** Read a saved transcript back in as the conversation. */ + LOAD("/load", "/resume"), /** Summarize the history and continue with the summary; the argument steers the summary. */ COMPACT("/compact"), /** Keep working on one task until it is done; see {@link TaskLoop}. */ diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/Transcript.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/Transcript.java index e27484a3a..c03cc9603 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/Transcript.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/Transcript.java @@ -37,6 +37,10 @@ public final class Transcript { private static final DateTimeFormatter FILE_STAMP = DateTimeFormatter.ofPattern("yyyy-MM-dd_HH-mm-ss"); + /** The shape {@link #format} writes: {@code [stamp] kind: text}. */ + private static final java.util.regex.Pattern HEAD = + java.util.regex.Pattern.compile("\\[(\\d{4}-\\d{2}-\\d{2} \\d{2}:\\d{2}:\\d{2})\\] (\\w+): (.*)"); + /** Who said it. */ public enum Kind { /** What the user typed. */ @@ -164,6 +168,66 @@ public Path save(Path directory, @Nullable String name) throws IOException { return file; } + /** + * Read back a transcript that was written by {@link #save} or by {@code --transcript}. + * + *

A line that does not start with a stamp belongs to the entry above it: an answer keeps its + * newlines when it is written, so an entry is not the same thing as a line. Reading line by line + * would turn one answer into several, each of them nonsense on its own. + * + *

Anything before the first stamped line is ignored rather than guessed at — a file that is not + * a transcript yields no entries instead of one wrong one. + * + * @param text the file content + * @return the entries, oldest first + */ + public static List parse(String text) { + List parsed = new java.util.ArrayList<>(); + StringBuilder pending = new StringBuilder(); + LocalDateTime at = null; + Kind kind = null; + for (String line : text.split("\\r?\\n", -1)) { + java.util.regex.Matcher head = HEAD.matcher(line); + if (head.matches()) { + flush(parsed, at, kind, pending); + at = LocalDateTime.parse(head.group(1), STAMP); + kind = kindOf(head.group(2)); + pending.setLength(0); + pending.append(head.group(3)); + } else if (kind != null) { + pending.append(System.lineSeparator()).append(line); + } + } + flush(parsed, at, kind, pending); + return List.copyOf(parsed); + } + + private static void flush( + List parsed, @Nullable LocalDateTime at, @Nullable Kind kind, StringBuilder pending) { + if (at != null && kind != null && !pending.toString().isBlank()) { + parsed.add(new Entry(at, kind, pending.toString().strip())); + } + } + + private static Kind kindOf(String label) { + for (Kind candidate : Kind.values()) { + if (candidate.label().equals(label)) { + return candidate; + } + } + return Kind.NOTE; + } + + /** + * Replace everything recorded with what was read from a file. + * + * @param loaded the entries to keep + */ + public void replaceWith(List loaded) { + entries.clear(); + entries.addAll(loaded); + } + private void appendLive(Entry entry) { if (liveFile == null) { return; diff --git a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt index a0a0103d1..a9d637efa 100644 --- a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt +++ b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt @@ -16,6 +16,7 @@ Commands (everything else is sent to the model): /cls wipe the screen, keep the conversation (/clear-screen, or Ctrl-L) /save [name] write what was said, with timestamps, into the workspace (/transcript) /retry ask the last question again, dropping the answer (/again) + /load read a saved transcript back in as the conversation (/resume) /exit leave (/quit) The input line sits in the box at the bottom and is there while the agent works: you can type at diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java index e481b3b5b..b6ff7c634 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java @@ -36,6 +36,8 @@ */ class LocalAgentTest { + private static final String NL = System.lineSeparator(); + @TempDir Path workspace; @@ -312,6 +314,67 @@ void endOfInputInsideABlockEndsTheSession() throws Exception { assertThat(backend.requests(), hasSize(0)); } + @Test + void loadingASavedTranscriptMakesItTheConversationAgain() throws Exception { + // Save in one session, load in the next: the model is sent what was said before, so it can be + // asked to carry on rather than to start over. + ScriptedBackend first = new ScriptedBackend((call, request) -> ScriptedBackend.textTurn("the earlier answer")); + try (OpenAiCompatServer server = server(first)) { + AgentOptions options = AgentOptions.parse(new String[] { + "--base-url", "http://127.0.0.1:" + server.getPort() + "/v1", "--workspace", workspace.toString() + }); + LocalAgent.run( + options, + new StringReader("the earlier question" + NL + "/save earlier.txt" + NL + "/exit" + NL), + new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8), + new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8)); + } + + ScriptedBackend second = new ScriptedBackend((call, request) -> ScriptedBackend.textTurn("carrying on")); + ByteArrayOutputStream out = new ByteArrayOutputStream(); + try (OpenAiCompatServer server = server(second)) { + AgentOptions options = AgentOptions.parse(new String[] { + "--base-url", "http://127.0.0.1:" + server.getPort() + "/v1", "--workspace", workspace.toString() + }); + LocalAgent.run( + options, + new StringReader("/load earlier.txt" + NL + "and then?" + NL + "/exit" + NL), + new PrintStream(out, true, StandardCharsets.UTF_8), + new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8)); + } + + assertThat(out.toString(StandardCharsets.UTF_8), containsString("loaded")); + JsonNode messages = second.requests().get(0).path("messages"); + assertThat("system, the earlier pair, and the new question", messages.size(), is(4)); + assertThat(messages.get(1).path("content").asText(), is("the earlier question")); + assertThat(messages.get(2).path("role").asText(), is("assistant")); + assertThat(messages.get(2).path("content").asText(), is("the earlier answer")); + assertThat(messages.get(3).path("content").asText(), is("and then?")); + } + + @Test + void loadingSomethingThatIsNotThereSaysSoAndChangesNothing() throws Exception { + ScriptedBackend backend = new ScriptedBackend((call, request) -> ScriptedBackend.textTurn("ok")); + ByteArrayOutputStream out = new ByteArrayOutputStream(); + try (OpenAiCompatServer server = server(backend)) { + AgentOptions options = AgentOptions.parse(new String[] { + "--base-url", "http://127.0.0.1:" + server.getPort() + "/v1", "--workspace", workspace.toString() + }); + + LocalAgent.run( + options, + new StringReader("/load nowhere.txt" + NL + "still working?" + NL + "/exit" + NL), + new PrintStream(out, true, StandardCharsets.UTF_8), + new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8)); + } + assertThat(out.toString(StandardCharsets.UTF_8), containsString("cannot read")); + assertThat("the session carries on", backend.requests(), hasSize(1)); + assertThat( + "with nothing loaded into it", + backend.requests().get(0).path("messages").size(), + is(2)); + } + @Test void failedTurnExitsNonZero() throws Exception { ScriptedBackend backend = new ScriptedBackend((call, request) -> ScriptedBackend.textTurn("never")); diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/TranscriptTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/TranscriptTest.java index 82a8d79ad..aba997165 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/TranscriptTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/TranscriptTest.java @@ -22,6 +22,8 @@ */ class TranscriptTest { + private static final String LINE_SEPARATOR = System.lineSeparator(); + @TempDir Path directory; @@ -130,6 +132,49 @@ void aFileThatCannotBeWrittenDoesNotEndTheSession() { assertThat("kept in memory regardless", transcript.size(), is(1)); } + @Test + void whatWasWrittenCanBeReadBack() { + Transcript written = new Transcript(); + written.add(Transcript.Kind.USER, "the question"); + written.add(Transcript.Kind.AGENT, "the answer"); + + List read = Transcript.parse(written.render()); + + assertThat(read.size(), is(2)); + assertThat(read.get(0).kind(), is(Transcript.Kind.USER)); + assertThat(read.get(0).text(), is("the question")); + assertThat(read.get(1).kind(), is(Transcript.Kind.AGENT)); + assertThat(read.get(1).text(), is("the answer")); + assertThat( + "the time survives, to the second the file records it in", + read.get(0).at(), + is(written.entries().get(0).at().truncatedTo(java.time.temporal.ChronoUnit.SECONDS))); + } + + @Test + void anAnswerWithNewlinesComesBackAsOneEntry() { + // An entry is not a line: an answer keeps its newlines when it is written, so reading line by + // line would turn one answer into several, each of them nonsense on its own. + Transcript written = new Transcript(); + written.add(Transcript.Kind.AGENT, "first line" + LINE_SEPARATOR + "second line"); + + List read = Transcript.parse(written.render()); + + assertThat(read.size(), is(1)); + assertThat(read.get(0).text(), containsString("first line")); + assertThat(read.get(0).text(), containsString("second line")); + } + + @Test + void aFileThatIsNotATranscriptYieldsNothing() { + // Rather than one wrong entry: a guess here would be replayed to the model as if it were said. + assertThat( + Transcript.parse("just some notes" + LINE_SEPARATOR + "and more") + .size(), + is(0)); + assertThat(Transcript.parse("").size(), is(0)); + } + @Test void clearingForgetsTheSession() { Transcript transcript = new Transcript(); From 64dbb1d775fe60934cb27de05557a9a2fac684f8 Mon Sep 17 00:00:00 2001 From: Bernard Ladenthin Date: Sun, 27 Sep 2026 10:45:13 +0200 Subject: [PATCH 36/36] llama-atmosphere-agent: fix a wrapped line in the help text Left behind by an earlier edit: the sentence about typing during a turn broke mid-clause. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01E2h8gXyyE5UeimkL9vQv9G --- .../main/resources/net/ladenthin/llama/atmosphere/help.txt | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt index a9d637efa..690c4c2b7 100644 --- a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt +++ b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt @@ -19,9 +19,9 @@ Commands (everything else is sent to the model): /load read a saved transcript back in as the conversation (/resume) /exit leave (/quit) -The input line sits in the box at the bottom and is there while the agent works: you can type at -any time. A line typed -during a turn stops it and is sent as the next message. Shift+Tab switches the approval mode. +The input line sits in the box at the bottom and is there while the agent works: you can type +at any time. A line typed during a turn stops it and is sent as the next message. +Shift+Tab switches the approval mode. In manual mode every tool that writes or runs a command asks first: [y]es runs it once, [n]o tells the model the user cancelled it, [a]uto stops asking for the rest of the session.